A processing method and a browser
Patent Information
- Application Number
- CN202610607404.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-07
Smart Images

Figure CN122529067A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a processing method and a browser. Background Technology
[0002] With the development of artificial intelligence (AI) technology, users' needs have expanded beyond simply engaging in dialogue with large models or intelligent agents. This involves users inputting questions or requests within the agent's interface, and the large model or agent providing responses. During this process, users can observe the large model's thought process and its responses through the agent's interface.
[0003] Users want intelligent agents or large models to help them perform more operations. Summary of the Invention
[0004] This application provides a processing method and a browser.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a processing method, the method comprising: A first program, which operates based on a first mode, obtains input information, which represents natural language. The first program can call at least one first model. The first program runs in the second mode; Based on the input information from the second mode response; activate the second mode. in, In response to input information, the execution flow output by the second model includes invoking and displaying the second program, wherein the second model belongs to at least one first model; The process of responding to input information based on the second mode does not affect the input operations obtained by the first program based on the first mode to change the display content of the first program.
[0006] This application provides a browser, which is stored in a storage device and executed by a processor, including: Based on the first mode, the program obtains input information, which represents natural language. The first program can call at least one first model. Run in the second mode; the first mode is different from the second mode. Based on the second mode response input information; in, In response to input information, the execution flow output by the second model includes invoking and displaying the second program, wherein the second model belongs to at least one first model; The process of responding to input information based on the second mode does not affect the input operations obtained by the first program based on the first mode to change the display content of the first program. Attached Figure Description
[0007] Figure 1 This is a flowchart illustrating a processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating an exemplary interactive input channel configuration provided in an embodiment of this application; Figure 3 This is an exemplary structural diagram of the relationship between virtual and physical devices provided in an embodiment of this application; Figure 4 This is an exemplary flowchart illustrating the automatic control mode of a browser as provided in an embodiment of this application; Figure 5 This is a flowchart illustrating an exemplary mode switching method provided in an embodiment of this application; Figure 6 This is a schematic diagram of the browser structure provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0008] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0009] This application provides a processing method, such as... Figure 1 As shown, the process includes the following steps S101 and S104: Step S101: The first program, running based on the first mode, obtains input information, which represents natural language. The first program can call at least one first model.
[0010] In the embodiments of this application, the first program is a program running on an electronic device. For example, the first program can be various types of applications (e.g., browsers, office software, entertainment applications), intelligent agents (e.g., Doubao), or intelligent operating systems (e.g., Windows operating systems). In the embodiments of this application, the first program supports a first mode of operation. When the first program operates in the first mode, it acquires input information, which is natural language input by the user. For example, the user can input the input information in the form of text or voice.
[0011] In the embodiments of this application, the input information is information representing user needs. After receiving the input information, the first program generates a corresponding demand result based on the user needs. For example, the input information could be "Please help me organize the recent price trends of item A".
[0012] In embodiments of this application, the first program may call at least one first model, which may include a large language model and a special function model (e.g., an image processing model, a semantic analysis model, a visual understanding model, and an audio / video processing model).
[0013] Step S102: The first program runs in the second mode.
[0014] In the embodiments of this application, the first program has at least two modes: a first mode and a second mode, and the two modes can run simultaneously. Specifically, the first program can run the first mode and / or the second mode, and the two modes do not conflict when the first program runs both simultaneously.
[0015] For example, when the first program is run in the first mode, the first program can also run in the second mode. The first mode and the second mode do not interfere with each other during operation.
[0016] For example, the second mode is an automatic execution mode. In this mode, the automatic execution mode generates automated execution plans based on a large model or intelligent agent and automatically completes cross-application operations. It utilizes user input in natural language to generate execution flows from relevant models, automatically completing cross-application operations. For instance, a user can input a request in natural language, which will generate an execution flow from the relevant model and automatically launch the corresponding application to complete tasks such as data retrieval (calling a search engine application), content editing (calling an editing application), or task processing.
[0017] In the embodiments of this application, if the first program is an application of various types, then each application embeds an AI agent that can support the first program to run in a second mode. That is, each application not only retains the original functions of the original application, but also takes over some of the system's capabilities (e.g., system tools) and AI functions; then, the original application can support the first program to run in a first mode.
[0018] In the embodiments of this application, if the first program itself is an intelligent agent, then the intelligent agent can support the first program to run in the second mode. For cases where the second mode does not need to be started, the intelligent agent can run based on the first mode, that is, the intelligent agent can run the first mode and the second mode at the same time.
[0019] In the embodiments of this application, if the first program is a smart operating system, then the operating system can run both the first mode and the second mode simultaneously.
[0020] Step S103: Respond to input information based on the second mode; wherein, in response to the input information, the execution flow output by the second model includes initiating and displaying the second program, and the second model belongs to at least one first model; the process of responding to input information based on the second mode does not affect the input operation obtained by the first program running based on the first mode to change the display content of the first program.
[0021] In the embodiments of this application, the input information can be information that needs to be executed by the first program running the second mode. Therefore, when the first program is running the first mode, the input information is received and executed by running the second mode of the first program.
[0022] For example, the input information might be information that needs to be performed across applications, such as needing to launch and display a second program.
[0023] In the embodiments of this application, the first program responds to input information based on a second mode. The execution flow of responding to the input information is determined by the execution flow output by the second model, rather than by manual user operation; that is, the second mode is the mode in which the first program can perform operations based on the model. For example, the second mode can be an automatic execution mode.
[0024] In the embodiments of this application, the second model may be a target model (e.g., a large model or an executable program) or other model that can call the second program and output the execution flow based on the display content of the second program.
[0025] In the embodiments of this application, the second program is a program that runs on an electronic device and is different from the first program. For example, if the first program is an AI browser / intelligent agent / intelligent operating system, the second program may be a chat, payment, ticket booking, or other similar program.
[0026] In embodiments of this application, in response to input information, a first program can invoke a second model for execution. For example, if the input information is "Please help me send a message C to chat partner B in chat software A," then the first program invokes the second model to execute the task indicated by the input information. The second program determines the execution flow required to complete the task indicated by the input information based on the content displayed on the interface, and then executes it. The second model is at least one of the models in the first model.
[0027] In the embodiments of this application, if the input information is "Please help me send a message C to chat object B in chat software A", then when the first program runs the second mode, it is necessary to start chat software A and display the interface of chat software A. That is, when chat software A executes the input information, the second program needs to be called in the execution flow output by the second model.
[0028] In the embodiments of this application, during the process of the first program responding to input information based on the second mode, it does not affect the input operations obtained by the first program based on the first mode to change the display content of the first program. When the first program is running based on the first mode, it also needs to receive input operations based on the display content of the first program. Therefore, when the first program is running in the first mode, it also needs to display the display content corresponding to the first mode. For example, the first mode can be a manual execution mode.
[0029] In the embodiments of this application, the first program needs to display the corresponding display content in both the first mode and the second mode. Furthermore, the process of responding to input information based on the second mode does not affect the input operation obtained by the first program based on the first mode to change the display content of the first program. This indicates that displaying different content in the two modes is necessary and independent of each other.
[0030] In this way, the first program runs in both the first and second modes simultaneously and independently. That is, while the second program is executing a series of processes with the help of the second model, it can also obtain input operations to change the content displayed by the first program based on the first mode, thereby enabling the first program to execute multiple tasks at the same time and improving the flexibility of processing.
[0031] The second mode is the automatic execution mode. In its running state, the automatic execution mode can generate automated execution plans based on large models or intelligent agents and automatically complete cross-application operations. The automated execution plan requires operation of the target program's visual interface. During execution, the generated execution flow often needs to launch and display the target application's window, relying on this window to obtain display status and input focus, thereby completing subsequent interactive operations. By setting different modes, and ensuring that the process of responding to input information based on the automatic execution mode does not affect the input operations obtained by the first program running in the first mode to change the display content of the first program, it indicates that the display and input operations between the two modes are independent of each other. This avoids the automatic execution mode occupying display resources and input channels in the first mode, allowing users to perform other tasks unrelated to the automatic execution mode's input response during the execution of automated tasks. Thus, it is not necessary to wait for the automated task to complete before performing other tasks unrelated to the automated task. By setting the first and second modes to run independently and without interference, users can perform other tasks based on the first mode while executing automated tasks based on the automatic execution mode, and the windows of applications launched by other tasks executed based on the first mode will not affect the display status and input focus of the automatic execution mode. In this way, automated task execution will not lose display resources and input channels, avoiding forced interruption or error reporting of automated tasks. That is, the window that initiates and operates other program windows based on the automatic execution mode can independently obtain the corresponding display state and input focus. Therefore, when the user performs input operations on the display content of a certain program based on the first mode, it will not be affected by the automatic execution mode, thus enabling the user to execute two modes in parallel on the same device and independently complete their respective tasks. In other words, the embodiments of this application solve the problem of the automatic execution mode initiating and displaying the target application window, and relying on this window to obtain the foreground display state and input focus affecting the user's parallel execution of other tasks on the same device. During the execution of an automated task, based on the second mode running state, the user does not need to wait for the automated task to complete before performing other tasks unrelated to the automated task. The embodiments of this application can run the first mode and the second mode simultaneously. Even if the user triggers the execution of other tasks during the automated task, the first mode responds to the other tasks, and the electronic device responds to the window of the application launched by the other task to obtain the foreground display state and input focus. This will not cause the automated task execution to lose display resources and input channels, leading to forced interruption or error reporting of the automated task.Similarly, when the automatic execution process invokes and operates other program windows and needs to obtain the foreground display state and input focus, since the automatic execution process is based on the second mode and runs independently of the first mode, the user's operation to input the display content of a certain program will not be interfered with or even interrupted.
[0032] In some embodiments, the interactive input channel of the first mode is isolated from the interactive input channel of the second mode, so that the interactive input channel in the first mode will not be occupied by the second mode.
[0033] In the embodiments of this application, the first mode and the second mode each have their own interactive input channels, and the two interactive input channels are isolated from each other. This ensures that the interactive input channel in the first mode will not be occupied by the second mode, so that the user can change the input operation of the displayed content of the first program based on the interactive input channel of the first mode.
[0034] In the embodiments of this application, the interactive input channel of the first mode is different from the interactive input channel of the second mode. This may be because the input interactive channels of the first mode and the second mode use different input devices. For example, the interactive input channel may be an interactive input channel built based on input devices such as keyboards and mice.
[0035] For example, when the input devices are a keyboard and a mouse, the interaction input channel in the first mode can be keyboard 1 and mouse 1, and the interaction channel in the second mode can be keyboard 2 and mouse 2. The input channel constructed by keyboard 1 and mouse 1 does not affect the input channel constructed by keyboard 2 and mouse 2. Thus, in the first mode, control is performed based on the interaction input channel constructed by keyboard 1 and mouse 1, and in the second mode, control is performed based on the interaction input channel constructed by keyboard 2 and mouse 2, achieving mutual isolation between the interaction input channels of the first mode and the second mode.
[0036] In the embodiments of this application, the first program runs in a first mode, obtaining input information based on the interactive input channel constructed by keyboard 1 and mouse 1, and the first program runs in a second mode, responding to input information based on the interactive input channel constructed by keyboard 2 and mouse 2.
[0037] For example, the first mode is the manual execution mode, i.e., the browsing mode, and the second mode is the automatic execution mode. The manual execution mode and the automatic execution mode each have their own interactive input channels, and the two interactive input channels are isolated from each other. This ensures that the interactive input channels in different modes are independent of each other, so that they can respond to the input focus independently and avoid being occupied.
[0038] In this way, the interactive input channels in the first mode and the interactive input channels in the second mode are isolated from each other, and different modes can respond to their respective interactive input channels. This ensures that the interactive input channel in the first mode is not occupied by the second mode, and the two interactive input channels respond independently.
[0039] In some embodiments, such as Figure 2 As shown, the interactive input channel of the first mode is isolated from the interactive input channel of the second mode, including the following steps S201 and S202: Step S201: Configure the virtual input device in the second mode to receive keyboard input and / or mouse input driven by automatic interactive operation, and output the automatic interactive operation by the third mode.
[0040] In the embodiments of this application, the second mode is configured with a virtual input device. For example, the virtual input device may include a virtual keyboard and a virtual mouse. Alternatively, the virtual input device may be different in different application scenarios. For example, in a game scenario, a virtual game controller may also be a virtual input device, or in a voice control scenario, a virtual microphone may also be a virtual input device.
[0041] In the embodiments of this application, a virtual input device is configured in the second mode, and each execution operation in the execution flow output by the second model can be implemented through the virtual input device. Exemplarily, the virtual input device can be implemented by installing a driver.
[0042] For example, if the input information is "Please help me send a message C to chat partner B in chat software A", the execution flow output by the second model could be: open chat software A, find chat partner B and open the dialog box, edit message C, and send message C. Then, using the configured virtual input device, the keyboard or mouse driven by the interactive operation driver automatically executes the corresponding process. For example, double-clicking to open chat software A with the virtual mouse, swiping the interface of chat software A to find chat partner B with the virtual mouse, clicking on the area where chat partner B is located to open the dialog box, editing message C in the dialog box using the virtual keyboard, and finally clicking the send button with the virtual mouse to send message C. In this case, the virtual input device executing all the automatic interactive operations is output by the third model, and the first program runs the virtual input device driver to implement the corresponding automatic interactive operations.
[0043] In the embodiments of this application, the automated execution of interactive operations is an operation output by a third model. The third model may be the same as or different from the second model. The third model may determine the corresponding automated execution of interactive operations based on the execution flow output by the second model.
[0044] For example, determining the automated execution operation based on the execution flow could be: opening the dialog box corresponding to chat object B, the corresponding automated execution operation is: clicking the area where chat object B is located; opening chat software A, the corresponding automated execution operation is: double-clicking the icon of chat software A.
[0045] In the automatic execution mode, all operations are performed using virtual input devices without user intervention, which lays the foundation for the first program to achieve multi-mode parallelism.
[0046] Step S202: Configure the physical input device in the first mode, and receive user interactive input and output the response result of the interactive input.
[0047] In the embodiments of this application, the first mode is configured with a physical input device, which serves as a medium for receiving user interactive input. The user inputs interactive input operations through the physical input device.
[0048] In the embodiments of this application, in the first mode, the user inputs interactive input through a physical input device, and the first program receives the interactive input operation from the physical input device and outputs the response result of the interactive input. For example, if the first program is an application embedded with an intelligent agent, then the first program's receipt of the interactive input operation from the physical input device can be performed by the operating system responding to the interactive input operation; that is, the interactive input operation from the physical input device has an operating system response; or, regardless of whether the first program itself is an intelligent agent or executed by an intelligent operating system, the receiving function of the physical input device can be executed by the operating system.
[0049] For example, physical input devices may include physical keyboards and physical mice, or they may be different in different application scenarios. For example, in a gaming scenario, a physical game controller may also be a physical input device, or in a voice input scenario, a microphone may also be a physical input device.
[0050] For example, if a user wants to send a message C to chat partner B in chat software A, the user needs to double-click to open chat software A with the physical mouse, move the physical mouse to find chat partner B, and click to open the dialog box with the physical mouse. The user then edits message C using the physical keyboard and clicks the send message button with the physical mouse to send message C to chat partner B. It can be seen that when the first program runs in the first mode, the user needs to complete the user's needs step by step based on the display content and physical input device of the first program running in the first mode.
[0051] In the embodiments of this application, the first program runs in a first mode, which can receive input information based on an interactive input channel constructed using a physical input device; the first program runs in a second mode, which can execute the automatic interactive operation output by the third model based on an interactive input channel constructed using a virtual input device.
[0052] In this way, isolated interactive input channels can be built using virtual input devices and physical input devices to enable different modes of the first program to run simultaneously and independently.
[0053] In some embodiments, the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, so that the first program can obtain the automatic execution of interactive operations in the second mode, and determine the corresponding instruction or character input based on the encoding information of the automatic execution of interactive operations.
[0054] In the embodiments of this application, the virtual input device is implemented in the form of a driver. Therefore, the operating system can obtain the corresponding automatic interactive operation performed by the first program through the virtual input device. Thus, it is necessary to design that the encoding information of the operation performed on the virtual input device is different from that of the operation performed on the physical input device in order to distinguish the input of different types of input devices.
[0055] In the embodiments of this application, the key code information of the physical input device can be its own key code information, while the key code information of the virtual input device can be different from the key code information of the physical input device. Furthermore, the relationship between the key code information of the virtual input device and the real instructions or characters is pre-stored in the first program. In this way, after the first program receives the automatic execution interactive operation input by the virtual input device, it can determine the corresponding instruction or character input based on the encoding information of the automatic execution interactive operation.
[0056] For example, when the first program starts the second mode, the first program can monitor the steps and status of the automated operation in real time through the intelligent model. When the automated interactive operation is input through the virtual input device, the intelligent model can sense and obtain the corresponding automated interactive operation, and then determine the corresponding instruction or character input to respond.
[0057] In the embodiments of this application, if the first program is an intelligent operating system, then the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, which also enables the operating system to distinguish input in different modes and respond accordingly to achieve mutual isolation. However, when the first program is various types of applications or intelligent agents, the operating system cannot recognize the key code information corresponding to the virtual input device, so it can choose not to recognize it. This also achieves the isolation of the interactive input channel constructed between the virtual input device and the physical input device.
[0058] In the embodiments of this application, different interactive input channels are constructed between the automatic execution mode and the manual execution mode through virtual input devices and physical input devices, respectively. This allows the input of the two modes to be independent and distinguishable, enabling the two modes to run independently without interruption. The basis for achieving independent operation of the two modes is to bind different input devices to different modes and set different key codes for different input devices to distinguish input operations from different devices, thus ensuring the isolation of the interactive input channels.
[0059] Thus, the first program runs in the first mode, receiving input information through an interactive input channel built on a physical input device. The first program runs in the second mode, executing the automatic interactive operation output by the third model through an interactive input channel built on a virtual input device. Because different key codes are set for different input devices, the two interactive input channels respond separately, ensuring complete isolation between the interactive input channels of the first mode and the interactive input channels of the second mode.
[0060] In some embodiments, when performing step S103 above, the following steps may also be performed: the second program outputs the display content of the second program based on the first window which is in a non-visible state, so that the display of the first program running in the first mode is not affected, wherein the display content of the second program is used to input the third model, and the third model infers the automatic execution of interactive operations.
[0061] In the embodiments of this application, when the first program runs in the second mode in response to input information, if the execution flow output by the second model needs to call the second program, the third model needs to infer the automatic execution of interactive operations based on the display content of the second program. Therefore, the second program needs to be displayed. In the first mode, the user's manual operation also needs to be received based on the display content. That is to say, the running status of the first program needs to be displayed in both the first and second modes. Therefore, the second program started in the second mode can be output in an invisible first window. In this way, the display of the first program running in the first mode can be guaranteed not to be affected.
[0062] In the embodiments of this application, by making the first window displayed by the second program invisible relative to the current display screen, the display of the first program running in the current first mode is not affected. An exemplary operation can be to set the display transparency of the first window to 0, so that the first window is invisible to the user, and the display of the first program running by the user in the first mode will not be affected.
[0063] In the embodiments of this application, since the first program can be aware of the second mode when it is running, the display of the first window does not need to be set with focus, but only serves a display function. In this way, the display state and input focus of the first mode will not be affected, so that the third model can obtain the display content of the second program based on the first window. In this way, the requirement to display the display content of the second program in the second mode is met, and the display of the first program running in the first mode will not be affected.
[0064] In the embodiments of this application, if the display of the first program running in the first mode overlaps with the first window, user misoperation may occur on the first window. Since the first program running in the second mode will be monitored, it is possible to respond to the user's input operation on the content displayed in the first mode of the first program running in the first mode, but not to respond to the user's operation on the first window, that is, to intercept the user's operation in the first window.
[0065] In the embodiments of this application, the implementation of the first program responding to input information based on the second mode can be as follows: the first program can call the second model to execute. For example, if the input information is "Please help me send a message C to chat object B in chat software A", then the execution flow output by the first program calling the second model can be: open chat software A, find chat object B and open the dialog box, edit message C, and send message C. Then, how to open chat software A, where chat software A is located in the first window, or where chat object B is located in the chat list, the position of the send button, etc., need to be determined with the help of the display content of the first window. For each execution flow output by the second model, the display position and display type of the first window when executing the flow are determined based on the display content of the first window.
[0066] In the embodiments of this application, during the process of the first program responding to input information based on the second mode, the interactive input channel for responding to input information is isolated from the interactive input channel of the first program based on the first mode. In this way, the automatic execution of interactive operations determined based on the content displayed in the first window in the second mode is executed by the interactive input channel corresponding to the second mode, while the user interactive input based on the displayed content in the first mode is input through the interactive input channel corresponding to the first mode.
[0067] In the embodiments of this application, during the process of the first program responding to input information based on the second mode, the third model, for each execution flow output by the second model, determines the display position and display type of the first window corresponding to the execution of that flow based on the displayed content of the first window. Then, based on the display position and display type (e.g., button, icon, or text), the automatically executed interactive operation (click, double-click, swipe) is determined and executed via a virtual input device. In the first mode, the user runs the content displayed in the first mode, executes the corresponding operation through a physical input device, and outputs the response result of the interactive operation.
[0068] In the embodiments of this application, since the key code information between the virtual input device and the physical input device is different when the first program runs the first mode and the second mode, the first program can obtain the automatic execution of interactive operations based on the key code information and respond within the first window based on the real instructions for the automatic execution of interactive operations.
[0069] Thus, by displaying independently in the first and second modes, the inference operations required in the second mode based on the displayed content are satisfied, thereby improving the flexibility of automated processing.
[0070] In some embodiments, the second mode configures a virtual screen, and the first mode configures a physical screen; the second program outputs the display content of the second program based on the first window which is in an invisible state, including: outputting the display content included in the first window of the second program based on the virtual screen, so that the first window is in an invisible state.
[0071] In the embodiments of this application, the first mode is configured with a physical screen and the second mode is configured with a virtual screen. In this way, different modes are configured with different display screens. When the first program runs the second mode, it needs to call the second program to be displayed, which can be displayed on the virtual screen.
[0072] In the embodiments of this application, different display channels are constructed between the automatic execution mode and the manual execution mode through virtual and physical screens, respectively, so that the display of the two modes can be independent and distinguishable, thereby enabling the two modes to run independently without interruption. The basis for achieving independent operation of the two modes is that different display devices are bound to different modes, ensuring isolation on the display channels.
[0073] In the embodiments of this application, a physical screen is configured in a first mode and a virtual screen is configured in a second mode. The display of the second program that the first program needs to call and display in the second mode on the virtual screen is also invisible to the user, so as not to affect the operation and display of the first program based on the first mode.
[0074] In the embodiments of this application, the content displayed in the first mode is displayed based on the physical screen, and the content displayed in the second mode is displayed based on the virtual screen. The second mode outputs the display content of the first window of the second program based on the virtual screen, so that the first window is in an invisible state. At this time, the virtual screen is equivalent to an extension screen of the physical screen.
[0075] In the embodiments of this application, the virtual screen can form a set of input and output interaction channels in the second mode with the interactive input channel of the second mode. That is, the display content of the first window is displayed through the virtual screen, and the execution flow output by the second model is operated on the displayed content based on the interactive input channel of the second mode. The physical screen can form a set of input and output interaction channels in the first mode with the interactive input channel of the first mode. That is, the content that the first program needs to display in the first mode is displayed through the physical screen, and then the content displayed on the physical screen is operated through the interactive input channel of the first mode.
[0076] In the embodiments of this application, the virtual screen can form an input and output interaction channel in the second mode with the virtual input device, that is, the display content of the first window is displayed through the virtual screen, and the execution flow of the second model output by the virtual input device is used to operate on the displayed content; the physical screen can form an input and output interaction channel in the first mode with the physical input device, that is, the content to be displayed when the first program runs in the first mode is displayed through the physical screen, and then the content displayed on the physical screen is operated through the physical input device.
[0077] In the embodiments of this application, the virtual screen can form an input and output interaction channel in the second mode with the virtual input device. That is, the display content of the first window is displayed through the virtual screen, and the execution flow of the second model output by the virtual input device is used to operate on the displayed content. The physical screen can form an input and output interaction channel in the first mode with the physical input device. That is, the content to be displayed when the first program runs in the first mode is displayed through the physical screen, and the content displayed on the physical screen is operated through the physical input device. Since the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, the two sets of input and output devices are isolated from each other and operate independently.
[0078] For example, such as Figure 3As shown, the physical display (corresponding to the physical screen) 31 is equipped with a physical keyboard / mouse (corresponding to the physical input device) 32, and the virtual display (corresponding to the virtual screen) 33 is equipped with a virtual keyboard / mouse (corresponding to the virtual input device) 34. The physical display 31 and the physical keyboard / mouse 32 are for manual operation by the user, while the virtual display 33 and the virtual keyboard / mouse 34 are for AI operation. The display content 35 of the first mode of the first program is displayed on the physical display 31, and the display content 36 of the second mode of the second program is displayed on the virtual display 33.
[0079] For example, if the first program responds to input information based on the second mode, and the input information is "Please help me send a message C to chat partner B in chat software A", then the first program obtains the display image on the current virtual screen (e.g., by taking a screenshot), and then inputs the third model based on the display image and the input information to obtain the automatic interactive operation (double-clicking the icon of chat software A) input by the third model for the virtual screen. Then, the second program is launched, and then the display image after double-clicking the icon of chat software A on the virtual screen is captured, and then the third model is input again based on the display image and the input information to obtain the automatic interactive operation (clicking the dialog box of chat partner B) input by the third model, until the task indicated by the input information is completed, and the final result is obtained: a message C has been sent to chat partner B in chat software A.
[0080] For example, such as Figure 4 As shown, for Figure 3 The implementation flow of the first program running the second mode is shown below. Taking the AI browser as an example, the first program is discussed here, including the following steps S401 to S405: Step S401: Enter the AI browser's automatic control mode.
[0081] Here, the AI browser operates in second mode to respond to input.
[0082] Step S402: Start the virtual monitor.
[0083] Here, a monitor driver is pre-embedded in the AI browser to implement a virtual monitor.
[0084] For example, the process of launching a virtual display includes steps S4021 to S4027: Step S4021: Create a display object.
[0085] Here, a virtual display object is initialized as the basic container for all subsequent operations.
[0086] Step S4022: Configure monitor parameters.
[0087] Here, you set specific display parameters for the virtual monitor, such as resolution, refresh rate, and color depth.
[0088] Step S4023: Implement the lddCxMonitorCreate interface.
[0089] Here, the virtual display is registered to the system driver layer, completing the initialization.
[0090] Step S4024: Implement the lddCxMonitorArrival interface.
[0091] Here, device connection events are handled, the virtual display driver is loaded, and the device state is initialized.
[0092] Step S4025: The operating system detects the virtual display.
[0093] Here, the operating system actively scans and identifies new devices through the registered interface.
[0094] Step S4026: Create a graphics buffer sequence.
[0095] Here, a sequence of graphics buffers is created based on specific display parameters to manage rendering and display synchronization.
[0096] Step S4027: Activate display.
[0097] Here, the virtual display function is enabled to complete the final output of the graphic data.
[0098] Step S403: Display the AI browser window in automatic control mode as a virtual monitor through the system interface.
[0099] Here, the AI browser window in automated operation mode is displayed on the virtual monitor.
[0100] Step S404: Start the virtual keyboard / mouse.
[0101] Here, a Human Interface Device (HID) driver is pre-embedded in the AI browser to implement a virtual keyboard / mouse.
[0102] For example, the process of launching a virtual keyboard / mouse includes steps S4041 to S4043: Step S4041: Block / filter real keyboard / mouse operations.
[0103] Here, in automatic control mode, the AI browser blocks / filters real keyboard / mouse operations and only accepts virtual keyboard / mouse operations.
[0104] Step S4042: Capture and parse the custom virtual keyboard / mouse operation encoding information.
[0105] Here, after obtaining the virtual keyboard / mouse operations, the custom virtual keyboard / mouse operation encoding information is parsed based on the corresponding relationship.
[0106] Step S4043: Convert into AI browser keyboard / mouse operation commands.
[0107] Here, the parsed custom virtual keyboard / mouse operation encoding information is converted into AI browser keyboard / mouse operation instructions.
[0108] Step S405: Associate the keyboard / mouse operations of the AI browser in automatic control mode with the virtual keyboard / mouse.
[0109] In this way, after the virtual keyboard / mouse operation is associated with the virtual keyboard / mouse, it responds to the content indicated by the corresponding AI browser keyboard / mouse operation command.
[0110] In this way, the first program implements a virtual monitor and virtual keyboard / mouse through an embedded driver. The only resources required for the driver to run are used, without the need for separate configuration of memory, disk, or central processing unit. Furthermore, the virtual monitor and virtual keyboard / mouse perform operations independently of the physical monitor and physical keyboard / mouse, freeing up the physical monitor and physical keyboard / mouse for the user to perform other operations. That is, by building a set of virtual input and output devices, the AI browser that automatically performs tasks can use this set of virtual keyboard, mouse, and monitor, thus not affecting the user's behavior on the real screen and improving the flexibility of processing.
[0111] In some embodiments, when performing step S102 above, the following steps may also be performed: if a target request for calling the window is detected, run the second mode; wherein the window is used to output the display content of the second program.
[0112] In the embodiments of this application, a monitoring program is provided in the first program, which can detect whether a target request for calling the window is generated.
[0113] For example, if the execution flow of input information requires calling a window (e.g., displaying the display content of a second program (launching WeChat to display the WeChat interface)) to know the current display content, and then determine the operation location and operation type based on the display content, then the first program will run the second mode to respond to the input information if the target request for calling the window is detected.
[0114] In the embodiments of this application, the first program is embedded with an intelligent agent or is itself an intelligent agent. The intelligent agent can monitor the entire process of the execution of the first program. If it detects that it needs to be implemented across applications, it is considered that a target request for calling the window has been detected, and the second mode is run.
[0115] For example, if the first program is an intelligent agent such as Doubao or an embedded intelligent agent, and the input information is "Please help me check the current weather", it can be achieved without calling other applications. In this case, the result can be directly output in the first program. In this case, it is considered that no target request to generate a call window has been detected, and the second mode will not be run. Alternatively, if the first program is any type of application, and the traditional functions of the application can achieve the query function, it is also considered that no target request to generate a call window has been detected, and the second mode will not be run.
[0116] In the embodiments of this application, if there are pre-set trigger conditions for the activation of the second mode, such as inputting the input information in an AI assistant embedded in an application, or setting a mode activation shortcut key, then after receiving the corresponding operation, it is considered that a target request to generate a call window has been detected, and the first program runs the second mode to display it through the window.
[0117] Thus, by setting the start conditions for the second mode and binding the start conditions to the window to be invoked, it is shown that the second mode needs to display corresponding content in order to respond to input information, providing a basis for automated execution.
[0118] In some embodiments, such as Figure 5 As shown, the first program can also execute the following steps S501 and S502: Step S501: Based on the second mode response input information process, obtain the viewing request.
[0119] In the embodiments of this application, the viewing request can be a pre-set viewing trigger condition, such as setting a viewing button, a viewing shortcut key, or other triggering methods.
[0120] For example, if a view button or view shortcut is received while the second mode is running, it is considered that a view request has been received.
[0121] Step S502: Respond to the viewing request and output the progress status of the process execution flow based on the second mode response input information in the first mode.
[0122] In the embodiments of this application, if a viewing request is received, the progress status of the process execution flow of the second mode responding to the input information can be output in the first mode, that is, the display content of the virtual screen is displayed on the physical screen so that the user can view the progress status of the process execution flow of the second mode responding to the input information on the physical screen.
[0123] For example, a scheduler can be set in the first program. The trigger condition for this scheduler is whether a viewing request is received. If a request is received, the scheduler is run, and the display content of the first program running in the second mode is displayed on the physical screen. Figure 3 As shown, the state is switched via the intelligent scheduler 37. The scheduler program is stored in the intelligent scheduler.
[0124] For example, the first program enters the AI automatic control mode, switches the AI browser window to the extended screen where the virtual monitor is located, and switches the focus of the physical keyboard and mouse to the virtual keyboard and mouse; if the user needs to temporarily check the control status of the AI browser, the AI browser window is switched back to the physical screen.
[0125] In this way, users can freely switch between the content of the first mode and the second mode, which not only meets the user's requirement for automatic background task operation, but also allows the user to keep track of the task progress status at any time, thus improving the user experience.
[0126] In one embodiment of this application, a first program running in a first mode obtains input information representing natural language; if a target request for calling a window is detected, a second mode is run; wherein, the window is used to output the display content of the second program; the second program outputs the display content of the second program based on the first window which is in a non-visible state, so that the display of the first program running in the first mode is not affected; the display content of the second program is used to input a third model, and the third model infers the automatic execution of interactive operations; the interactive input channel of the first mode and the interactive input channel of the second mode are isolated from each other, so that the interactive input channel in the first mode is not occupied by the second mode.
[0127] In another embodiment of this application, a first program running in a first mode obtains input information representing natural language. The first program can invoke at least one first model. The first program runs in a second mode. Based on the second mode's response to input information, the execution flow output by the second model includes invoking and displaying the second program, where the second model belongs to at least one first model. The display content of the second program's first window is output based on a virtual screen, making the first window invisible so that the display of the first program running in the first mode is not affected. The display content of the second program is used to input a third model, which infers the automatic execution of interactive operations. The second mode configures a virtual input device to receive keyboard and / or mouse input driven by the automatic execution of interactive operations, with the automatic execution of interactive operations being output by the third model. The first mode configures a physical input device and receives user interactive input and outputs the response result of the interactive input.
[0128] In another embodiment of this application, a first program running in a first mode obtains natural language input information, and the first program can call at least one first model; the first program runs in a second mode; it responds to input information based on the second mode; in response to the input information, the execution flow output by the second model includes invoking and displaying the second program, the second model belonging to at least one first model; the display content of the first window of the second program is output based on a virtual screen, making the first window invisible, so that the display of the first program running in the first mode is not affected, the display content of the second program is used to input a third model, and the third model infers the automatic execution of interactive operations; the second mode configures virtual input. The device receives keyboard and / or mouse input driven by automatic interactive operation in a second mode, and the automatic interactive operation is output by a third model; the first mode configures the physical input device, receives user interactive input, and outputs the response result of the interactive input; the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, so that the first program can obtain the automatic interactive operation in the second mode, and determine the corresponding instruction or character input based on the encoding information of the automatic interactive operation; based on the process of responding to input information in the second mode, a viewing request is obtained; in response to the viewing request, the progress status of the process of responding to input information in the second mode is output in the first mode.
[0129] This application provides a processing method in which a first program, running in a first mode, obtains input information, the input information representing natural language, and the first program can call at least one first model; the first program runs a second mode; and responds to input information based on the second mode; wherein, in response to the input information, the execution flow output by the second model includes initiating and displaying the second program, and the second model belongs to at least one first model; the process of responding to input information based on the second mode does not affect the input operation obtained by the first program running in the first mode to change the display content of the first program. In the processing method provided by this application, the interactive input channels in the first mode and the interactive input channels in the second mode are isolated from each other, and different modes can respond to their respective interactive input channels, thereby ensuring that the interactive input channel in the first mode is not occupied by the second mode, and the two interactive input channels respond independently.
[0130] like Figure 6 As shown, this application embodiment provides a browser 61, which is stored in a storage device 62. The browser 61 is executed by a processor 63, including: running based on a first mode, obtaining input information, the input information representing natural language, and a first program capable of calling at least one first model; running a second mode; the first mode being different from the second mode; and responding to input information based on the second mode. The execution flow output by the second model in response to the input information includes invoking and displaying a second program, the second model belonging to at least one first model. The process of responding to input information based on the second mode does not affect the input operations obtained by the first program running based on the first mode to change the display content of the first program.
[0131] In one embodiment of this application, the first mode of the browser 61 is a browsing mode, the second mode of the browser 61 is an intelligent agent automatic execution mode, and the process of responding to the input information based on the second mode is in a non-visible state.
[0132] In the embodiments of this application, the interactive input channel of the first mode is isolated from the interactive input channel of the second mode, so that the interactive input channel in the first mode will not be occupied by the second mode.
[0133] In embodiments of this application, the second mode configures a virtual input device to receive keyboard input and / or mouse input driven by automatically executing interactive operations, wherein the automatically executed interactive operations are output by a third model; the first mode configures a physical input device to receive user interactive input and output the response result of the interactive input.
[0134] In the embodiments of this application, the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, so that the first program can obtain the automatic execution interactive operation in the second mode, and determine the corresponding instruction or character input based on the encoding information of the automatic execution interactive operation.
[0135] In an embodiment of this application, the browser 62 is executed by the processor 63 including: the second program outputs the display content of the second program based on the first window in a non-visible state, so that the display of the first program running in the first mode is not affected, wherein the display content of the second program is used to input the third model, and the automatic execution of the interactive operation is inferred from the third model.
[0136] In embodiments of this application, the second mode configures a virtual screen, and the first mode configures a physical screen; the browser 62 is executed by the processor 63 including: outputting the display content of the first window of the second program based on the virtual screen, so that the first window is in a non-visible state.
[0137] In embodiments of this application, the browser 62 is executed by the processor 63 including: if a target request for a calling window is detected, running the second mode; wherein the window is used to output the display content of the second program.
[0138] In an embodiment of this application, the browser 62 is executed by the processor 63 as follows: a process of responding to the input information based on the second mode to obtain a viewing request; and in response to the viewing request, the progress status of the process of responding to the input information based on the second mode is output in the first mode.
[0139] This application provides a browser in which a first program running in a first mode obtains input information, the input information representing natural language, and the first program can call at least one first model; the first program runs a second mode; and responds to input information based on the second mode; wherein, in response to input information, the execution flow output by the second model includes invoking and displaying the second program, and the second model belongs to at least one first model; the process of responding to input information based on the second mode does not affect the input operation obtained by the first program running in the first mode to change the display content of the first program. In the browser provided in this application, the interactive input channels in the first mode and the interactive input channels in the second mode are isolated from each other, and different modes can respond to their respective interactive input channels, thereby ensuring that the interactive input channel in the first mode is not occupied by the second mode, and the two interactive input channels respond independently.
[0140] like Figure 7As shown, this application embodiment provides an electronic device 7, which includes a storage device 62 and a processor 63. The processor 63 can run a program corresponding to a browser 61 stored in the storage device 62 to achieve: running based on a first mode to obtain input information, the input information representing natural language, the first program being able to call at least one first model; running a second mode; the first mode being different from the second mode; responding to input information based on the second mode; wherein, in response to input information, the execution flow output by the second model includes calling and displaying the second program, the second model belonging to at least one first model; the process of responding to input information based on the second mode does not affect the input operation obtained by the first program running based on the first mode to change the display content of the first program.
[0141] This application provides a computer-readable storage medium storing one or more computer programs, which can be executed by one or more processors to implement the above-described processing method. The computer-readable storage medium can be transient or non-transient.
[0142] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described processing method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium; in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0143] In some embodiments, the storage medium may be a computer-readable storage medium, which may be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), ferromagnetic random access memory (FRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disk-read-only memory (CD-ROM); or it may be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0144] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0145] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0148] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0149] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file containing other programs or data, for example, in one or more scripts within a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files storing one or more modules, subroutines, or code sections). As an example, executable instructions may be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0150] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0151] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not performed.
[0152] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A processing method, the method comprising: A first program, operating based on a first mode, obtains input information, the input information representing natural language, and the first program is capable of calling at least one first model; The first program runs in the second mode; The input information is responded to based on the second mode; in, In response to the input information, the execution flow output by the second model includes invoking and displaying a second program, wherein the second model belongs to the at least one first model; The process of responding to the input information based on the second mode does not affect the input operation obtained by the first program based on the first mode to change the display content of the first program.
2. The processing method according to claim 1, wherein the interactive input channel of the first mode and the interactive input channel of the second mode are isolated from each other, so that the interactive input channel in the first mode will not be occupied by the second mode.
3. The processing method according to claim 2, wherein the interactive input channel of the first mode and the interactive input channel of the second mode are isolated from each other, comprising: The second mode configures a virtual input device to receive keyboard input and / or mouse input driven by automatically executing interactive operations, wherein the automatically executed interactive operations are output by the third model; The first mode configures a physical input device and receives user interaction input and outputs the response result of the interaction input.
4. In the processing method according to claim 3, the key code information corresponding to the virtual input device is different from the key code information corresponding to the physical input device, so that the first program can obtain the automatic execution interactive operation in the second mode, and determine the corresponding instruction or character input based on the encoding information of the automatic execution interactive operation.
5. The processing method according to any one of claims 1 to 4, wherein the step of responding to the input information based on the second mode further comprises: The second program outputs its display content based on the first window which is in a non-visible state, so that the display of the first program running in the first mode is not affected. The display content of the second program is used to input the third model, and the automatic execution of the interactive operation is inferred from the third model.
6. The processing method according to claim 5, wherein the second mode configures a virtual screen, and the first mode configures a physical screen; The second program outputs its display content based on the first window, which is in a non-visible state, including: The first window of the second program is output based on the virtual screen, making the first window invisible.
7. The processing method according to claim 1, wherein the first program runs in a second mode, comprising: If a target request that generates a calling window is detected, run the second mode; The window is used to output the display content of the second program.
8. The processing method according to claim 1, further comprising: Based on the process of responding to the input information in the second mode, a viewing request is obtained; In response to the viewing request, the progress status of the process execution flow that responds to the input information based on the second mode is output in the first mode.
9. A browser, the browser being stored in a storage device, the browser being executed by a processor comprising: The program operates based on a first mode and obtains input information, which represents natural language. The first program can call at least one first model. Run in second mode; The first mode is different from the second mode; The input information is responded to based on the second mode; in, In response to the input information, the execution flow output by the second model includes invoking and displaying a second program, wherein the second model belongs to the at least one first model; The process of responding to the input information based on the second mode does not affect the input operation obtained by the first program based on the first mode to change the display content of the first program.
10. The browser according to claim 9, wherein the first mode of the browser is a browsing mode, the second mode of the browser is an intelligent agent automatic execution mode, and the process of responding to the input information based on the second mode is in a non-visible state.