Interactive interface processing method and device based on Markov process
Through the interactive interface processing method based on Markov process, the problem of insufficient interface understanding ability of the model in complex scenarios is solved, and higher data processing accuracy and task adaptability are achieved.
Patent Information
- Application Number
- CN202510521610.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the interactive interface processing method has poor interface understanding ability in complex scenarios, resulting in inaccurate positioning and insufficient understanding.
The interactive interface processing method based on the Markov process is adopted to model the interface interaction tasks through the Markov decision-making process, determine the state action pair of the decision steps, and calculate the process reward based on the preset step reward mechanism to finally determine the optimal agent strategy.
It improves data processing accuracy and task adaptability, can understand and process interactive interfaces more accurately in complex scenarios, and improves the comprehensive performance and generalization capabilities of the model.
Smart Images

Figure CN120066354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of interface interaction, and in particular, to an interactive interface processing method and device based on a Markov process. Background Art
[0002] Traditional user interface (UI) understanding and positioning training tasks are usually designed around question-and-answer pairs (Q&A), aiming to help the model understand the elements and their functions in the UI page through the Q&A form. In this task design, the overall function of each UI element and page is usually expressed through a series of question-and-answer pairs.
[0003] Although these tasks can help the model achieve certain effects in simple scenarios, they usually cannot handle complex interactive interfaces. Especially when there are multiple similar elements in the interface or the interactivity of the elements is not clear, the model's ability to understand and position the UI is often insufficient.
[0004] Therefore, it can be seen that the interactive interface processing method in the related technology has the technical problem that the model has poor understanding ability of the interface in complex scenarios. Summary of the Invention
[0005] The present invention provides an interactive interface processing method and device based on a Markov process, which are used to solve the defect that the model has poor understanding ability of the interface in the interactive interface processing method in the prior art, and improve the data processing accuracy and task adaptability.
[0006] The present invention provides an interactive interface processing method based on a Markov process, including the following steps.
[0007] Obtain the interface interaction task input by the user; based on the Markov decision process, model the interface interaction task through the current interactive interface understanding model to obtain a plurality of decision steps; determine a plurality of state-action pairs corresponding to each decision step, where the state-action pair is used to simulate the state transition of the user action and the interface interaction; based on a preset step reward mechanism, determine the action selection reward, coordinate positioning reward, and output result reward of each state-action pair of each decision step to obtain the process reward of each state-action pair; use the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; based on the target state-action pairs corresponding to the plurality of decision steps, determine the optimal agent strategy corresponding to the interface interaction task.
[0008] According to an interactive interface processing method based on a Markov process provided by the present invention, the interface interaction task is modeled based on a Markov decision process to obtain a plurality of decision steps, including: determining a state space of a target interface corresponding to the interface interaction task, where the state space includes a plurality of states, and the states are used to represent element attributes of the target interface; determining an action space corresponding to each state in the state space, where the action space includes a plurality of actions, and the actions are used to represent operation types for the element attributes; determining updated states obtained by respectively executing each action in the action space for each state in the state space, and using the set corresponding to the updated states as an updated state space; determining an immediate reward for a decision step corresponding to respectively executing each action in the action space for each state in the state space; determining a state transition probability based on the state space, the action space, and the updated state space; and determining a plurality of decision steps based on the state transition probability and the interface interaction task.
[0009] According to an interactive interface processing method based on a Markov process provided by the present invention, the method further includes: determining, according to a state value function, an expected value of a cumulative immediate reward corresponding to executing a target agent policy in a target state:
[0010] Wherein, represents the state value function, represents the expected value, represents a termination time step, represents a discount factor, represents the immediate reward corresponding to the target agent policy, represents setting the starting state as the target state , represents the target agent policy.
[0011] According to an interactive interface processing method based on a Markov process provided by the present invention, the method further includes: determining, according to an action value function, an expected value of a cumulative immediate reward corresponding to executing a target agent policy after executing a target action in the target state:
[0012] Wherein, represents the action value function, represents the expected value, represents a termination time step, represents a discount factor, represents the immediate reward corresponding to the target agent policy, Indicates setting the starting state to the target state , Indicates setting the starting action to the target action , Indicates the target agent policy.
[0013] According to an interactive interface processing method based on a Markov process provided by the present invention, based on a preset step reward mechanism, determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step, and obtaining a process reward for each state-action pair, including: obtaining a correct operation, a correct coordinate, and a correct output corresponding to the interface interaction task; respectively taking each state-action pair of each decision step as the current state-action pair, and performing the following operations to obtain the process reward of the current state-action pair: determining the action selection, coordinate positioning, and output result of the current state-action pair; determining the action selection reward of the current state-action pair based on the similarity between the correct operation and the action selection; determining the coordinate positioning reward of the current state-action pair based on the similarity between the correct coordinate and the coordinate positioning; determining the output result reward of the current state-action pair based on the similarity between the correct output and the output result; taking the sum of the action selection reward, the coordinate positioning reward, and the output result reward as the process reward of the current state-action pair.
[0014] According to an interactive interface processing method based on a Markov process provided by the present invention, the method further includes: Determining the current state of the interactive interface understanding model for the target decision step of the interface interaction task, performing multiple action samplings on the target decision step to obtain a set of target state-action pairs and an immediate process reward corresponding to the set of target state-action pairs; Determining the current sampling action distribution probability of the set of target state-action pairs under the current agent policy; Determining the historical sampling action distribution probability of the set of target state-action pairs under the historical agent policy; Based on the current sampling action distribution probability and the historical sampling action distribution probability, determining a policy optimization target for the target decision step:
[0015] Wherein, represents the number of samplings of the multiple action samplings; represents the current agent policy in the current state and the probability of generating the th sampling action under the interface interaction task; ; represents the probability of generating the th sampling action under the current state and the interface interaction task; ; represents the advantage function of the th sampling action; represents the th sampling action;
[0016] wherein, represents the process reward of the th sampling action, represents the mean value of the process rewards of multiple action samplings, represents the standard deviation of the process rewards of multiple action samplings.
[0017] The present invention also provides an interactive interface processing device based on a Markov process, including the following modules: an acquisition module for acquiring an interface interaction task input by a user; a modeling module for modeling the interface interaction task based on a Markov decision process through a current interactive interface understanding model to obtain a plurality of decision steps; a state-action module for determining a plurality of state-action pairs corresponding to each decision step, wherein the state-action pair is used to simulate the state transition between a user action and an interface interaction; a reward module for determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step based on a preset step reward mechanism to obtain the process reward of each state-action pair; a determination module for using the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; and a decision module for determining an optimal agent policy corresponding to the interface interaction task based on the target state-action pairs respectively corresponding to the plurality of decision steps.
[0018] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method for processing an interactive interface based on a Markov process as described in any one of the above is implemented.
[0019] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for processing an interactive interface based on a Markov process as described in any one of the above is implemented.
[0020] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the method for processing an interactive interface based on a Markov process as described in any one of the above.
[0021] The method and device for processing an interactive interface based on a Markov process provided by the present invention model an interface interaction task through a Markov decision process to obtain multiple decision steps, which can clearly plan the task process; determine the state-action pairs of each decision step to accurately simulate the interaction state transition between the user and the interface; calculate the process rewards of each state-action pair based on a preset step reward mechanism, which can quantitatively evaluate the value of different actions; select the state-action pair with the maximum process reward as the target state-action pair to optimize each step of the decision; and finally determine the optimal agent strategy based on the target state-action pair, which can efficiently and accurately complete the interface interaction task. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 is a flowchart showing the method for processing an interactive interface based on a Markov process provided by the present invention.
[0024] Figure 2 is a flowchart showing the modeling of an interface interaction task based on a Markov decision process provided by the present invention.
[0025] Figure 3 is a flowchart showing the determination of the process rewards of decision steps provided by the present invention.
[0026] Figure 4 is a general flowchart showing the method for processing an interactive interface based on a Markov process provided by the present invention.
[0027] Figure 5 is a structural diagram showing the device for processing an interactive interface based on a Markov process provided by the present invention.
[0028] Figure 6 is a physical structural diagram showing the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Apparently, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0030] Traditional UI understanding and positioning training task designs are usually constructed around question-and-answer pairs (Q&A), aiming to help the model understand the elements and their functions in the UI page through the Q&A form. In this task design, the overall functions of each UI element and the page are usually expressed through a series of question-and-answer pairs. Traditional design solutions mainly focus on generating simple and basic question-and-answer pairs, such as asking about text content, text position, element category, element position, etc. These tasks establish a one-to-one Q&A relationship, enabling the model to extract the information and functions of specific elements from the visual information of the UI page. For example, for a button, the task might be to ask: "What is the text on the button?" The model needs to generate the correct answer based on the element content in the page, such as "Submit" or "Login". Similarly, for an input box, the task might be to ask: "Where is the input box located?" The model needs to provide the bounding box coordinates of the input box, or "What is the input box hint?" The model would then answer "Please enter your username", etc.
[0031] In addition, traditional task design solutions also include some more complex question-and-answer pairs, such as multi-turn dialogue understanding, especially when the UI page involves interactions among multiple elements or user input. Such tasks require the model to be able to understand the interaction logic between elements and conduct multi-turn Q&A. However, these complex tasks still mainly rely on the functional understanding of single elements and often lack in-depth understanding of the overall page layout, complex relationships between elements, and interaction behaviors.
[0032] Although these traditional task designs have achieved certain effects in the recognition and positioning of UI elements, their limitations are also relatively obvious. Since the tasks are usually designed around individual elements, the model often lacks a comprehensive understanding of the entire UI page. Especially when there are strong interactions between elements or the page structure is complex, traditional task designs may not be able to effectively solve the problems of insufficient page understanding and inaccurate element positioning by the model.
[0033] Traditional UI model training methods usually adopt supervised training or self-supervised training methods. Self-supervised training does not require artificial data labels and trains the model's understanding ability by designing specific tasks. Such tasks include masked text reconstruction, masked image reconstruction, text and image patch alignment, etc. Through these tasks, the model can self-learn the structure of the UI page and the relationships between elements. For example, the masked text reconstruction task requires the model to predict the occluded text content, thus promoting the model's semantic understanding of the page content. Similarly, the masked image reconstruction task helps the model understand the visual layout of the UI page. The advantage of self-supervised training is that it does not need to rely on artificially labeled data labels, but its disadvantage is that usually a large amount of data is required for training to ensure that the model can extract effective features from a large number of unlabeled samples.
[0034] In contrast, supervised training relies on artificially labeled or labeled data generated by large models and usually trains the model by designing specific tasks. These tasks may involve the localization of UI elements, the identification of element types, the extraction of text content, etc. The goal of supervised training is to minimize the difference between the model output and the data labels as much as possible, thereby improving the model's performance on specific tasks. In this way, the model updates its parameters through the backpropagation algorithm to improve prediction accuracy. Although supervised training can achieve higher accuracy on specific tasks, it relies on a large amount of high-quality data labels, and the process of data annotation is usually very time-consuming and expensive.
[0035] It can be seen that the traditional task design is too single, resulting in deficiencies in the model's element understanding and localization. Most task designs focus on the basic UI element recognition and localization, such as identifying text content, locating the positions of UI elements, classifying element types, etc. Although these tasks can help the model achieve certain results in simple scenarios, they usually cannot handle complex UI interfaces, especially when there are multiple similar elements or the interactivity of elements is not clear in the page, the model's understanding and localization capabilities are often insufficient. For example, multiple similar buttons or input boxes may be misjudged or located incorrectly, resulting in the model's inability to accurately understand the page's functions. In addition, traditional task design often ignores the relevance and overall structure between UI elements and fails to deeply consider complex factors such as the interaction logic of the page and the overall layout of functional modules. Therefore, when facing complex and dynamically changing UI interfaces, the model may be unable to correctly understand the user's intention or perform incorrect operations.
[0036] Traditional training methods usually rely on fixed labeled data, which limits the generalization ability of the model. Since fixed manually labeled data or labeled data generated by large models are used during the training process, the model's learning tends to overly rely on these specific labeling trajectories. This approach easily leads to overfitting of the model to specific labeled samples and makes it difficult to adapt to new and unknown UI design styles or layout changes. Moreover, the fixed labeling data trajectories may not be optimal. They are usually based on a specific design or task requirement and may not consider all possible UI element configurations or page structures. Therefore, the model may perform poorly in some special UI design scenarios and cannot effectively handle diverse design requirements, restricting its scalability and adaptability in practical applications.
[0037] In the method for processing an interaction interface based on a Markov process proposed in the present invention, to address the problem of the single UI model training task currently, a multi-task design is proposed to improve the model's ability in different tasks.
[0038] First, to solve the problem of inaccurate element positioning by the model, a question-and-answer task focusing on element positions will be designed. For example, questions asking about the relative positions of two elements are designed to help the model learn the spatial relationships between elements, rather than just the absolute positions of the elements. In this way, the model can better understand the layout and relative positions of elements on the page, thereby improving the accuracy of positioning.
[0039] Second, to address the deficiency of the model in differentiating similar elements, a question-and-answer task focusing on similarity comparison will be designed. For example, questions asking about the differences between two similar elements or finding the element most similar to a certain element are designed. These tasks can help the model improve its differentiating ability when faced with multiple similar elements, enabling it to accurately identify and locate different elements even if they are similar in appearance or function.
[0040] Finally, to address the problem of the model's insufficient overall understanding of the page, a question-and-answer task focusing on element and text alignment will be designed. For example, questions asking whether a certain element matches the text or which piece of text a certain element matches are designed. Such tasks can enhance the model's understanding ability of the page content, help it understand the semantic relationship between elements and text, and further improve the model's overall understanding and layout inference ability of the entire UI page.
[0041] Through the multi-task design, the model can not only perform well in each task but also share the knowledge learned between different tasks, enhancing the overall performance and generalization ability.
[0042] By introducing multi-task learning mechanisms such as position Q&A, similarity comparison Q&A, and element-text alignment Q&A, the model's capabilities in element localization, similarity discrimination, and overall page understanding have been significantly improved. Compared with traditional single-task models, the multi-task mechanism designed in the present invention can share knowledge among tasks, improving the model's comprehensive performance and generalization ability on different tasks.
[0043] Optionally, the Markov process-based interactive interface processing method in the embodiments of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device. Taking the execution of the Markov process-based interactive interface processing method in this embodiment by a server as an example.
[0044] Figure 1 is a flowchart of the Markov process-based interactive interface processing method provided by the present invention, as Figure 1 shown, the method includes the following: Step 101, obtain the interface interaction task input by the user.
[0045] In the embodiments of the present invention, first, the original input of the user is captured through a multi-modal interface (such as touch, voice, buttons, etc.), then standardized processing and task type recognition are performed, and at the same time, context features are constructed by combining the current interface state and historical operation records, and finally, a structured interface interaction task is output.
[0046] Step 102, through the current interactive interface understanding model, model the interface interaction task based on the Markov decision process to obtain multiple decision steps.
[0047] In the embodiments of the present invention, the interactive interface understanding model, that is, the UI understanding model, is a technology based on artificial intelligence, aiming to automatically identify, analyze, and understand the elements (such as buttons, texts, icons, layouts, etc.) and their interaction logics in the interface by analyzing the visual, structural, or code information of the user interface (User Interface, UI).
[0048] In the embodiments of the present invention, during the process of the current interactive interface understanding model understanding the interface interaction task, the interface interaction task can be regarded as a Markov decision process (Markov Decision Process, MDP), where each decision step corresponds to a state transition of the user's interaction with the interface. The Markov decision process model mainly consists of four parts: state space, action space, reward function, and state transition probability.
[0049] Refer to Figure 2 , Figure 2 is a flowchart of modeling the interface interaction task based on the Markov decision process provided by the present invention.
[0050] According to an interactive interface processing method based on a Markov process provided by the present invention, an interface interaction task is modeled based on a Markov decision process to obtain multiple decision steps, including: Step 201, determine the state space of the target interface corresponding to the interface interaction task, where the state space includes multiple states, and the state is used to represent the element attributes of the target interface.
[0051] Step 202, determine the action space corresponding to each state in the state space, where the action space includes multiple actions, and the action is used to represent the operation type for the element attributes.
[0052] Step 203, determine each state in the state space, and obtain the updated state obtained by executing each action in the action space respectively, and use the set corresponding to the updated state as the updated state space.
[0053] Step 204, determine the immediate reward of the decision step corresponding to each action executed by each state in the state space respectively.
[0054] Step 205, determine the state transition probability based on the state space, the action space, and the updated state space.
[0055] Step 206, determine multiple decision steps based on the state transition probability and the interface interaction task.
[0056] In the embodiment of the present invention, the state space defines all possible states of the UI interface at a certain moment. In the understanding of the UI interface, the state space contains the visual information of the interface elements. The target state can be described as , that is, the state of the UI interface at time step t. Since the UI interface consists of a large number of complex interface elements, including text, buttons, menus, input boxes, etc., each interface element can be described by its position, attributes, etc. For example, the content of the input box, whether the button is clickable, whether the page is loaded completely, etc. can all be regarded as part of the state. For a specific UI interface, the target state (state space) can be defined as: .
[0057] Among them, the element attribute represents the state or feature of each interface element in each UI interface.
[0058] In the embodiment of the present invention, the action space defines all possible actions that the agent can take in each state. In the context of the UI interface, the actions may include clicking a button, scrolling the page, selecting a menu item, entering text, etc. Each action can affect the state of the current UI interface. The action space can be expressed as , i.e., at time step t, the set of all actions that the agent can take. Each target action will result in a change in the target state , that is, a transition to a new state (i.e., an updated state) .
[0059] For operations in the UI interface, a finite action space (operation space) can be defined, which contains all possible operation types. For example: .
[0060] In the embodiments of the present invention, the immediate reward for each decision step corresponding to each action in the action space is determined by the reward function for each state in the state space; that is, the reward function is used to evaluate the effect of the agent performing a certain action in a certain state, reflecting the quality of its behavior. In UI interface understanding, the reward can be designed to evaluate the quality of the behavior according to the information generated by the agent for interacting with the interface. The reward function can be expressed as: , where is the immediate reward after performing the target action in the target state .
[0061] In the Markov decision process, the probability of state transition represents the probability of transitioning from the current state (target state) to the next state (updated state) after performing the target action . In the UI interface, state transitions may be affected by various factors, such as user input, interface loading speed, external system response, etc.
[0062] State transitions can be modeled using a probability function: , which describes the likelihood of the system transitioning to the updated state after performing the target action in the target state .
[0063] The goal of the Markov decision process is to maximize the cumulative reward by optimizing the policy. In the context of UI interface understanding, the agent continuously interacts with the interface, obtains multiple decision steps based on the state transition probability and the interface interaction task, and thus learns an optimal target agent policy , that is, the rule for outputting the best action in each state.
[0064] In the embodiments of the present invention, the agent policy refers to an optimal set of rules generated based on the Markov decision process modeling, which guides the agent to perform actions in the interface interaction task.
[0065] The agent policy is comprehensively evaluated through dynamic programming and a multi-dimensional reward mechanism (action semantic matching, operation positioning accuracy, result feedback), and the action sequence that maximizes the cumulative reward is selected in each decision-making step. The agent policy not only ensures the efficiency of task execution (such as minimizing the number of steps and precise operations), but also can adapt to different interface states and user intentions through the reinforcement learning mechanism, solving the problem of insufficient generalization of traditional rule engines in complex scenarios, and finally achieving the goal of automated interaction across interface layouts.
[0066] Through the embodiments of the present invention, by establishing a complete Markov decision process model, the intelligent processing of interface interaction tasks is realized. It can automatically learn the optimal interaction policy and adapt to different interface interaction tasks and interface changes, significantly improving the interface understanding ability and decision-making reliability.
[0067] According to an interface processing method based on a Markov process provided by the present invention, the above method further includes: Determine the expected value of the cumulative immediate reward corresponding to the execution of the target agent policy in the target state according to the state value function:
[0068] Among them, represents the state value function, represents the expected value, represents the termination time step, represents the discount factor, represents the immediate reward corresponding to the target agent policy, represents setting the starting state as the target state , represents the target agent policy.
[0069] In the embodiments of the present invention, the state value function is used to represent the expectation of the cumulative immediate reward that can be obtained when acting according to the target agent policy in the target state .
[0070] In some embodiments, obtain the interface interaction task and the corresponding target interface, and clarify the current target state (such as the set of interface element attributes); determine the target agent policy for completing the interface interaction task (such as "click the button and enter text"); select the discount factor according to the nature of the task (such as emphasizing long-term rewards); determine the termination time step (for example, the task is completed, failed, or the maximum number of steps is reached).
[0071] Through the embodiments of the present invention, by calculating the state value function, the long-term benefits of different decision-making steps in a specific state can be quantitatively evaluated, and the immediate reward and future benefits are balanced through the discount factor to avoid short-sighted decision-making.
[0072] According to an interactive interface processing method based on a Markov process provided by the present invention, the above method further includes: Determine the expected value of the cumulative immediate reward corresponding to the execution of the target agent policy after executing the target action in the target state according to the action value function:
[0073] Wherein, represents the action value function, represents the expected value, represents the termination time step, represents the discount factor, represents the immediate reward corresponding to the target agent policy, represents setting the starting state as the target state , represents setting the starting action as the target action , represents the target agent policy.
[0074] In the embodiments of the present invention, the action value function is used to represent the expected value of the cumulative immediate reward that can be obtained when acting according to the target agent policy after executing the target action in the target state . After that, when acting according to the target agent policy
[0075] In some embodiments, obtain the interface interaction task and the corresponding target interface, and clarify the current target state (such as the set of interface element attributes); determine the target action expected to be executed in the target state; determine the target agent policy for completing the interface interaction task (such as "click the button and enter text"); select the discount factor according to the nature of the task (such as emphasizing long-term rewards); determine the termination time step (for example, the task is completed, failed, or the maximum number of steps is reached).
[0076] Through the embodiments of the present invention, the long-term benefits of executing a specific action in a specific state can be evaluated through the action value function, providing a quantitative basis for decision optimization.
[0077] Through multiple reinforcement learning algorithms in the embodiments of the present invention, the agent can continuously optimize the decision-making steps, so that the actions selected in each state can maximize the cumulative reward, thereby optimizing the understanding and interaction of the UI interface.
[0078] Step 103: Determine multiple state-action pairs corresponding to each decision step, where the state-action pairs are used to simulate the state transition of user actions and interface interactions.
[0079] In the embodiments of the present invention, the state-action pairs are used to describe how to update the state and return feedback after performing an action in a specific interface state. Specifically, it includes: State (s): The feature representation of the current interface (such as UI components, user input, system response, etc.).
[0080] Action (a): An executable operation (such as clicking a button, entering text, swiping the screen, etc.).
[0081] State transition (from s to s'): The probability distribution of entering a new state (updating the state) after performing the action.
[0082] Immediate reward (R): Measure the effectiveness of the action in completing the task (for example, +1 for successfully submitting a form, -0.5 for an incorrect operation).
[0083] Step 104: Based on a preset step reward mechanism, determine the action selection reward, coordinate positioning reward, and output result reward for each state-action pair of each decision step, and obtain the process reward for each state-action pair.
[0084] Currently, supervised learning often relies on labeled tags and is difficult to adapt to new UI layouts or task changes. Therefore, reinforcement learning based on reward feedback is proposed to improve the generalization ability of the model. Reinforcement learning can enable the model to learn how to handle different scenarios without explicit labels through a feedback-based approach, which is particularly suitable for optimizing UI interaction operations. The task of the UI model is to perform appropriate operations on the UI interface according to the user's instruction content. Usually, each operation not only includes the understanding of the current problem but also specific action selection and action parameters. For example, in different UI layouts or task scenarios, the model needs to determine whether to click, swipe, enter text, or select an option, and the execution result of each operation will affect the subsequent task execution and the final effect. The advantage of reinforcement learning is that it can assign an immediate reward signal to each operation, so that the model no longer adjusts the strategy solely based on the success or failure of the final result. Through a dynamic reward mechanism, the model can real-time perceive the quality of each operation, whether it is correct or not, whether it meets the user's expectations, or the execution efficiency. The reward signal can provide timely feedback. Such a mechanism effectively avoids the drawback of relying on the final success result to measure the model's performance, and at the same time provides richer learning signals for the model in complex UI tasks.
[0085] Reference Figure 3 ,Figure 3 It is a schematic flow chart of the process reward for determining the state-action pair provided by the present invention.
[0086] According to an interactive interface processing method based on a Markov process provided by the present invention, based on a preset step reward mechanism, determine the action selection reward, coordinate positioning reward, and output result reward for each state-action pair of each decision step, and obtain the process reward for each state-action pair, including: Step 301, obtain the correct operation, correct coordinates, and correct output corresponding to the interface interaction task.
[0087] Step 302, respectively take each state-action pair of each decision step as the current state-action pair, and perform the following operations to obtain the process reward of the current state-action pair.
[0088] Step 3021, determine the action selection, coordinate positioning, and output result of the current state-action pair.
[0089] Step 3022, determine the action selection reward of the current state-action pair based on the similarity between the correct operation and the action selection.
[0090] Step 3023, determine the coordinate positioning reward of the current state-action pair based on the similarity between the correct coordinates and the coordinate positioning.
[0091] Step 3024, determine the output result reward of the current state-action pair based on the similarity between the correct output and the output result.
[0092] Step 3025, take the sum of the action selection reward, coordinate positioning reward, and output result reward as the process reward of the current state-action pair.
[0093] In the embodiment of the present invention, the intelligent agent outputs three parts of content for each instruction (i.e., the interface interaction task), including the thinking about the user instruction and the current UI interface, the action to be taken next, and the specific parameters of the action. A corresponding reward signal is given to each part of the content, thereby solving the problem of sparse reward signals caused by relying solely on result feedback.
[0094] In the embodiment of the present invention, the preset step reward mechanism specifically includes the following content: Output format: For each part that meets the requirements, the reward signal +2.
[0095] Thinking: Think about the operation process required to achieve the goal on this page.
[0096] Action: Define the <action selection> reward method: If the output action does not belong to the defined operation space, the reward signal -10.
[0097] If the output action belongs to the operation space but is irrelevant to the correct operation, the reward signal is -5.
[0098] If the output action belongs to the operation space, is related to the operation action but not completely correct (e.g., the output is a single click but the correct answer is a double click), the reward signal is +5.
[0099] If the output action belongs to the operation space and is completely correct, the reward signal is +10.
[0100] Coordinate positioning: At what position to execute <action> Define the reward method for <coordinate positioning>: For single click, double click and long press operations, the agent should output the specific operation coordinates.
[0101] If the output coordinates are not within the bounding box of the correct answer and not within the horizontal and vertical regions where the bounding box is located, the reward signal is -10.
[0102] If the output coordinates are not within the bounding box of the correct answer but within the horizontal or vertical region where the bounding box is located, the reward signal is -5.
[0103] If the output coordinates are within the bounding box of the correct answer, the reward signal is +5.
[0104] If the distance between the output coordinates and the center of the bounding box of the correct answer is less than 50, the reward signal is +10.
[0105] If the distance between the output coordinates and the center of the bounding box of the correct answer is greater than 50 and less than 100, the reward signal is +8.
[0106] If the distance between the output coordinates and the center of the bounding box of the correct answer is greater than 100 and less than 200, the reward signal is +6.
[0107] If the distance between the output coordinates and the center of the bounding box of the correct answer is greater than 200 and less than 400, the reward signal is +2.
[0108] For swipe and long press swipe operations, the starting coordinates and ending coordinates of the operation should be output. The rewards for the two coordinates refer to the above criteria.
[0109] Output result: For the input operation, the agent should output the specific text content, and the reward signal can be given by calculating the cosine similarity between the output text and the correct answer.
[0110] Define the reward method for <output result>: If the cosine similarity between the output text and the correct answer is less than 0.2, the reward signal is +0.
[0111] If the cosine similarity between the output text and the correct answer is greater than 0.2 and less than 0.4, the reward signal is +2.
[0112] If the cosine similarity between the output text and the correct answer is greater than 0.4 and less than 0.6, the reward signal is +4.
[0113] If the cosine similarity between the output text and the correct answer is greater than 0.6 and less than 0.8, the reward signal is +6.
[0114] If the cosine similarity between the output text and the correct answer is greater than 0.8, the reward signal is +8.
[0115] Through the embodiments of the present invention, by introducing reinforcement learning based on reward feedback, higher flexibility and adaptability can be brought to UI operation tasks. Especially when facing unknown UI changes or tasks, reinforcement learning can effectively improve the generalization ability of the model by continuously exploring and adjusting the strategy through feedback, so as to better cope with the increasingly complex actual application scenarios.
[0116] Step 105: Take the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step.
[0117] In each decision step, select the state-action pair that maximizes the process reward as the target state-action pair for this decision step.
[0118] In the embodiments of the present invention, through process reward feedback, the reinforcement learning model continuously adjusts its strategy, that is, gradually learns to determine the optimal state-action pair as the target state-action pair under the given UI interface state. To ensure effective learning, the balance between exploration and exploitation is crucial. Exploration means trying different interaction actions under uncertain circumstances to discover which behaviors can maximize future rewards; while exploitation is to select the most effective actions based on the knowledge already obtained. As the number of interactions increases, it will gradually tend to exploit known successful experiences while maintaining a certain degree of exploration to cope with changes in UI interface design or task objectives.
[0119] In the present invention, the process is defined as three parts: <action selection>, <coordinate positioning>, <output result>. When the corresponding process is completed, the corresponding special marker is output, and the corresponding reward method is used to calculate the obtained reward. The final process reward of each state-action pair is the sum of the rewards of the three parts.
[0120] Step 106: Based on the target state-action pairs corresponding to the multiple decision steps, determine the optimal agent strategy corresponding to the interface interaction task.
[0121] In the embodiment of the present invention, based on the target state-action pairs of all decision steps, a final policy is generated through a policy optimization algorithm (such as policy iteration, Q-Learning, PPO).
[0122] Through the above steps of the embodiment of the present invention, a Markov decision process is used to model the interface interaction task to obtain multiple decision steps, which can clearly plan the task process; determining the state-action pairs of each decision step can accurately simulate the interaction state transition between the user and the interface; calculating the process reward of each state-action pair based on a preset step reward mechanism can quantitatively evaluate the value of different actions; selecting the state-action pair with the largest process reward as the target state-action pair can optimize the decision of each step; finally, determining the optimal agent policy based on the target state-action pair can efficiently and accurately complete the interface interaction task.
[0123] According to an interface interaction processing method based on a Markov process provided by the present invention, the above method further includes: Determine the current state of the interface interaction understanding model for the target decision step of the interface interaction task, perform multiple action samplings on the target decision step to obtain a set of target state-action pairs and the immediate process rewards corresponding to the set of target state-action pairs; Determine the current sampling action distribution probability of the set of target state-action pairs under the current agent policy; Determine the historical sampling action distribution probability of the set of target state-action pairs under the historical agent policy; Based on the current sampling action distribution probability and the historical sampling action distribution probability, determine the policy optimization objective of the target decision step:
[0124] Among them, represents the number of sampling times of multiple action samplings; represents the current agent policy in the current state and the interface interaction task to generate the th sampling action probability; represents the historical agent policy in the current state and the interface interaction task to generate the th sampling action probability; represents the advantage function of the th sampling action;
[0125] Among them, represents the The process reward of a sampling action represents the mean of the process rewards of multiple action samplings and represents the standard deviation of the process rewards of multiple action samplings
[0126] In an embodiment of the present invention, during the optimization process of the model, a reinforcement learning algorithm is used to evaluate the value of the state-action pair
[0127] For the UI interface (I) and the interface interaction task Question (Q), the current agent policy of the interface interaction understanding model is Model (M) (i.e., the current interface interaction understanding model), and the historical agent policy of the interface interaction understanding model is Model old (M old ) (i.e., the historical interface interaction understanding model)
[0128] For any decision step in the interface interaction task Q, use the historical interface interaction understanding model M old to directly sample a set of answers (i.e., historical state-action pairs): {o 1 , o 2 , o 3 , o 4 ,..., o G}, and obtain the reward score corresponding to each answer based on the above reward calculation method: {r 1 , r 2 , r 3 , r 4 ,..., r G}. And calculate the reinforcement learning optimization objective:
[0129] wherein, represents the number of sampling times of multiple action samplings; represents the current agent policy in the current state , the interface interaction task to generate the th sampling action probability; represents the historical agent policy in the current state , the interface interaction task to generate the th sampling action probability; represents the advantage function of the th sampling action;
[0130] wherein, represents the The process reward for one sampling action represents the mean of the process rewards for multiple action samplings represents the standard deviation of the process rewards for multiple action samplings
[0131] Reference Figure 4 , Figure 4 is the overall flowchart of the interactive interface processing method based on the Markov process provided by the present invention, which includes: multi-task data construction, UI interface MDP modeling, designing a step-level reward function, and optimizing the model policy by reinforcement learning
[0132] Through the embodiments of the present invention, the optimization process through reward feedback is self-adaptive and can handle the problems brought about by UI interface changes. For example, when the UI design is updated or changed, it can continue to adjust its policy by interacting with the new interface to adapt to the new interface layout or interaction logic. Therefore, reinforcement learning can ensure that the UI understanding model can continuously improve the interaction efficiency and task completion rate when facing a constantly changing interface. Finally, after multiple rounds of optimization, it can complete the task with the fewest steps and the optimal policy, thereby realizing the efficient understanding and operation of the UI interface
[0133] Through the above embodiments of the present invention, to improve the ability of the UI understanding model in positioning and similarity discrimination, a multi-task training framework is designed. By designing specific training objectives for different tasks, the model can optimize multiple tasks simultaneously in the same training process, such as positioning of interface elements, discrimination of content similarity, etc., thereby improving the comprehensive performance of the model. Introduce a process reward mechanism, and by designing different reward signals for different task stages, solve the problem of sparse reward signals in the current result reward model. In this way, the model can obtain immediate feedback at each step of task completion, optimize its decision-making process, and improve the understanding ability and operation efficiency of the UI interface
[0134] Through the multi-task design and reinforcement learning method based on reward feedback of the present invention, compared with the prior art, it has the following outstanding effects Improve data processing accuracy and task adaptability By introducing multi-task learning mechanisms such as position question answering, similarity comparison question answering, and element-text alignment question answering, the model's capabilities in element localization, similarity discrimination, and overall page understanding have been significantly improved. Compared with traditional single-task models, the multi-task mechanism designed in the present invention can share knowledge between tasks, improving the model's comprehensive performance and generalization ability on different tasks. The reinforcement learning model based on reward feedback no longer relies on explicit annotation labels, but optimizes the model strategy in real time through a dynamic reward mechanism. Compared with traditional supervised learning methods, the present invention has higher adaptability in the face of unknown UI layouts and task changes, and can significantly improve the model's execution ability in complex scenarios.
[0135] Reduce resource consumption and development costs: The reinforcement learning method reduces the dependence on large-scale manually annotated data through reward-based policy optimization. This method can effectively learn under unsupervised or weakly supervised conditions, thereby reducing the data annotation cost. The knowledge sharing mechanism of the multi-task design enables the model to quickly transfer existing knowledge in new tasks or new layout scenarios, significantly shortening the development cycle and saving model training time and computing resources.
[0136] Improve user interaction experience and operation efficiency: The reinforcement learning model can optimize the UI operation strategy in real time through the reward feedback mechanism, significantly improving the accuracy and efficiency of user interaction operations and reducing the occurrence of misoperations. Whether it is clicking, swiping, or input operations, the model can provide better solutions in dynamic scenarios. The reinforcement learning method of the present invention enables the model to dynamically adapt to different UI layouts and task requirements, ensuring that users can still obtain a smooth and efficient operation experience when facing complex or changing interfaces.
[0137] Enhance the stability and robustness of the model: Multi-task learning improves the stability of the model in complex scenarios by modeling multi-dimensional features such as relationships between elements, similarities, and semantic alignments. Even when the appearance or layout of page elements changes, the model can still maintain high accuracy. The immediate feedback mechanism of reinforcement learning can correct incorrect operations in a timely manner, ensuring the robustness and reliability of the model, especially when the task difficulty or environmental complexity increases.
[0138] In summary, the present invention is significantly superior to the prior art in terms of data processing accuracy, resource savings, task adaptability, user experience, and model robustness, and can provide a new solution for the intelligent operation and optimization of UI interfaces.
[0139] The following describes the interactive interface processing device based on the Markov process provided by the present invention. The interactive interface processing device based on the Markov process described below can be correspondingly referred to the interactive interface processing method based on the Markov process described above.
[0140] Reference Figure 5 , Figure 5 is a schematic structural diagram of the interactive interface processing device based on the Markov process provided by the present invention.
[0141] An acquisition module 501, configured to acquire an interface interaction task input by a user; A modeling module 502, configured to model the interface interaction task based on the Markov decision process through a current interactive interface understanding model, and obtain a plurality of decision steps; A state-action module 503, configured to determine a plurality of state-action pairs corresponding to each decision step, where the state-action pair is used to simulate the state transition between a user action and an interface interaction; A reward module 504, configured to determine an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step based on a preset step reward mechanism, and obtain a process reward for each state-action pair; A determination module 505, configured to use the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; A decision module 506, configured to determine an optimal agent policy corresponding to the interface interaction task based on the target state-action pairs respectively corresponding to the plurality of decision steps.
[0142] Specifically, the above-mentioned interactive interface processing device based on the Markov process provided by the present invention can implement all the method steps implemented by the above-mentioned interactive interface processing method embodiment based on the Markov process, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiment are not specifically described herein again.
[0143] Figure 6 is a schematic physical structure diagram of the electronic device provided by the present invention, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute an interactive interface processing method based on a Markov process. The method includes: obtaining an interface interaction task input by a user; modeling the interface interaction task based on a Markov decision process through a current interactive interface understanding model to obtain a plurality of decision steps, where each decision step corresponds to a state transition of a user interacting with the interface; determining a plurality of state-action pairs corresponding to each decision step, where the state-action pair is used to simulate the state transition of a user action and the interface interaction; based on a preset step reward mechanism, determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step in the plurality of decision steps to obtain a process reward for each state-action pair of each decision step; using the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; and determining an optimal agent policy corresponding to the interface interaction task based on the target state-action pairs respectively corresponding to the plurality of decision steps.
[0144] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0145] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the Markov process-based interactive interface processing method provided by the above-mentioned various methods. The method includes: obtaining an interface interaction task input by a user; modeling the interface interaction task based on a Markov decision process through a current interactive interface understanding model to obtain a plurality of decision steps, where each decision step corresponds to a state transition of a user interacting with the interface; determining a plurality of state-action pairs corresponding to each decision step, where the state-action pair is used to simulate the state transition of a user action and the interface interaction; based on a preset step reward mechanism, determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step among the plurality of decision steps to obtain a process reward for each state-action pair of each decision step; taking the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; and determining an optimal agent strategy corresponding to the interface interaction task based on the target state-action pairs respectively corresponding to the plurality of decision steps.
[0146] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the Markov process-based interactive interface processing method provided by the above-mentioned various methods. The method includes: obtaining an interface interaction task input by a user; modeling the interface interaction task based on a Markov decision process through a current interactive interface understanding model to obtain a plurality of decision steps, where each decision step corresponds to a state transition of a user interacting with the interface; determining a plurality of state-action pairs corresponding to each decision step, where the state-action pair is used to simulate the state transition of a user action and the interface interaction; based on a preset step reward mechanism, determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step among the plurality of decision steps to obtain a process reward for each state-action pair of each decision step; taking the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; and determining an optimal agent strategy corresponding to the interface interaction task based on the target state-action pairs respectively corresponding to the plurality of decision steps.
[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Markov process-based interactive interface processing method, characterized in that: include: Interface interaction tasks to obtain user input; Through the current interactive interface understanding model, the interface interaction task is modeled based on the Markov decision process to obtain multiple decision steps; Determine a plurality of state-action pairs corresponding to each decision step, wherein the state-action pairs are used to simulate state transitions of user actions and interface interactions; Based on a preset step reward mechanism, determine the action selection reward, coordinate positioning reward and output result reward for each state-action pair of each decision step, and obtain the process reward for each state-action pair; Taking the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; Based on the target state-action pairs corresponding to the multiple decision steps, the optimal agent strategy corresponding to the interface interaction task is determined.
2. The interactive interface processing method based on Markov process according to claim 1, characterized in that: The interface interaction task is modeled based on the Markov decision process to obtain multiple decision steps, including: Determine a state space of a target interface corresponding to the interface interaction task, wherein the state space includes a plurality of states, and the states are used to represent element attributes of the target interface; Determine an action space corresponding to each state in the state space, wherein the action space includes a plurality of actions, and the actions are used to represent the operation type for the element attribute; Determine each state in the state space, execute each action in the action space respectively to obtain an updated state, and use the set corresponding to the updated state as the updated state space; Determine, for each state in the state space, an immediate reward for executing a decision step corresponding to each action in the action space; Determining a state transition probability based on the state space, the action space, and the updated state space; Based on the state transition probability and the interface interaction task, a plurality of decision steps are determined.
3. The interactive interface processing method based on Markov process according to claim 2 is characterized in that: The method further comprises: According to the state value function, determine the expected value of the cumulative immediate reward corresponding to executing the target agent strategy in the target state: ; in, represents the state-value function, represents the expected value, represents the termination time step, represents the discount factor, represents the immediate reward corresponding to the target agent strategy, Indicates the starting state Set to the target state , Represents the target agent strategy.
4. The interactive interface processing method based on Markov process according to claim 3 is characterized in that: The method further comprises: According to the action value function, the expected value of the cumulative immediate reward corresponding to the execution of the target agent strategy after executing the target action in the target state is determined: ; in, represents the action-value function, represents the expected value, represents the termination time step, represents the discount factor, represents the immediate reward corresponding to the target agent strategy, Indicates the starting state Set to the target state , Indicates that the action will be started Set as the target action , Represents the target agent strategy.
5. The interactive interface processing method based on Markov process according to claim 1, characterized in that: The step reward mechanism based on the preset step is used to determine the action selection reward, coordinate positioning reward and output result reward for each state-action pair of each decision step, and obtain the process reward for each state-action pair, including: Obtaining the correct operation, correct coordinates, and correct output corresponding to the interface interaction task; Take each state-action pair of each decision step as the current state-action pair, and perform the following operations to obtain the process reward of the current state-action pair: Determine the action selection, coordinate positioning and output result of the current state-action pair; Determining an action selection reward for the current state-action pair based on the similarity between the correct operation and the action selection; Determining a coordinate positioning reward for the current state-action pair based on the similarity between the correct coordinates and the coordinate positioning; Determining an output result reward of the current state-action pair based on the similarity between the correct output and the output result; The sum of the action selection reward, the coordinate positioning reward and the output result reward is used as the process reward of the current state-action pair.
6. The interactive interface processing method based on Markov process according to claim 1, characterized in that: The method further comprises: Determine the current state of the interactive interface understanding model at the target decision step of the interface interaction task, perform multiple action sampling on the target decision step, and obtain a target state-action pair set and an immediate process reward corresponding to the target state-action pair set; Determine the current sampled action distribution probability of the target state-action pair set under the current agent strategy; Determine the historical sampled action distribution probability of the target state-action pair set under the historical agent strategy; Based on the current sampling action distribution probability and the historical sampling action distribution probability, the strategy optimization target of the target decision step is determined: ; in, Indicates the sampling times of the multiple action samplings; Represents the current agent strategy In the current state , the interface interaction task Next generation Sampling Actions probability; Represents the historical agent strategy In the current state , the interface interaction task Next generation Sampling Actions probability; Indicates Advantage function of each sampled action; ; in, Indicates The process reward of sampling actions, represents the mean of the rewards of multiple action sampling processes, Represents the standard deviation of the reward over multiple action samples.
7. An interactive interface processing device based on Markov process, characterized in that: include: An acquisition module is used to acquire the interface interaction tasks input by the user; A modeling module, used to understand the model through the current interactive interface, model the interface interaction task based on the Markov decision process, and obtain multiple decision steps; A state-action module, used to determine a plurality of state-action pairs corresponding to each decision step, wherein the state-action pairs are used to simulate the state transition of user actions and interface interactions; A reward module, for determining an action selection reward, a coordinate positioning reward, and an output result reward for each state-action pair of each decision step based on a preset step reward mechanism, and obtaining a process reward for each state-action pair; A determination module, configured to take the state-action pair with the maximum process reward as the target state-action pair corresponding to the decision step; A decision module is used to determine the optimal intelligent agent strategy corresponding to the interface interaction task based on the target state-action pairs corresponding to the multiple decision steps.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the interactive interface processing method based on the Markov process as claimed in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the interactive interface processing method based on Markov process as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the interactive interface processing method based on Markov process as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Task planning method based on reinforcement learning under sequential logic constraint and related device
CN114265674A
Intelligent ship collision avoidance decision-making method driven by expert demonstration data
CN117687405A
Bus resource dynamic scheduling method based on spatial graph convolution and near-end strategy optimization
CN119169854A