Man-machine cooperation method and system based on robot process automation
By building a data perception and knowledge migration processing network, the efficient human-computer collaboration of the RPA system in a dynamic environment is achieved, and the problem of flexibility and insufficient knowledge migration of traditional RPA systems in complex process processing is solved, improving the system's adaptability and collaboration efficiency.
Patent Information
- Application Number
- CN202510749034.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Traditional RPA systems lack flexibility in dealing with dynamically changing complex business processes, lack knowledge transfer capabilities, and low human-computer collaboration efficiency.
Build a data-aware processing network and a knowledge migration processing network, generate task processing strategies through data acquisition and semantic analysis, robot units handle structured tasks, manual operation units handle unstructured tasks, and real-time communication and collaboration through information interaction channels, using graph embedding algorithms and scene migration models to achieve knowledge migration.
It improves the adaptability and generalization capabilities of the RPA system in complex environments, optimizes the human-computer collaboration process, and improves task processing efficiency and coordination efficiency.
Smart Images

Figure CN120258747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to automation technology, and particularly to a human - machine collaboration method and system based on robotic process automation. Background Art
[0002] Robotic Process Automation (RPA) technology has made significant progress in automating various business processes in recent years, especially in processing structured data and rule - based tasks. Traditional RPA systems are good at performing predefined, repetitive tasks such as data entry, form filling, and report generation. These systems usually rely on explicit rules and workflows and have limitations when dealing with unstructured data or complex scenarios that require human judgment and intervention. The existing RPA systems mainly have the following three defects and deficiencies: Firstly, traditional RPA systems are difficult to adapt to dynamically changing environments. They usually require explicit rules and instructions, and when faced with unexpected situations or process changes, a large amount of reprogramming and configuration is required. This limits their flexibility in dealing with complex and unpredictable business processes.
[0003] Secondly, traditional RPA systems lack effective knowledge transfer capabilities. They can usually only process specific types of data and tasks and are difficult to apply the learned knowledge to new scenarios. This leads to an increase in the cost of repeated development and deployment and limits the wide application of RPA technology.
[0004] Finally, traditional RPA systems are insufficient in human - machine collaboration. They usually regard humans and machines as independent individuals and lack effective communication and collaboration mechanisms. This results in low efficiency and limits the potential of human - machine collaboration. In many actual application scenarios, humans and machines need to complete tasks together, such as in the case of dealing with unstructured data or situations that require human judgment and intervention. Summary of the Invention
[0005] Embodiments of the present invention provide a human - machine collaboration method and system based on robotic process automation, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention, A human - machine collaboration method based on robotic process automation is provided, including: Constructing a data perception and processing network, where the data perception and processing network includes a data acquisition layer and a semantic understanding layer. The data acquisition layer collects user operation instruction data, system operation status data, and task environment data through distributed sensors. The semantic understanding layer performs semantic parsing on the collected data based on a deep neural network to obtain a task feature vector, matches the task feature vector with a preset business model to construct a task mapping matrix, and generates a task processing strategy; Build a human-machine collaborative processing framework based on a task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing exceeds the specified threshold, the robot unit sends the historical data analysis result and decision-making suggestion information to the manual operation unit, and the manual operation unit transmits the processing procedure information to the rule database of the robot unit; Build a knowledge transfer processing network. Use a graph embedding algorithm to extract features from the business rule data and expert experience data during the human-machine collaborative processing to obtain business feature vectors. Train a scenario transfer model based on the business feature vectors, and map the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer processing network also includes a verification module to verify the effectiveness of the scenario mapping data and output a verification result, and update the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaborative processing framework based on the verification result.
[0007] In an alternative embodiment, The steps of generating a task processing strategy by performing feature matching between the task feature vector and a preset business model to construct a task mapping matrix include: Build a model library containing multiple preset business models. Each preset business model contains a standard feature vector and a corresponding processing template. Calculate the similarity between the task feature vector and the standard feature vector based on contrastive learning, map the task feature vector to the feature space corresponding to the standard feature vector with the highest similarity, and construct a task mapping matrix based on the mapping result; Input the task mapping matrix into a hierarchical decoder. The hierarchical decoder includes a task decomposition layer, a resource allocation layer, and an execution sequence layer connected in sequence. The task decomposition layer generates a task decomposition strategy based on the task mapping matrix, the resource allocation layer generates a resource allocation strategy based on the task mapping matrix and the task decomposition strategy, and the execution sequence layer generates an execution sequence strategy based on the task mapping matrix, the task decomposition strategy, and the resource allocation strategy; Calculate a constraint adaptation factor based on the task feature vector, the resource status output of the resource allocation layer, and the time constraint parameter. Based on the constraint adaptation factor, perform constraint adjustment on the task decomposition strategy, resource allocation strategy, and execution sequence strategy respectively to obtain a task processing strategy; construct a dominance relationship matrix of the task processing strategy, calculate an optimal strategy combination based on the dominance relationship matrix, and determine the strategy weight according to the contribution degrees of the task decomposition strategy, resource allocation strategy, and execution sequence strategy to the optimal strategy combination; Input the policy weight and the optimal policy combination into a policy optimization network, where the policy optimization network includes a policy storage module and a policy sampling module. The policy storage module stores historical policy data, and the policy sampling module performs importance sampling on the historical policy data to obtain sampled policy data. Calculate policy evaluation parameters based on the sampled policy data, and optimize and update the network parameters of the hierarchical decoder according to the policy evaluation parameters.
[0008] In an alternative embodiment, Construct a human-machine collaborative processing framework based on a task processing strategy. The steps for a robot unit to process structured tasks that meet preset rules and for a manual operation unit to process unstructured tasks include: Obtain a task type feature vector based on the task processing strategy, calculate the entropy value of the task type feature vector using an information entropy calculation model, and divide the tasks into structured tasks and unstructured tasks according to the entropy value; Construct a policy mapping mechanism for the robot execution unit based on the task processing strategy. Map the task decomposition strategy in the task processing strategy to the robot basic action library to obtain an action correspondence matrix, determine action execution parameters according to the resource allocation strategy in the task processing strategy, generate a robot control instruction stream based on the execution sequence strategy in the task processing strategy, and the robot execution unit processes the structured tasks according to the robot control instruction stream; Construct an operator cognitive state vector including attention level, workload, and professionalism. Evaluate the operator cognitive state vector based on the resource allocation strategy in the task processing strategy, dynamically adjust the information display density and frequency of the manual operation interface according to the evaluation result, and process the unstructured tasks through the manual operation interface; Construct collaborative effectiveness evaluation indicators including execution efficiency, accuracy, and cognitive load. Determine the collaborative weight coefficient of the collaborative effectiveness evaluation indicators according to the policy weight in the task processing strategy, calculate the human-machine collaborative processing effectiveness score based on the collaborative weight coefficient, and use the human-machine collaborative processing effectiveness score as a feedback signal to optimize and update the task processing strategy.
[0009] In an alternative embodiment, The robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing process is greater than a specified threshold, the steps for the robot unit to send the historical data analysis result and decision-making advice information to the manual operation unit and for the manual operation unit to transmit the processing procedure information to the rule database of the robot unit include: The task execution trajectory is decomposed into time-frequency feature representations of different scales through wavelet transform. An adaptive attention mechanism is used to perform importance weighting on the time-frequency feature representations of different scales, and the weighted features are fused to obtain global deviation features. A recurrent neural network is used to model the global deviation feature sequence, and a multi-layer perceptron is used to calculate the deviation degree score. When the deviation degree score exceeds a preset deviation threshold, an anomaly detection signal is triggered; an anomaly pattern feature library is constructed based on historical execution data, and the type and cause of the current anomaly are matched from the anomaly pattern feature library based on the anomaly detection signal; A multi-level knowledge graph structure including a task layer graph, a resource layer graph, and a processing layer graph is constructed. The task layer graph describes the dependencies between tasks, the resource layer graph describes the distribution of device and personnel capabilities, and the processing layer graph describes the operation process. The node information in the multi-level knowledge graph structure is updated based on the anomaly type and cause; A graph neural network using a multi-head attention mechanism is used for inter-layer information transfer of the multi-level knowledge graph structure, and node feature representations are generated through normalized weights calculation between nodes; the node feature representations are combined with historical sequence features to construct a state vector; a policy network is trained based on the state vector, and the weighted combination of task completion degree and resource utilization rate is used as a reward signal. The network parameters are optimized using the policy gradient method to obtain a decision-making model; guided by the decision-making model, the access count, value estimation, and prior probability are used as search attributes, and the Monte Carlo tree search algorithm is used to generate an optimal decision path; a local causal graph including variable nodes and causal edges is constructed for the nodes in the optimal decision path to generate processing suggestions; The processing suggestions are represented as processing procedures of conditional action pairs, the importance of the processing procedures is evaluated, and the processing procedures with importance exceeding a preset importance threshold are updated to the rule database, and the update result of the rule database is used to adjust the preset deviation threshold.
[0010] In an alternative embodiment, The steps of constructing a knowledge transfer processing network, using a graph embedding algorithm to extract features from business rule data and expert experience data in the human-machine collaborative processing process to obtain business feature vectors, and training a scenario transfer model based on the business feature vectors to map the business feature vectors to a preset scenario feature space to generate scenario mapping data include: The knowledge transfer processing network includes a knowledge graph construction module, a feature extraction module, and a scenario transfer module; The knowledge graph construction module is used to construct a knowledge graph including rule nodes, experience nodes, and relationship edges. The rule nodes describe the robot process automation operation procedures in the form of a triple of trigger conditions, operation sequences, and constraint conditions. The experience nodes represent process optimization techniques in a structured form of abnormal scenarios, processing steps, and expected effects. The relationship edges describe the inclusion relationship, pre - condition relationship, and optimization relationship between the rule nodes and the experience nodes; The feature extraction module uses a graph embedding algorithm with a multi - head attention mechanism to extract features from the business rule data and expert experience data in the knowledge graph. Among them, the business rule data comes from the trigger conditions, operation sequences, and constraint conditions of the rule nodes, and the expert experience data comes from the abnormal scenarios, processing steps, and expected effects of the experience nodes. By calculating the attention scores of the node set directly connected by relationship edges to the current processing node and performing feature aggregation, a business feature vector is obtained, and the business feature vector contains the interaction relationship information between the robot process automation operation nodes; The scenario migration module trains a scenario migration model based on the business feature vector. The scenario migration model includes a discriminator network and a generator network. The discriminator network is used to distinguish between the real process samples and the generated migration samples of the target scenario, and the generator network is used to map the business feature vector of the source scenario to a preset scenario feature space; A contrast learning mechanism is adopted to align the feature spaces of the source scenario and the target scenario. The same processing steps in the source scenario and the target scenario are constructed as positive sample pairs, and the different processing steps are constructed as negative samples. By minimizing the contrast loss function, the feature representations of the same operation steps approach each other in the feature space, and scenario mapping data is obtained.
[0011] In an alternative embodiment, The steps for the scenario migration module to train a scenario migration model based on the business feature vector, where the scenario migration model includes a discriminator network and a generator network, and the discriminator network is used to distinguish between the real process samples and the generated migration samples of the target scenario, and the generator network is used to map the business feature vector of the source scenario to a preset scenario feature space, include: The generator network adopts a three-layer fully-connected neural network structure, and the layers are connected by ReLU activation functions. The last layer outputs the target scene feature vector through a hyperbolic tangent activation function. The discriminator network adopts a four-layer fully-connected network structure, and a batch normalization layer and a dropout layer are set after each hidden layer. A discriminator loss function is constructed based on the Wasserstein distance. The discriminator loss function includes an expected term for real samples, an expected term for generated samples, and a gradient penalty term. The discriminator network is trained through the discriminator loss function. A generator loss function is constructed. The generator loss function includes an adversarial loss term output by the discriminator and a semantic consistency loss term. The semantic consistency loss term is obtained by calculating the Euclidean distance between the source scene features and the generated features and the Euclidean distance between the features in the intermediate layer of the feature extraction network. A feature quality evaluation index is constructed. The feature quality evaluation index includes a feature consistency score and a task relevance score. The feature consistency score is calculated based on the feature space distance between the target scene features and the generated features, and the task relevance score is calculated based on the accuracy of the downstream task classifier. An alternating optimization strategy is adopted to train the scene migration model. The generator network is fixed, and the parameters of the discriminator network are optimized iteratively for multiple times. The discriminator network is fixed, and the parameters of the generator network are optimized once. The service feature vectors are divided into training scales according to the number of rule node triples. The number of rule node triples in the initial training scale is set to 1, the growth step value of the number of rule node triples is set, and the number of training epochs for each training scale is set. The scene migration model is trained based on the service feature vectors of the current training scale. When the feature quality evaluation index reaches a preset threshold, the number of rule node triples is increased according to the growth step value, and the scene migration model is continuously trained based on the increased training scale until the preset number of training times is reached. The generated target scene features are evaluated based on the feature quality evaluation index, and the weight coefficient of the semantic consistency loss term and the growth step value of the number of rule node triples are adjusted according to the evaluation result of the target scene features.
[0012] In an optional implementation manner, The knowledge transfer processing network further includes a verification module that validates the effectiveness of the scene mapping data and outputs a verification result. The steps of updating the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaboration processing framework based on the verification result include: Construct a verification matrix to verify the scenario mapping data. The verification matrix includes a semantic consistency vector, a structural similarity vector, and a temporal rationality vector. The semantic consistency vector is obtained by calculating the keyword matching degree of business rules before and after mapping. The structural similarity vector is calculated based on the edit distance of the graph structure. The temporal rationality vector is obtained by verifying the dependency relationship of the operation sequence. Construct multi-level verification metrics based on the verification matrix. The multi-level verification metrics include rule-level verification metrics and process-level verification metrics. The rule-level verification metrics are obtained by calculating rule similarity and internal rule consistency. The process-level verification metrics are obtained by calculating temporal constraint satisfaction and structural integrity. Weightedly combine the multi-level verification metrics with the business scenario coverage rate to generate a verification evaluation score. Update the parameters of the data-aware processing network based on the verification evaluation score, including: performing gradient update on the parameters of the perception layer using an adaptive learning rate, where the adaptive learning rate is dynamically adjusted according to the verification evaluation score; determining the adjustment direction and adjustment step size of the data-aware processing network parameters based on the difference between the verification evaluation score and a preset threshold. Construct a task assignment matrix, where the elements of the task assignment matrix represent the probability of task assignment to processing units. Calculate the update amount of the task assignment matrix based on the verification evaluation score and resource utilization rate; dynamically adjust the task assignment scheme in the human-machine collaborative processing framework according to the update amount of the task assignment matrix, and optimize the task assignment parameters in the human-machine collaborative processing framework based on the adjusted task assignment scheme.
[0013] In the second aspect of the embodiments of the present invention, Provide a human-machine collaborative system based on robotic process automation, including: A first unit for constructing a data-aware processing network, collecting user operation instruction data, system operation status data, and task environment data, performing semantic parsing on the collected data to obtain task feature vectors, and constructing a task mapping matrix by matching the task feature vectors with a preset business model to generate a task processing strategy. A second unit for constructing a human-machine collaborative processing framework based on the task processing strategy, where the robotic unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robotic unit and the manual operation unit are connected through an information interaction channel. When the deviation value during task processing is greater than a specified threshold, the robotic unit sends historical data analysis results and decision-making suggestion information to the manual operation unit, and the manual operation unit transmits processing procedure information to the rule database of the robotic unit. A third unit, for constructing a knowledge transfer processing network, uses a graph embedding algorithm to extract features from the business rule data and expert experience data in the human-machine collaborative processing process to obtain business feature vectors, trains a scenario transfer model based on the business feature vectors, and maps the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer processing network further includes a verification module, which verifies the validity of the scenario mapping data and outputs a verification result, and updates the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaborative processing framework based on the verification result.
[0014] In the third aspect of the embodiments of the present invention, a kind of electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0015] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0016] Through the data perception processing network, the present invention can quickly and accurately understand the user operation instructions and the system operation state, and combine with the task environment data to generate an optimal task processing strategy. By using the human-machine collaborative processing framework, the robot can efficiently process structured tasks, while the human focuses on unstructured tasks, and the two work together to improve the task processing efficiency.
[0017] When there are deviations in the task processing process of the present invention, the robot can provide the historical data analysis results and decision-making suggestions to the human, and the human can adjust the processing procedures according to the actual situation and update the robot rule database, so that the system can adapt to the complex and changeable task environment. The knowledge transfer processing network can transfer the business rules and expert experience in the human-machine collaborative processing process to a new scenario, improving the generalization ability and adaptability of the system.
[0018] Through the information interaction channel, the robot and the human can communicate and cooperate in real time to achieve complementary advantages. The verification module of the knowledge transfer processing network can verify the validity of the scenario mapping data, and optimize the parameters of the data perception processing network and the human-machine collaborative processing framework according to the verification result, thereby continuously optimizing the human-machine collaborative process and improving the collaborative efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flowchart of the human-machine collaborative method based on robotic process automation according to the embodiments of the present invention; Figure 2 This is a schematic structural diagram of a human - machine collaboration system based on robotic process automation according to an embodiment of the present invention. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0022] Figure 1 This is a schematic flowchart of a human - machine collaboration method based on robotic process automation according to an embodiment of the present invention. As Figure 1 shown, the method includes: Construct a data perception and processing network, collect user operation instruction data, system operation status data, and task environment data, perform semantic parsing on the collected data to obtain task feature vectors, match the task feature vectors with a preset business model to construct a task mapping matrix, and generate a task processing strategy; Construct a human - machine collaboration processing framework based on the task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing process is greater than the specified threshold, the robot unit sends historical data analysis results and decision - making suggestion information to the manual operation unit, and the manual operation unit transmits processing procedure information to the rule database of the robot unit; Construct a knowledge transfer and processing network, use a graph embedding algorithm to extract feature vectors of business rule data and expert experience data during the human - machine collaboration processing process to obtain business feature vectors, train a scenario transfer model based on the business feature vectors, and map the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer and processing network further includes a verification module, which verifies the validity of the scenario mapping data and outputs a verification result, and updates the processing parameters of the data perception and processing network and the task allocation parameters of the human - machine collaboration processing framework based on the verification result.
[0023] In an alternative implementation manner, The steps of matching the task feature vectors with a preset business model to construct a task mapping matrix and generating a task processing strategy include: Construct a model library containing multiple preset business models. Each preset business model includes a standard feature vector and a corresponding processing template. Calculate the similarity between the task feature vector and the standard feature vector based on contrastive learning, map the task feature vector to the feature space corresponding to the standard feature vector with the highest similarity, and construct a task mapping matrix based on the mapping result; Input the task mapping matrix into a hierarchical decoder. The hierarchical decoder includes a task decomposition layer, a resource allocation layer, and an execution sequence layer connected in sequence. The task decomposition layer generates a task decomposition strategy based on the task mapping matrix, the resource allocation layer generates a resource allocation strategy based on the task mapping matrix and the task decomposition strategy, and the execution sequence layer generates an execution sequence strategy based on the task mapping matrix, the task decomposition strategy, and the resource allocation strategy; Calculate a constraint adaptation factor according to the task feature vector, the resource status output of the resource allocation layer, and the time constraint parameter. Based on the constraint adaptation factor, respectively perform constraint adjustment on the task decomposition strategy, the resource allocation strategy, and the execution sequence strategy to obtain a task processing strategy; construct a dominance relationship matrix of the task processing strategy, calculate an optimal strategy combination based on the dominance relationship matrix, and determine strategy weights according to the contribution degrees of the task decomposition strategy, the resource allocation strategy, and the execution sequence strategy to the optimal strategy combination; Input the strategy weights and the optimal strategy combination into a strategy optimization network. The strategy optimization network includes a strategy storage module and a strategy sampling module. The strategy storage module stores historical strategy data, the strategy sampling module performs importance sampling on the historical strategy data to obtain sampled strategy data, calculates a strategy evaluation parameter based on the sampled strategy data, and optimizes and updates the network parameters of the hierarchical decoder according to the strategy evaluation parameter.
[0024] Exemplarily, in the semantic parsing stage, the BERT pre-trained model is used to extract context semantic features for text-based instructions, Word2Vec is used to map discrete operations into word vectors, and the LSTM network is used to process the temporal features for numerical data. Finally, the attention mechanism is used to perform weighted fusion on the three types of features to obtain the task feature vector, achieving a unified expression of the user's intention. Then, a model library is constructed. The model library contains multiple preset business models, and each business model contains a set of standard feature vectors and corresponding processing templates. For example, a business model for "customer complaints" may have standard feature vectors including features such as "complaint type", "complaint urgency", "customer value", etc., and the corresponding processing template stipulates the standard process for handling this type of complaint. Suppose the model library contains three business models: "customer complaints", "order processing", and "product consultation", and the standard feature vectors of each model are [0.8, 0.2, 0.5], [0.1, 0.9, 0.3], and [0.3, 0.4, 0.8] respectively.
[0025] Perform feature matching. Calculate the similarity between the input task feature vector and the standard feature vectors of each business model in the model library. Methods such as cosine similarity can be used to calculate the similarity. Map the task feature vector to the feature space corresponding to the standard feature vector with the highest similarity. For example, if the feature vector of a task is [0.7, 0.3, 0.6] and has the highest similarity with the "customer complaints" model, then this task is mapped to the feature space of the "customer complaints" model. Suppose the calculated similarities are 0.95, 0.75, and 0.85 respectively, then this task is mapped to the "customer complaints" model.
[0026] Construct a task mapping matrix. This matrix reflects the mapping relationship between tasks and each business model. For example, if a task is mapped to the "customer complaints" model, then the value at the position corresponding to the "customer complaints" model in the matrix is 1, and the values at other positions are 0. Taking the above case as an example, the task mapping matrix is [1, 0, 0].
[0027] Input the task mapping matrix into the hierarchical decoder. The hierarchical decoder consists of a task decomposition layer, a resource allocation layer, and an execution sequence layer. The task decomposition layer decomposes the task into subtasks according to the task mapping matrix. For example, decompose the "customer complaints" task into subtasks such as "record complaint information", "investigate the cause of the complaint", and "formulate a solution". The resource allocation layer allocates resources according to the task mapping matrix and the task decomposition strategy. For example, allocate the subtask of "investigate the cause of the complaint" to the customer service staff and the subtask of "formulate a solution" to the technical staff. The execution sequence layer generates an execution sequence according to the task mapping matrix, the task decomposition strategy, and the resource allocation strategy. For example, first execute "record complaint information", then execute "investigate the cause of the complaint", and finally execute "formulate a solution".
[0028] Calculate the constraint adaptation factor based on the task feature vector, the resource status output of the resource allocation layer, and the time constraint parameter. For example, if resources are scarce or time is tight, the constraint adaptation factor is small. Based on the constraint adaptation factor, constraint adjustments are made to the task decomposition strategy, resource allocation strategy, and execution sequence strategy to obtain the final task processing strategy. For example, if time is tight, the steps of "investigating the cause of the complaint" can be simplified.
[0029] Construct a dominance relationship matrix for the task processing strategies. This matrix reflects the mutual influence relationship between different strategies. For example, the "investigating the cause of the complaint" strategy must be executed before the "formulating a solution" strategy. Calculate the optimal strategy combination based on the dominance relationship matrix. For example, select the strategy combination with the highest execution efficiency and meeting the constraint conditions. Determine the strategy weights according to the contribution degrees of the task decomposition strategy, resource allocation strategy, and execution sequence strategy to the optimal strategy combination. For example, if the "investigating the cause of the complaint" strategy has a greater impact on the final result, a higher weight is assigned to it.
[0030] Input the strategy weights and the optimal strategy combination into the strategy optimization network. The strategy optimization network includes a strategy storage module and a strategy sampling module. The strategy storage module stores historical strategy data. The strategy sampling module performs importance sampling on the historical strategy data to obtain sampled strategy data. For example, preferentially sample strategy data with a higher success rate. Calculate the strategy evaluation parameters based on the sampled strategy data. For example, calculate the average execution time and success rate of the strategy. Optimize and update the network parameters of the hierarchical decoder according to the strategy evaluation parameters. For example, if the success rate of a certain strategy is low, adjust the corresponding parameters in the hierarchical decoder to improve the success rate of this strategy.
[0031] Through feature matching and hierarchical decoding, the present invention can quickly generate processing strategies for specific tasks, avoiding the cumbersome process of manually formulating strategies, thereby improving task processing efficiency; dynamically allocate resources according to task requirements and resource status, avoiding resource waste and improving resource utilization rate; continuously learn and optimize strategies through the strategy optimization network, enabling the generated strategies to be more adaptable to the actual situation and improving the success rate of task processing.
[0032] In an alternative embodiment, Construct a human-machine collaborative processing framework based on the task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the steps for the manual operation unit to process unstructured tasks include: Obtain the task type feature vector based on the task processing strategy, calculate the entropy value of the task type feature vector based on the information entropy calculation model, and divide the tasks into structured tasks and unstructured tasks according to the entropy value; Construct a policy mapping mechanism for the robot execution unit based on the task processing policy, map the task decomposition policy in the task processing policy to the robot basic action library to obtain an action correspondence matrix, determine the action execution parameters according to the resource allocation policy in the task processing policy, generate a robot control instruction stream based on the execution sequence policy in the task processing policy, and the robot execution unit processes the structured task according to the robot control instruction stream; Construct an operator cognitive state vector including attention level, workload, and professionalism, evaluate the operator cognitive state vector based on the resource allocation policy in the task processing policy, dynamically adjust the information display density and frequency of the manual operation interface according to the evaluation result, and process the unstructured task through the manual operation interface; Construct a collaborative effectiveness evaluation index including execution efficiency, accuracy, and cognitive load, determine the collaborative weight coefficient of the collaborative effectiveness evaluation index according to the policy weight in the task processing policy, calculate the human-machine collaborative processing effectiveness score based on the collaborative weight coefficient, and use the human-machine collaborative processing effectiveness score as a feedback signal to optimize and update the task processing policy.
[0033] Exemplarily, first, it is necessary to analyze the input task and extract its type feature vector. For example, a task can be described by multiple features, such as data type (text, image, numerical value), data structure (structured, semi-structured, unstructured), task complexity (simple, medium, complex), etc. Quantify these features to form a task type feature vector. For example, [0.8, 0.2, 0.5] represents text data, semi-structured data, and medium complexity respectively. Then, use the information entropy calculation model to calculate the entropy value of this feature vector. The higher the entropy value, the greater the uncertainty of the task type, and it is more inclined to classify it as an unstructured task. Suppose the calculated entropy value is 0.6. According to the preset threshold (such as 0.5), it is determined that this task is an unstructured task and assigned to a human operator for processing. Otherwise, it is classified as a structured task and assigned to the robot unit for processing.
[0034] For the structured tasks assigned to the robot unit, the policy mapping mechanism in the framework transforms the predefined task processing policies into instructions executable by the robot. For example, the task processing policies include task decomposition policies, resource allocation policies, and execution sequence policies. Suppose a structured task is "move item A from location B to location C". The task decomposition policy decomposes it into subtasks such as "grab item A", "move to location B", "grab item A", "move to location C", "put down item A", etc. The resource allocation policy determines the resources required for each subtask, such as the robot arm, mobile platform, etc. The execution sequence policy determines the execution order of the subtasks. The policy mapping mechanism maps these policies to the robot's basic action library, such as "grab", "move", "put down", etc., and generates corresponding action parameters, such as the grabbing force, moving speed, target location, etc., finally forming a robot control instruction stream to control the robot to complete the task.
[0035] For the unstructured tasks assigned to human operators, the framework dynamically adjusts the information display of the human-machine operation interface according to the cognitive state of the operator. For example, construct an operator cognitive state vector, which includes indicators such as attention level, workload, and professionalism. Suppose the current operator's cognitive state vector is [0.9, 0.5, 0.8], representing a high attention level, medium workload, and high professionalism. Based on the resource allocation policy in the task processing policy, evaluate this cognitive state and adjust the information display density and frequency of the human-machine operation interface accordingly. For example, for an operator with a high attention level, the information display density can be increased; for an operator with a high workload, the information display frequency can be reduced. In this way, the human-machine interaction can be optimized and the work efficiency of the operator can be improved.
[0036] Finally, the framework evaluates the effectiveness of human-machine collaborative processing and optimizes the task processing policy according to the evaluation results. For example, construct collaborative effectiveness evaluation indicators, which include execution efficiency, accuracy, and cognitive load, etc. Suppose the weights of execution efficiency, accuracy, and cognitive load in the task processing policy are 0.4, 0.5, and 0.1 respectively. According to these weights, calculate the human-machine collaborative processing effectiveness score. For example, suppose the current execution efficiency is 0.9, accuracy is 0.8, and cognitive load is 0.6, then the collaborative effectiveness score is 0.4×0.9 + 0.5×0.8 + 0.1×0.6 = 0.82. Use this score as a feedback signal to optimize and update the task processing policy, such as adjusting the policy weights, optimizing the task decomposition policy, etc., so as to continuously improve the effectiveness of human-machine collaborative processing.
[0037] The present invention assigns structured tasks to robots for automatic processing and unstructured tasks to humans for processing, giving full play to their respective advantages, thereby improving the overall processing efficiency. Robots have a high accuracy in processing structured tasks, and humans have strong flexibility in processing unstructured tasks. The combination of the two can effectively improve the overall processing accuracy. Assigning cumbersome and repetitive structured tasks to robots for processing can effectively reduce the cognitive load and work intensity of operators, enabling them to focus on more creative unstructured tasks.
[0038] In an alternative embodiment, The robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing exceeds the specified threshold, the steps for the robot unit to send the historical data analysis result and decision-making advice information to the manual operation unit, and for the manual operation unit to transmit the processing procedure information to the rule database of the robot unit include: Decompose the task execution trajectory into time-frequency feature representations of different scales through wavelet transform, use an adaptive attention mechanism to perform importance weighting on the time-frequency feature representations of different scales, fuse the weighted features to obtain global deviation features, use a recurrent neural network to model the global deviation feature sequence, calculate the deviation degree score through a multi-layer perceptron, and trigger an anomaly detection signal when the deviation degree score exceeds the preset deviation threshold; construct an anomaly pattern feature library based on historical execution data, and match the type and cause of the current anomaly from the anomaly pattern feature library based on the anomaly detection signal; Construct a multi-level knowledge graph structure including a task layer graph, a resource layer graph, and a processing layer graph. The task layer graph describes the dependencies between tasks, the resource layer graph describes the distribution of device and personnel capabilities, and the processing layer graph describes the operation process. Update the node information in the multi-level knowledge graph structure based on the anomaly type and cause; Use a graph neural network with a multi-head attention mechanism to perform inter-layer information transfer on the multi-level knowledge graph structure, and generate node feature representations through normalized weights calculation between nodes; combine the node feature representations with historical sequence features to construct a state vector; train a policy network based on the state vector, use the weighted combination of task completion and resource utilization as a reward signal, and optimize the network parameters using the policy gradient method to obtain a decision model; guided by the decision model, use access count, value estimation, and prior probability as search attributes, and use the Monte Carlo tree search algorithm to generate an optimal decision path; construct a local causal graph including variable nodes and causal edges for the nodes in the optimal decision path to generate processing suggestions; The processing procedure that represents the processing suggestions as condition-action pairs is evaluated for importance, and the processing procedures with importance exceeding the preset importance threshold are updated to the rule database, and the update result of the rule database is used to adjust the preset deviation threshold.
[0039] Exemplarily, deviation detection is performed on the task execution trajectory. The task execution trajectory is decomposed into feature representations at different time scales and frequency scales. For example, the motion trajectory of the robotic arm is decomposed into features at different frequencies such as speed and acceleration, and features within different time periods. Then, weighting is performed according to the importance of features at different scales. For example, features with drastic speed changes are given higher weights. The weighted features are fused to obtain global deviation features, and a recurrent neural network is used to model the global deviation feature sequence. For example, a long short-term memory network (LSTM) is used to learn the dependencies between historical deviation features. Finally, the deviation degree score is calculated through a multi-layer perceptron. For example, the output of the LSTM network is used as the input of the multi-layer perceptron, and a deviation score between 0 and 1 is calculated. When the deviation degree score exceeds the preset deviation threshold (e.g., 0.8), an anomaly detection signal is triggered.
[0040] After the anomaly detection signal is triggered, the system performs anomaly matching based on the anomaly pattern feature library constructed from historical execution data. This feature library contains feature representations of various known anomaly types. For example, the robotic arm getting stuck, sensor failure, etc. By comparing the similarity between the features of the current anomaly and the anomaly patterns in the feature library, the type and cause of the current anomaly are determined. For example, if the features of the current anomaly are highly similar to the feature pattern of "the robotic arm getting stuck", it is determined that the current anomaly type is "the robotic arm getting stuck".
[0041] The system updates the multi-level knowledge graph structure. This structure includes a task layer graph, a resource layer graph, and a processing layer graph. The task layer graph describes the dependencies between tasks. For example, the "screw tightening" task depends on the "screw grasping" task. The resource layer graph describes the distribution of device and personnel capabilities. For example, robotic arm A can perform grasping and screw tightening operations, and operator B is good at handling sensor failures. The processing layer graph describes the operation process. For example, the process for handling a stuck robotic arm is: stop the movement of the robotic arm, check the cause of the jam, eliminate the fault, and restart the robotic arm. According to the anomaly type and cause, the node information in the multi-level knowledge graph structure is updated. For example, if robotic arm A gets stuck, the status of robotic arm A in the resource layer graph is updated to "fault".
[0042] The system uses a graph neural network to perform inter-layer information transfer on the multi-level knowledge graph structure. By calculating the normalized weights between nodes, node feature representations are generated. For example, the node feature representation of robotic arm A includes information such as its status, capabilities, and connection relationships with other nodes. The node feature representation is combined with historical sequence features to construct a state vector. For example, the node feature representation of robotic arm A is combined with the execution trajectory features of robotic arm A over a past period of time.
[0043] Based on the state vector, the system trains a policy network. The weighted combination of task completion degree and resource utilization rate is used as the reward signal. For example, the weight of the task completion degree is 0.7, and the weight of the resource utilization rate is 0.3. The policy gradient method is used to optimize the network parameters to obtain a decision-making model. This decision-making model can output the best action policy according to the current state.
[0044] Guided by the decision-making model, the system uses the Monte Carlo tree search algorithm to generate an optimal decision-making path. The visit count, value estimation, and prior probability are used as search attributes. For example, the higher the visit count of a node, the more reliable its value estimation. By simulating different action paths and evaluating the value of each path, the path with the highest value is finally selected as the optimal decision-making path.
[0045] For the nodes in the optimal decision-making path, a local causal graph containing variable nodes and causal edges is constructed to generate processing suggestions.
[0046] The processing suggestions are represented as processing procedures of conditional action pairs. For example, "if the robotic arm gets stuck, stop the movement of the robotic arm". The importance of the processing procedures is evaluated, and the processing procedures whose importance exceeds the preset importance threshold are updated to the rule database. The update result of the rule database is used to adjust the preset deviation threshold. For example, if the new processing procedure is proven to be effective, the preset deviation threshold can be appropriately reduced.
[0047] Through the collaborative work of the robotic unit and the manual operation unit, the present invention can detect and process task anomalies faster; through the multi-level knowledge graph and the graph neural network, the system can better understand the relationships between tasks, resources, and processing flows, so as to be able to more effectively handle various abnormal situations and enhance the robustness of the system; the system stores the processing procedures in the rule database and continuously updates and optimizes them, thus realizing the accumulation and reuse of knowledge.
[0048] In an alternative embodiment, The steps of constructing a knowledge transfer processing network, using a graph embedding algorithm to extract feature vectors of business rule data and expert experience data in the human-machine collaborative processing process to obtain business feature vectors, training a scenario transfer model based on the business feature vectors, and mapping the business feature vectors to a preset scenario feature space to generate scenario mapping data include: The knowledge transfer processing network includes a knowledge graph construction module, a feature extraction module, and a scenario transfer module; The knowledge graph construction module is used to construct a knowledge graph including rule nodes, experience nodes, and relationship edges. The rule nodes describe the robot process automation operation procedures in the form of a triple of trigger conditions, operation sequences, and constraint conditions. The experience nodes represent process optimization techniques in a structured form of abnormal scenarios, processing steps, and expected effects. The relationship edges describe the inclusion relationship, pre - condition relationship, and optimization relationship between the rule nodes and the experience nodes; The feature extraction module uses a graph embedding algorithm with a multi - head attention mechanism to extract features from the business rule data and expert experience data in the knowledge graph. Among them, the business rule data comes from the trigger conditions, operation sequences, and constraint conditions of the rule nodes, and the expert experience data comes from the abnormal scenarios, processing steps, and expected effects of the experience nodes. By calculating the attention scores of the node set directly connected by relationship edges to the current processing node and performing feature aggregation, a business feature vector is obtained. The business feature vector contains the interaction relationship information between the robot process automation operation nodes; The scenario transfer module trains a scenario transfer model based on the business feature vector. The scenario transfer model includes a discriminator network and a generator network. The discriminator network is used to distinguish between real process samples and generated transfer samples of the target scenario, and the generator network is used to map the business feature vector of the source scenario to a preset scenario feature space; A contrast learning mechanism is used to align the feature spaces of the source scenario and the target scenario. The same processing steps in the source scenario and the target scenario are constructed as positive sample pairs, and the different processing steps are constructed as negative samples. By minimizing the contrast loss function, the feature representations of the same operation steps approach each other in the feature space, obtaining scenario mapping data.
[0049] Exemplarily, the knowledge transfer processing network is used to optimize robot process automation (RPA). By integrating expert experience into RPA rules, cross - scenario process migration and adaptation are achieved. This network mainly includes three modules: knowledge graph construction, feature extraction, and scenario transfer.
[0050] Construct a knowledge graph. The knowledge graph represents RPA operation procedures and expert experience in a graph structure. The nodes in the graph are divided into rule nodes and experience nodes. Rule nodes describe RPA operation procedures in the form of triples (trigger conditions, operation sequences, constraint conditions). For example, "If an invoice email is received, extract invoice information, and the invoice amount must be greater than 0". Experience nodes represent process optimization techniques in a structured form (abnormal scenarios, handling steps, expected effects). For example, "If the invoice format is incorrect, contact the supplier to confirm and ensure the accuracy of invoice information". The edges in the graph describe the relationships between rule nodes and experience nodes, including inclusion relationships (experience nodes are supplements to rule nodes), pre - condition relationships (one rule node must be executed before another rule node), and optimization relationships (experience nodes optimize rule nodes).
[0051] Perform feature extraction. Use a graph embedding algorithm with a multi - head attention mechanism to extract the features of business rule data and expert experience data in the knowledge graph. Business rule data comes from the triples of rule nodes, and expert experience data comes from the structured information of experience nodes. The feature extraction process focuses on the nodes directly related to the current processing node. By calculating attention scores and aggregating features, a business feature vector containing information about the interaction relationships of RPA operation nodes is generated. For example, when processing the "Extract invoice information" node, it will pay attention to the "Received invoice email" node with a "pre - condition relationship" and the "If the invoice format is incorrect" node with an "optimization relationship", calculate the attention scores of these nodes, and aggregate their features with the features of the "Extract invoice information" node to generate the business feature vector of this node.
[0052] Perform scenario migration. Train a scenario migration model based on the business feature vector. This model includes a discriminator network and a generator network. The generator network maps the business feature vector of the source scenario to a preset scenario feature space to generate migration samples. The discriminator network distinguishes between real process samples and generated migration samples in the target scenario. Use a contrast learning mechanism to align the feature spaces of the source scenario and the target scenario. Construct positive sample pairs from the same processing steps in the source scenario and the target scenario. For example, if there is an "Extract invoice information" step in both the source scenario and the target scenario, then these two steps are constructed as a positive sample pair. Construct negative sample pairs from different processing steps. For example, the "Send email notification" step in the source scenario and the "System entry" step in the target scenario. By minimizing the contrast loss function, the feature representations of the same operation steps approach each other in the feature space, and finally, scenario mapping data is obtained. For example, migrate the model for processing the "Online order" scenario to the "Offline order" scenario. Through contrast learning, the feature representations of the "Order processing" step in the two scenarios approach each other, thus realizing cross - scenario knowledge transfer.
[0053] Through knowledge transfer, the RPA process of the present invention can adapt to new scenarios without reconfiguration, reducing the development and maintenance costs of the RPA process; by integrating expert experience, optimizing the RPA operation procedures, improving the execution efficiency of the RPA process, and reducing manual intervention; by handling abnormal scenarios, improving the fault tolerance of the RPA process, and enhancing its stability in complex environments.
[0054] In an alternative embodiment, The scenario migration module trains a scenario migration model based on the business feature vectors. The scenario migration model includes a discriminator network and a generator network. The discriminator network is used to distinguish between real process samples and generated migration samples of the target scenario. The steps for the generator network to map the business feature vectors of the source scenario to a preset scenario feature space include: The generator network adopts a three-layer fully connected neural network structure, and the layers are connected by ReLU activation functions. The last layer outputs the target scenario feature vectors through a hyperbolic tangent activation function. The discriminator network adopts a four-layer fully connected network structure, and a batch normalization layer and a dropout layer are set after each hidden layer; a discriminator loss function is constructed based on the Wasserstein distance. The discriminator loss function includes an expected term for real samples, an expected term for generated samples, and a gradient penalty term. The discriminator network is trained through the discriminator loss function; a generator loss function is constructed. The generator loss function includes an adversarial loss term output by the discriminator and a semantic consistency loss term. The semantic consistency loss term is obtained by calculating the Euclidean distance between the source scenario features and the generated features and the Euclidean distance between the intermediate layer features of the feature extraction network; A feature quality evaluation index is constructed. The feature quality evaluation index includes a feature consistency score and a task relevance score. The feature consistency score is calculated based on the feature space distance between the target scenario features and the generated features, and the task relevance score is calculated based on the accuracy of the downstream task classifier; The scene transfer model is trained using an alternating optimization strategy. The generator network is fixed, and the parameters of the discriminator network are optimized through multiple iterations. Then, the discriminator network is fixed, and the parameters of the generator network are optimized once. The business feature vectors are divided into training scales according to the number of rule node triples. The number of rule node triples for the initial training scale is set to 1, the growth step value of the number of rule node triples is set, and the number of training epochs for each training scale is set. The scene transfer model is trained based on the business feature vectors of the current training scale. When the feature quality evaluation index reaches a preset threshold, the number of rule node triples is increased according to the growth step value, and the scene transfer model is continuously trained based on the increased training scale until the preset number of training times is reached. The generated target scene features are evaluated based on the feature quality evaluation index, and the weight coefficient of the semantic consistency loss term and the growth step value of the number of rule node triples are adjusted according to the evaluation results of the target scene features.
[0055] Exemplarily, process sample data of the source scene and the target scene are collected, and the corresponding business feature vectors are extracted. For example, the source scene can be the product recommendation scene of an e-commerce platform, and the target scene can be the friend recommendation scene of a social platform. The business feature vectors can include information such as the age, gender, hobbies, and purchase history of users. Suppose we have collected 10,000 source scene samples and 10,000 target scene samples, and each sample is represented as a 20-dimensional feature vector.
[0056] A scene transfer model is constructed, which includes a generator network and a discriminator network. The generator network adopts a three-layer fully connected neural network structure, with ReLU activation functions connected between each layer, and the hyperbolic tangent activation function is used in the last layer to output the target scene feature vector. For example, the input of the generator network is the 20-dimensional feature vector of the source scene, and the output is the 20-dimensional feature vector of the target scene. The first layer has 50 neurons, the second layer has 100 neurons, and the third layer has 20 neurons. The discriminator network adopts a four-layer fully connected network structure, with a batch normalization layer and a dropout layer set after each hidden layer. For example, the input of the discriminator network is a 20-dimensional feature vector, and the output is a value between 0 and 1, representing the probability that the input feature belongs to the target scene. The first layer has 100 neurons, the second layer has 50 neurons, the third layer has 20 neurons, and the fourth layer has 1 neuron.
[0057] Next, define the loss function and train the scene transfer model. The discriminator loss function is constructed based on the Wasserstein distance, including the expected term of real samples, the expected term of generated samples, and the gradient penalty term. The generator loss function includes the adversarial loss term output by the discriminator and the semantic consistency loss term. The semantic consistency loss term is obtained by calculating the Euclidean distance between the source scene features and the generated features, as well as the Euclidean distance between the features in the middle layer of the feature extraction network. An alternating optimization strategy is used to train the scene transfer model, that is, fix the generator network and optimize the parameters of the discriminator network iteratively for multiple times, and then fix the discriminator network and optimize the parameters of the generator network once.
[0058] During the training process, a strategy of gradually increasing the scale of training data is adopted. The business feature vectors are divided into training scales according to the number of rule node triples. The number of rule node triples in the initial training scale is set to 1, and the growth step value of the number of rule node triples is set, for example, the step value is 10. The number of training epochs for each training scale also needs to be set, for example, set to 100 epochs. Train the scene transfer model based on the business feature vectors of the current training scale. When the feature quality evaluation index reaches the preset threshold, increase the number of rule node triples according to the growth step value, and continue to train the scene transfer model based on the increased training scale until the preset number of training times is reached, for example, the preset number of training times is 1000 epochs. The feature quality evaluation index includes the feature consistency score and the task relevance score. The feature consistency score is calculated based on the feature space distance between the target scene features and the generated features, and the task relevance score is calculated based on the accuracy of the downstream task classifier. For example, the generated target scene features can be used for the user recommendation task, and the accuracy of the recommendation is calculated as the task relevance score. During the training process, adjust the weight coefficient of the semantic consistency loss term and the growth step value of the number of rule node triples according to the results of the feature quality evaluation index.
[0059] By combining the Wasserstein distance and gradient penalty, the present invention can effectively stabilize the training process of the GAN (Generative Adversarial Network), avoid problems such as mode collapse, and thus generate more diverse and realistic target scene features; the introduction of the semantic consistency loss term ensures the semantic consistency between the generated features and the source scene features, thereby improving the effectiveness of cross-scene knowledge transfer; by gradually increasing the scale of training data and dynamically adjusting parameters, the model can better adapt to different data distributions and task requirements, thereby improving the generalization ability of the model.
[0060] In an alternative embodiment, The knowledge transfer processing network further includes a verification module that validates the validity of the scenario mapping data and outputs a verification result. The steps of updating the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaboration processing framework based on the verification result include: Construct a verification matrix to verify the scenario mapping data. The verification matrix includes a semantic consistency vector, a structural similarity vector, and a temporal rationality vector. The semantic consistency vector is obtained by calculating the keyword matching degree of business rules before and after mapping. The structural similarity vector is calculated based on the edit distance of the graph structure. The temporal rationality vector is obtained by verifying the dependency relationship of the operation sequence. Construct a multi-level verification index based on the verification matrix. The multi-level verification index includes a rule layer verification index and a process layer verification index. The rule layer verification index is obtained by calculating the rule similarity and the internal consistency of the rules. The process layer verification index is obtained by calculating the temporal constraint satisfaction degree and the structural integrity. Weightedly combine the multi-level verification index with the business scenario coverage rate to generate a verification evaluation score. Update the parameters of the data perception processing network based on the verification evaluation score, including: performing gradient update on the perception layer parameters using an adaptive learning rate, where the adaptive learning rate is dynamically adjusted according to the verification evaluation score; determining the adjustment direction and adjustment step size of the data perception processing network parameters based on the difference between the verification evaluation score and a preset threshold. Construct a task allocation matrix, where the elements of the task allocation matrix represent the probability of tasks being assigned to processing units. Calculate the update amount of the task allocation matrix based on the verification evaluation score and the resource utilization rate; dynamically adjust the task allocation scheme in the human-machine collaboration processing framework according to the update amount of the task allocation matrix, and optimize the task allocation parameters in the human-machine collaboration processing framework based on the adjusted task allocation scheme.
[0061] Exemplarily, obtain the scenario mapping data. The scenario mapping data refers to converting the data in the actual business scenario into a data form that the knowledge transfer processing network can understand and process. For example, in an order processing scenario of an e-commerce platform, the scenario mapping data may include order information, user information, product information, etc.
[0062] Validate the validity of the scenario mapping data. The verification module uses the verification matrix to verify the scenario mapping data. The verification matrix consists of three vectors: a semantic consistency vector, a structural similarity vector, and a temporal rationality vector.
[0063] The calculation method of the semantic consistency vector is as follows: extract the keywords in the business rules before and after mapping, and calculate the matching degree of the keywords. For example, if the business rule before mapping is "Orders with an amount greater than 100 yuan are free of shipping fees", and the business rule after mapping is "Orders with an amount exceeding 100 yuan are free of postage", then the keywords "Order amount", "100 yuan", and "Free of shipping fees / Free of postage" can be extracted, and their matching degrees can be calculated.
[0064] The calculation method of the structural similarity vector is as follows: convert the business rules before and after mapping into graph structures, and calculate the edit distance of the graph structures. The edit distance refers to the minimum number of operations required to convert one graph into another (such as adding nodes, deleting nodes, modifying the connections between nodes). The smaller the edit distance, the higher the structural similarity.
[0065] The calculation method of the temporal rationality vector is as follows: verify the dependency relationships of the operation sequences. For example, in an order processing process, the order placement operation must be before the payment operation, and the shipping operation must be after the payment operation. If the operation sequence does not satisfy these dependency relationships, it is considered temporally unreasonable.
[0066] Construct multi-level verification metrics based on the verification matrix. The multi-level verification metrics include rule layer verification metrics and process layer verification metrics.
[0067] The calculation method of the rule layer verification metrics is as follows: calculate the rule similarity and the internal consistency of the rules. The rule similarity refers to the degree of similarity of the business rules before and after mapping. The internal consistency of the rules refers to whether the logic of the business rules themselves is consistent.
[0068] The calculation method of the process layer verification metrics is as follows: calculate the temporal constraint satisfaction degree and the structural integrity. The temporal constraint satisfaction degree refers to whether the operation sequence satisfies the temporal constraints. The structural integrity refers to whether the structure of the flowchart is complete.
[0069] Weightedly combine the multi-level verification metrics with the business scenario coverage rate to generate a verification evaluation score. The business scenario coverage rate refers to the proportion of the business scenarios covered by the verification module. For example, if the verification module covers 80% of the business scenarios, the business scenario coverage rate is 0.8.
[0070] Update the parameters of the data-aware processing network based on the verification evaluation score. The parameters of the data-aware processing network include the parameters of the perception layer. The update method for the parameters of the perception layer is as follows: Gradient update the parameters of the perception layer using an adaptive learning rate. The adaptive learning rate is dynamically adjusted according to the verification evaluation score. If the verification evaluation score is high, it indicates that the performance of the model is good, and the learning rate can be decreased; if the verification evaluation score is low, it indicates that the performance of the model is poor, and the learning rate can be increased. Meanwhile, determine the adjustment direction and adjustment step size of the parameters of the data-aware processing network according to the difference between the verification evaluation score and the preset threshold. If the verification evaluation score is lower than the preset threshold, the parameters need to be adjusted to improve the performance of the model.
[0071] Construct a task assignment matrix. The elements of the task assignment matrix represent the probability of tasks being assigned to processing units. Calculate the update amount of the task assignment matrix based on the verification evaluation score and the resource utilization rate. The resource utilization rate refers to the utilization rate of processing units. Dynamically adjust the task assignment scheme in the human-machine collaborative processing framework according to the update amount of the task assignment matrix, and optimize the task assignment parameters in the human-machine collaborative processing framework based on the adjusted task assignment scheme.
[0072] For example, assume that the verification evaluation score is 0.9 and the resource utilization rate is 0.8, then the update amount of the task assignment matrix can be set to 0.1. This means that 10% of the tasks will be reassigned to different processing units.
[0073] Through the validity verification of the scenario mapping data by the verification module in the present invention, the accuracy and efficiency of knowledge transfer can be improved, and the risk of incorrect transfer can be reduced; by dynamically adjusting the task assignment scheme in the human-machine collaborative processing framework, the efficiency of human-machine collaboration can be optimized, and the overall processing performance can be improved; by gradient-updating the parameters of the data-aware processing network using an adaptive learning rate, the self-adaptability of the system can be enhanced, enabling it to better adapt to different business scenarios.
[0074] Figure 2 This is a schematic structural diagram of a human-machine collaborative system based on robotic process automation according to an embodiment of the present invention, as Figure 2 shown, the system includes: A first unit, configured to construct a data-aware processing network, collect user operation instruction data, system operation status data, and task environment data, perform semantic parsing on the collected data to obtain task feature vectors, match the task feature vectors with a preset business model to construct a task mapping matrix, and generate a task processing strategy; A second unit for constructing a human-machine collaborative processing framework based on a task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing exceeds the specified threshold, the robot unit sends the historical data analysis result and decision-making suggestion information to the manual operation unit, and the manual operation unit transmits the processing procedure information to the rule database of the robot unit. A third unit for constructing a knowledge transfer processing network, which uses a graph embedding algorithm to extract feature vectors of business rule data and expert experience data during the human-machine collaborative processing to obtain business feature vectors, trains a scenario transfer model based on the business feature vectors, and maps the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer processing network further includes a verification module that validates the effectiveness of the scenario mapping data and outputs a verification result, and updates the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaborative processing framework based on the verification result.
[0075] In the third aspect of the embodiments of the present invention, A kind of electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0076] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0077] The present invention can be a method, device, system, and / or computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0078] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A human-machine collaboration method based on robotic process automation, characterized in that, Including: Construct a data-aware processing network to collect user operation instruction data, system operation status data, and task environment data, perform semantic parsing on the collected data to obtain task feature vectors, match the task feature vectors with a preset business model to construct a task mapping matrix, and generate a task processing strategy; Construct a human-machine collaborative processing framework based on the task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing process is greater than the specified threshold, the robot unit sends the historical data analysis result and decision-making suggestion information to the manual operation unit, and the manual operation unit transmits the processing procedure information to the rule database of the robot unit; Construct a knowledge transfer processing network, use the graph embedding algorithm to extract features from the business rule data and expert experience data during the human-machine collaborative processing process to obtain business feature vectors, train a scenario transfer model based on the business feature vectors, and map the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer processing network also includes a verification module to verify the validity of the scenario mapping data and output a verification result, and update the processing parameters of the data-aware processing network and the task allocation parameters of the human-machine collaborative processing framework based on the verification result.
2. The method according to claim 1, wherein The steps of matching the task feature vectors with a preset business model to construct a task mapping matrix and generate a task processing strategy include: Construct a model library containing standard feature vectors and processing templates, calculate the similarity between the task feature vectors and the standard feature vectors using contrastive learning, select the feature space with the highest similarity for mapping, and construct a task mapping matrix; Input the task mapping matrix into a hierarchical decoder, which includes a task decomposition layer, a resource allocation layer, and an execution sequence layer connected in sequence. The task decomposition layer generates a task decomposition strategy based on the task mapping matrix, the resource allocation layer generates a resource allocation strategy based on the task mapping matrix and the task decomposition strategy, and the execution sequence layer generates an execution sequence strategy based on the task mapping matrix, the task decomposition strategy, and the resource allocation strategy; Calculate a constraint adaptation factor based on the task feature vectors, the resource status output of the resource allocation layer, and the time constraint parameters, and perform constraint adjustment on the task decomposition strategy, the resource allocation strategy, and the execution sequence strategy respectively based on the constraint adaptation factor to obtain a task processing strategy; construct a dominance relationship matrix of the task processing strategy, calculate an optimal strategy combination based on the dominance relationship matrix, and determine the strategy weights according to the contribution degrees of the task decomposition strategy, the resource allocation strategy, and the execution sequence strategy to the optimal strategy combination; Input the strategy weights and the optimal strategy combination into a strategy optimization network, perform importance sampling based on historical strategy data, calculate strategy evaluation parameters, and then optimize and update the parameters of the hierarchical decoder.
3. The method according to claim 1, characterized in that, The steps of constructing a human-machine collaborative processing framework based on the task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks include: Obtain the task type feature vector based on the task processing strategy, calculate the entropy value of the task type feature vector based on the information entropy calculation model, and divide the task into a structured task and an unstructured task according to the entropy value; Construct a policy mapping mechanism for the robot execution unit based on the task processing strategy, map the task decomposition strategy in the task processing strategy to the robot basic action library to obtain an action correspondence matrix, determine the action execution parameters according to the resource allocation strategy in the task processing strategy, generate a robot control instruction stream based on the execution sequence strategy in the task processing strategy, and the robot execution unit processes the structured task according to the robot control instruction stream; Construct an operator cognitive state vector including attention level, workload, and professionalism, evaluate the operator cognitive state vector based on the resource allocation strategy in the task processing strategy, dynamically adjust the information display density and frequency of the manual operation interface according to the evaluation result, and process the unstructured task through the manual operation interface; Construct a collaborative effectiveness evaluation index including execution efficiency, accuracy, and cognitive load, determine the collaborative weight coefficient of the collaborative effectiveness evaluation index according to the strategy weight in the task processing strategy, calculate the human-machine collaborative processing effectiveness score based on the collaborative weight coefficient, and use the human-machine collaborative processing effectiveness score as a feedback signal to optimize and update the task processing strategy.
4. The method according to claim 1, wherein The robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing exceeds the specified threshold, the steps for the robot unit to send the historical data analysis result and decision-making advice information to the manual operation unit, and for the manual operation unit to transmit the processing procedure information to the rule database of the robot unit include: Extract the multi-scale time-frequency features of the task execution trajectory through wavelet transform, obtain the global deviation feature through adaptive attention weighted fusion, model the global deviation feature sequence using a recurrent neural network, calculate the deviation degree score through a multi-layer perceptron, and trigger an anomaly detection signal when the deviation degree score exceeds the preset deviation threshold; match the type and cause of the current anomaly from the anomaly pattern feature library based on the anomaly detection signal; Construct a multi-level knowledge graph structure including a task layer graph, a resource layer graph, and a processing layer graph. The task layer graph describes the dependencies between tasks, the resource layer graph describes the distribution of equipment and personnel capabilities, the processing layer graph describes the operation process, and update the node information in the multi-level knowledge graph structure based on the anomaly type and cause; The graph neural network with multi-head attention mechanism performs inter-layer information transfer on the multi-level knowledge graph structure, generates node feature representations, and combines them with historical sequence features to construct a state vector; trains a policy network based on the state vector, uses the weighted value of task completion and resource utilization as the reward signal, and optimizes to obtain a decision model through the policy gradient method; under the guidance of the decision model, uses access count, value estimation, and prior probability as search attributes, and adopts the Monte Carlo tree search algorithm to generate an optimal decision path; constructs a local causal graph containing variable nodes and causal edges for the nodes in the optimal decision path, and generates processing suggestions; Represents the processing suggestions as a processing procedure of conditional action pairs, evaluates the importance of the processing procedure, and updates the processing procedures with importance exceeding the preset importance threshold to the rule database, and the update result of the rule database is used to adjust the preset deviation threshold.
5. The method according to claim 1, wherein Constructs a knowledge transfer processing network, and uses a graph embedding algorithm to extract business feature vectors from business rule data and expert experience data in the human-machine collaborative processing process. The steps of training a scenario transfer model based on the business feature vectors and mapping the business feature vectors to a preset scenario feature space include: The knowledge transfer processing network includes a knowledge graph construction module, a feature extraction module, and a scenario transfer module; The knowledge graph construction module is used to construct a knowledge graph including rule nodes, experience nodes, and relationship edges. The rule nodes describe the robot process automation operation procedures in the form of a triple of trigger conditions, operation sequences, and constraint conditions. The experience nodes represent process optimization techniques in a structured form of abnormal scenarios, processing steps, and expected effects. The relationship edges describe the inclusion relationship, precondition relationship, and optimization relationship between the rule nodes and the experience nodes; The feature extraction module uses the multi-head attention graph embedding algorithm to extract business rule data and expert experience data from rule nodes and experience nodes respectively, and obtains business feature vectors containing the interaction relationships of operation nodes by calculating and aggregating the attention scores of adjacent nodes; The scenario transfer module trains a scenario transfer model based on the business feature vectors. The scenario transfer model includes a discriminator network and a generator network; Adopts a contrast learning mechanism to align the feature spaces of the source scenario and the target scenario, constructs positive sample pairs from the same processing steps in the source scenario and the target scenario, constructs different processing steps as negative samples, and makes the feature representations of the same operation steps approach in the feature space by minimizing the contrast loss function to obtain scenario mapping data.
6. The method according to claim 5, characterized in that The scenario transfer module trains a scenario transfer model based on the business feature vectors. The scenario transfer model includes a discriminator network and a generator network. The discriminator network is used to distinguish the real process samples of the target scenario and the generated transfer samples. The steps of the generator network for mapping the business feature vectors of the source scenario to a preset scenario feature space include: The generator network adopts a three-layer fully connected structure, uses ReLU activation and outputs the target scene features through tanh; the discriminator network adopts a four-layer fully connected structure, configured with batch normalization and dropout layers; the discriminator loss function is constructed based on the Wasserstein distance, including the expectation of real samples, the expectation of generated samples and the gradient penalty term; the generator loss function combines the adversarial loss output by the discriminator and the semantic consistency loss of the source scene features, where the semantic consistency is calculated by the Euclidean distance between the feature space and the intermediate layer features; Construct a feature quality evaluation index, which includes a feature consistency score and a task relevance score. The feature consistency score is calculated based on the feature space distance between the target scene features and the generated features, and the task relevance score is calculated based on the accuracy of the downstream task classifier; Adopt an alternating optimization strategy to train the scene migration model. Fix the generator network and optimize the parameters of the discriminator network iteratively for multiple times. Fix the discriminator network and optimize the parameters of the generator network once; divide the business feature vectors according to the number of rule node triples in the training scale, set the number of rule node triples in the initial training scale to 1, set the growth step value of the number of rule node triples, and set the number of training epochs for each training scale; train the scene migration model based on the business feature vectors of the current training scale. When the feature quality evaluation index reaches the preset threshold, increase the number of rule node triples according to the growth step value, and continue to train the scene migration model based on the increased training scale until the preset number of training times is reached; evaluate the generated target scene features based on the feature quality evaluation index, and adjust the weight coefficient of the semantic consistency loss term and the growth step value of the number of rule node triples according to the evaluation results of the target scene features.
7. The method according to claim 1, characterized in that, The knowledge transfer processing network further includes a verification module that validates the effectiveness of the scene mapping data and outputs a verification result. The steps of updating the processing parameters of the data perception processing network and the task allocation parameters of the human-machine collaboration processing framework based on the verification result include: Construct a verification matrix to verify the scene mapping data. The verification matrix includes a semantic consistency vector, a structural similarity vector and a temporal rationality vector. The semantic consistency vector is obtained by calculating the keyword matching degree of the business rules before and after mapping, the structural similarity vector is calculated based on the edit distance of the graph structure, and the temporal rationality vector is obtained by verifying the dependency relationship of the operation sequence; Construct a multi-level verification index based on the verification matrix. The multi-level verification index includes a rule layer verification index and a process layer verification index. The rule layer verification index is obtained by calculating the rule similarity and the internal consistency of the rules, and the process layer verification index is obtained by calculating the temporal constraint satisfaction degree and the structural integrity; Weightedly combine the multi-level verification index with the business scenario coverage rate to generate a verification evaluation score; Update the parameters of the data-aware processing network based on the verification evaluation score, including: performing gradient update on the parameters of the perception layer using an adaptive learning rate, where the adaptive learning rate is dynamically adjusted according to the verification evaluation score; determining the adjustment direction and adjustment step size of the data-aware processing network parameters based on the difference between the verification evaluation score and a preset threshold. Construct a task assignment matrix, where the elements of the task assignment matrix represent the probability of tasks being assigned to processing units, and calculate the update amount of the task assignment matrix based on the verification evaluation score and resource utilization rate; dynamically adjust the task assignment scheme in the human-machine collaborative processing framework according to the update amount of the task assignment matrix, and optimize the task assignment parameters in the human-machine collaborative processing framework based on the adjusted task assignment scheme.
8. A human-machine collaboration system based on robotic process automation for implementing the method according to any one of the preceding claims 1-7, characterized in that Including: The first unit is used to construct a data-aware processing network, collect user operation instruction data, system operation status data, and task environment data, perform semantic parsing on the collected data to obtain task feature vectors, match the task feature vectors with a preset business model to construct a task mapping matrix, and generate task processing strategies. The second unit is used to construct a human-machine collaborative processing framework based on the task processing strategy, where the robot unit processes structured tasks that meet preset rules, and the manual operation unit processes unstructured tasks; the robot unit and the manual operation unit are connected through an information interaction channel. When the deviation value during the task processing process is greater than the specified threshold, the robot unit sends the historical data analysis result and decision-making advice information to the manual operation unit, and the manual operation unit transmits the processing procedure information to the rule database of the robot unit. The third unit is used to construct a knowledge transfer processing network, extract business feature vectors from the business rule data and expert experience data during the human-machine collaborative processing process using a graph embedding algorithm, train a scenario transfer model based on the business feature vectors, and map the business feature vectors to a preset scenario feature space to generate scenario mapping data; the knowledge transfer processing network also includes a verification module, which validates the effectiveness of the scenario mapping data and outputs a verification result, and updates the processing parameters of the data-aware processing network and the task assignment parameters of the human-machine collaborative processing framework based on the verification result.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Epidemic prevention robot knowledge learning and migration method and system
CN112231489A
Robot control system and method, storage medium, controller and robot
CN118927246A
Intelligent software robot standardization implementation method
CN119090436A
Igniter robot production line control method and system based on deep reinforcement learning
CN119579120A
Logistics robot control method and system based on artificial intelligence
CN120038760A
Cited By
Multi-physical field regulation and control method and system in aluminum product processing
CN120540253A
Active man-machine cooperation method driven by process knowledge and interactive semantics
CN120653970A
Intelligent identification method for electric power infrastructure line operation behavior
CN120853093A
An intelligent identification method for power infrastructure line operation behavior
CN120853093B
Iterative task execution method and system based on dynamic feedback and causal fault tolerance
CN121301846A