Task processing method and device, intelligent agent and storage medium

By combining the pre-trained language model with the knowledge graph, and using the knowledge graph to correct the execution action sequence output by the pre-trained language model, the problem of low success rate of agent processing tasks is solved, and the success rate of task processing and resource utilization efficiency is improved.

CN120216136APending Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312679.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The success rate of agents when processing tasks is low, mainly because the execution action sequence decomposed by the pre-trained language model does not match the actual scenario.

Method used

Combining the pre-trained language model and knowledge graph, the execution action sequence is obtained through semantic decomposition, and the knowledge graph is used to correct the mismatched execution actions to obtain the corrected execution action sequence.

Benefits of technology

The success rate of the agent processing task is improved, making the final execution action sequence more in line with the actual application scenario, reducing the number of times of generating prompt information, and reducing the consumption of agent resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216136A_ABST
    Figure CN120216136A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method and device, an intelligent agent and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an execution instruction of a to-be-processed task, carrying out the semantic decomposition of the execution instruction through a pre-training language model, so as to obtain an execution action sequence of the to-be-processed task, and correcting the execution action sequence of the to-be-processed task through the knowledge graph to obtain a corrected execution action sequence, and executing the to-be-processed task according to the corrected execution action sequence. According to the method, the pre-training language model is combined with the knowledge graph model, and the knowledge graph is utilized to correct the execution action sequence output by the pre-training language model, so that the finally obtained execution action sequence is more in line with an actual application scene, and therefore, the user experience is improved. The problem that the task processing success rate of an intelligent agent is low in the related technology can be solved, and the technical effect of improving the task processing success rate is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, apparatus, electronic device, and storage medium for task processing. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence technology has been widely applied in various fields, and it is of great significance for an intelligent agent to process tasks through artificial intelligence technology.

[0003] In related technologies, an intelligent agent usually decomposes a task to be processed through a pre-trained language model to obtain an execution action sequence, and then processes the task to be processed according to the execution action sequence. However, the success rate of task processing by the intelligent agent in the above manner is relatively low. Summary of the Invention

[0004] This application provides a method, apparatus, intelligent agent, and storage medium for task processing, so as to at least solve the problem that the success rate of task processing by an intelligent agent in related technologies is relatively low.

[0005] This application provides a method for task processing, including:

[0006] Obtaining an execution instruction of a task to be processed;

[0007] Based on a pre-trained language model, semantically decomposing the execution instruction to obtain an execution action sequence of the task to be processed;

[0008] Based on a knowledge graph, correcting the execution actions in the execution action sequence that are different from the execution actions pre-stored in the knowledge graph to obtain a corrected execution action sequence;

[0009] Processing the task to be processed according to the corrected execution action sequence.

[0010] This application also provides a device for task processing, including:

[0011] An obtaining module, configured to obtain an execution instruction of a task to be processed;

[0012] A processing module, configured to semantically decompose the execution instruction based on a pre-trained language model to obtain an execution action sequence of the task to be processed;

[0013] A correction module, configured to correct the execution actions in the execution action sequence that are different from the execution actions pre-stored in the knowledge graph based on the knowledge graph to obtain a corrected execution action sequence;

[0014] The processing module is further configured to process the task to be processed according to the corrected execution action sequence.

[0015] The present application also provides an agent, including: a memory for storing a computer program; a processor for implementing the steps of the processing method of any one of the above tasks when executing the computer program.

[0016] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the processing method of any one of the above tasks.

[0017] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of the processing method of any one of the above tasks.

[0018] Through the present application, after obtaining the execution instruction of the task to be processed, the execution instruction is semantically decomposed by a pre-trained language model to obtain the execution action sequence of the task to be processed, and the execution action sequence of the task to be processed is corrected by a knowledge graph to obtain the corrected execution action sequence, and then the task to be processed is executed according to the corrected execution action sequence. The method of the present application combines the pre-trained language model with the knowledge graph model, and uses the knowledge graph to correct the execution action sequence output by the pre-trained language model, so that the finally obtained execution action sequence is more in line with the actual application scenario. Therefore, the problem of low success rate of task processing by an agent in the related art can be solved, and the technical effect of improving the success rate of task processing can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0021] Figure 2 It is a schematic flowchart of a method for processing a task provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic flowchart of a method for correcting an execution action different from a pre-stored execution action in a knowledge graph in an execution action sequence provided by an embodiment of the present application;

[0023] Figure 4 It is a schematic flowchart of a method for updating a knowledge graph provided by an embodiment of the present application;

[0024] Figure 5Exemplary schematic diagram of a method for processing a task provided by an embodiment of the present application;

[0025] Figure 6 Structural schematic diagram of a device for processing a task provided by an embodiment of the present application;

[0026] Figure 7 Structural schematic diagram of an agent provided by the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0028] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0030] With the rapid development of artificial intelligence technology, agents are widely used in multiple scenarios because they can assist users in completing diverse tasks. For example, in the home service scenario, an agent can use artificial intelligence technology to cook various delicious foods for the user according to different recipes to perform a cooking task, or automatically clean the room for the user to perform a cleaning task, etc. Another example is that in the industrial service scenario, an agent can use artificial intelligence technology to adapt to assembly and inspection tasks brought about by product upgrades and replacements.

[0031] In the related art, after an agent obtains a task to be processed, it will use a pre-trained language model that has been trained to perform semantic decomposition on the task to be processed, resulting in multiple subtasks. For example, after the agent obtains the task instruction of "frying steak", the pre-trained language model performs semantic decomposition to obtain multiple subtasks, which are steps such as "preparing the steak", "preparing the seasonings", "cutting the steak", and "frying the steak in a baking oven". The agent starts to execute the production of braised pork according to the multiple decomposed subtasks.

[0032] Although the pre-trained language model can decompose tasks, as the agent faces continuous changes in new tasks and scenarios, the generated subtasks are prone to mismatches with the actual scenario, resulting in a low success rate for the agent to process tasks. For example, if the kitchen only includes an ordinary stove, when it comes to the step of "frying the steak in a baking oven", the agent cannot complete the production of fried steak because it cannot use the baking oven, resulting in cooking failure.

[0033] Therefore, in view of the above technical problems in the prior art, the inventor found during the research that if a knowledge graph is used to pre-store the execution actions that match the current actual scenario in the knowledge graph, when the multiple subtasks decomposed by the pre-trained language model, that is, the execution action sequence, are corrected through the knowledge graph to obtain an execution action sequence that matches the current actual scenario, the success rate of the agent in task processing can be improved. Therefore, this application proposes a method for collaborative operation of a knowledge graph and a pre-trained language model to process tasks. Specifically, the agent obtains the execution instruction of the task to be processed, and based on the pre-trained language model, performs semantic decomposition on the execution instruction to obtain the execution action sequence of the task to be processed. Then, based on the knowledge graph, the execution actions in the execution action sequence that are different from the execution actions pre-stored in the knowledge graph are corrected to obtain a corrected execution action sequence, and the task to be processed is processed based on the corrected execution action sequence.

[0034] To facilitate the understanding of this application, the following will be described through Figure 1 an example application scenario. Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application scenario provided by an embodiment of this application. In this scenario, there is an agent 01, and in the agent, there are a pre-trained language model module 011, a knowledge graph module 012, a decision-making module 013, and an execution module 014.

[0035] Specifically, after the decision-making module 013 obtains the execution instruction of the task to be processed input by the user 02, it sends the execution instruction to the pre-trained language model module 011. The pre-trained language model module 011 decomposes the execution instruction according to its own decomposition ability to obtain the execution action sequence of the task to be processed, and sends the execution action sequence to the knowledge graph module 012. The knowledge graph module 012 corrects the execution action sequence according to the execution actions stored in advance by itself to obtain the corrected execution action sequence, and sends the corrected execution action sequence to the execution module 014. The execution module 014 processes the task to be processed according to the corrected execution action sequence.

[0036] It can be understood that the above examples are only for illustration and do not limit the present application.

[0037] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for processing a task provided by an embodiment of the present application. The execution subject of this method is an intelligent agent, and the intelligent agent can be a robot, etc. This method may include the following steps:

[0039] S201. Obtain the execution instruction of the task to be processed.

[0040] In the present application, before obtaining the execution instruction of the task to be processed, a knowledge graph is constructed in advance according to the application scenario. The knowledge graph is a technical model that describes entities in the real world and their relationships in a structured form. It constructs a semantic network by associating scattered data to better understand and reason about knowledge. The construction process of the knowledge graph can be as follows:

[0041] Construct the state graph G s and the attribute graph G k , where the state graph is used to record the instance information of objects in a certain scenario. Taking the kitchen scenario as an example, the state graph can be, for example, (potato1, obj_location, cabinet), indicating that the potato is placed in the cabinet. According to the state graph, the entity can be determined, and the entity in the above example is potato1. The attribute graph is used to record the attributes and action capabilities of object classes. The state graph can be, for example, (potato, peelable, true), indicating that the entity "potato" has the attribute "peelable" and the attribute value is "true". According to the attribute graph, the attribute information of the object class can be determined, so as to associate the attribute information with the corresponding entity.

[0042] Optionally, the triple data model format of the Resource Description Framework (RDF) can be adopted, and the state graph and property graph can be encoded in the Turtle format (.ttl file).

[0043] Entities are identified from the constructed state graph and property graph, relationships are identified based on the identified entities, and attributes are identified based on the identified entities and relationships. Then, the identified entities, relationships, and attributes are imported into a preset graph database to generate a knowledge graph.

[0044] Optionally, the user can input the task to be processed manually or by voice on the display interface of the intelligent agent, so that the intelligent agent can obtain the execution instruction of the task to be processed.

[0045] Exemplarily, the execution instruction of the task to be processed obtained by the intelligent agent can be "cooking shredded pork with shredded potatoes".

[0046] S202. Based on the pre-trained language model, semantically decompose the execution instruction to obtain the execution action sequence of the task to be processed.

[0047] In the pre-trained language model, "pre-training" is a strategy for training deep learning models. Its core lies in using a large-scale dataset to preliminarily train the model so that the model can learn general feature representations. This process is similar to the basic learning stage of humans before learning new knowledge, where they accumulate experience through extensive reading and observation.

[0048] Pre-trained language model: Generally, it refers to designing language model training tasks based on a large-scale corpus (including language training materials such as sentences and paragraphs), training a large-scale neural network algorithm structure to learn and implement. The finally obtained large-scale neural network algorithm structure and parameters are the pre-trained language model. For subsequent other tasks, feature extraction or task fine-tuning can be performed on this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain models adapted to other tasks. By pre-training on a large-scale corpus, the neural language representation model can learn powerful language representation capabilities and can extract rich syntactic and semantic information from the text. The pre-trained language model can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, or directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain downstream-specific models.

[0049] The neural network algorithm structure for pre-training a language model can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Long Short-Term Memory network (LSTM), etc., or it can be a model constructed with an attention network, such as a transformer, a Bidirectional Encoder Representations from Transformers (BERT), a Generative Pretrained Transformer (GPT), a Contrastive Language-Image Pretraining (CLIP), etc. This application does not make any limitations here. An attention network refers to a network model that uses an attention mechanism for training. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence, enabling the model to finally obtain a more accurate output.

[0050] "Fine-tuning" involved in the above content refers to further training on a dataset for a specific task based on the use of a pre-trained model to adjust the model parameters so that it can better adapt to the target task. During fine-tuning, most layers of the pre-trained model are usually frozen, and only the newly added layers or a small number of key layers are trained. Doing so can not only retain the features learned by the pre-trained model but also quickly adapt to the specific requirements of the new task. In addition, selecting appropriate learning rates and the number of training epochs is also the key to successful fine-tuning.

[0051] In this embodiment, the pre-trained language model can be a Large Language Model (LLM). A large language model is an artificial intelligence model based on deep learning technology that realizes the understanding, generation, and prediction of natural language through the training of a large amount of text data.

[0052] Based on a large amount of knowledge and common tasks used in the pre-training, the pre-trained language model decomposes the execution instruction into specific execution actions, and organizes the decomposed actions into an execution action sequence according to the logical relationship and sequence between the actions.

[0053] Exemplarily, for the execution instruction "make shredded potato with stir-fried pork" of the above task to be processed, the corresponding execution action sequence can include "find potatoes", "find pork", "peel the potatoes", "shred the potatoes", "slice the pork", "stir-fry in the pan", etc.

[0054] S203. Based on the knowledge graph, correct the execution action sequence that is different from the pre-stored execution actions in the knowledge graph to obtain a corrected execution action sequence.

[0055] Among them, the pre-stored execution actions in the knowledge graph can be determined according to the state graph and attribute graph in the knowledge graph.

[0056] The knowledge graph compares and verifies the execution actions included in the execution action sequence with the pre-stored execution actions in itself. If there are contents in the execution action sequence that do not match the pre-stored execution actions in itself, correct the unmatched contents to obtain a corrected execution action sequence.

[0057] S204. Process the task to be processed according to the corrected execution action sequence.

[0058] After the intelligent agent obtains the corrected execution action sequence, it sequentially executes according to each execution action included in the corrected execution action sequence according to the order of each execution action, so as to complete the processing of the task to be processed.

[0059] In the above embodiments of the present application, after the intelligent agent obtains the execution instruction of the task to be processed, it semantically decomposes the execution instruction through a pre-trained language model to obtain the execution action sequence of the task to be processed, and corrects the execution action sequence of the task to be processed through the knowledge graph to obtain a corrected execution action sequence, and then executes the task to be processed according to the corrected execution action sequence. The method of the present application combines the pre-trained language model with the knowledge graph model, and uses the knowledge graph to correct the execution action sequence output by the pre-trained language model, so that the finally obtained execution action sequence is more in line with the actual application scenario. Therefore, it can solve the problem that the success rate of the intelligent agent in processing tasks is relatively low in the related art, and achieve the technical effect of improving the success rate of task processing.

[0060] Further, on the basis of the above embodiments, the process of correcting the execution action sequence that is different from the pre-stored execution actions in the knowledge graph based on the knowledge graph to obtain a corrected execution action sequence is described in detail through the following embodiments.

[0061] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a method for correcting the execution action sequence that is different from the pre-stored execution actions in the knowledge graph provided by the embodiments of the present application. The method may include the following steps:

[0062] S301. Compare each execution action in the execution action sequence with the pre-stored execution actions in the knowledge graph, and determine the execution actions to be corrected from the execution action sequence.

[0063] For any execution action in the execution action sequence, the knowledge graph compares this action with the pre-stored execution actions in the knowledge graph, and determines the execution actions that fail to match as the execution actions to be corrected.

[0064] Exemplarily, assume that an execution action in the execution action sequence is "shred potatoes with a grater", but the pre-stored content in the knowledge graph does not include "grater", only includes "kitchen knife". Therefore, the knowledge graph will determine the execution action of "shred potatoes with a grater" as the execution action to be corrected in the execution action sequence.

[0065] S302. According to the execution action to be corrected, determine the pre-stored target execution action corresponding to the execution action to be corrected in the knowledge graph.

[0066] The knowledge graph determines the corresponding pre-stored target execution action according to the execution action to be corrected.

[0067] Taking the above example, assume that the execution action to be corrected is "shred potatoes with a grater". Since the knowledge graph does not include "grater" but only includes "kitchen knife", the determined corresponding pre-stored target execution action is "shred potatoes with a kitchen knife".

[0068] S303. According to the target execution action, correct the execution action to be corrected to obtain a corrected execution action sequence.

[0069] Optionally, the intelligent agent corrects the execution action to be corrected to the target execution action, and according to the positional relationship of the execution action to be corrected in the execution action sequence, determines the target position where the target execution action is located, updates the target execution action at the target position, and repeats the above steps, so as to correct the execution action to be corrected in the execution action sequence to obtain a corrected execution action sequence.

[0070] By adopting the method of directly comparing execution actions and considering the context relationship, the process for the intelligent agent to obtain a corrected execution action sequence is made more simple and intuitive.

[0071] Taking the above example, assume that the determined pre-stored target execution action is "shred potatoes with a kitchen knife". The intelligent agent replaces "shred potatoes with a grater" with "shred potatoes with a kitchen knife". At the same time, the intelligent agent considers the context relationship of this execution action in the execution action sequence, determines the position of the execution action of "shred potatoes with a kitchen knife" in the execution action sequence, and updates it to "shred potatoes with a kitchen knife" at this position, so as to obtain a corrected execution action sequence.

[0072] In the above embodiments of the present application, by comparing each execution action in the execution action sequence with the execution actions pre-stored in the knowledge graph, the execution actions to be corrected are determined from the execution action sequence, and according to the execution actions to be corrected, the pre-stored target execution actions corresponding to the execution actions to be corrected in the knowledge graph are determined. Furthermore, according to the target execution actions, the execution actions to be corrected are corrected to obtain a corrected execution action sequence. In the method of this embodiment, the knowledge graph improves the feasibility and accuracy of obtaining the corrected execution action sequence by correcting the execution actions. Furthermore, according to the accurate execution action sequence, the tasks to be processed are also more accurate.

[0073] In the above embodiment, if the agent cannot determine the pre-stored target execution action corresponding to the execution action to be corrected in the knowledge graph according to the execution action to be corrected, a prompt message is generated, where the prompt message is used to request to regenerate a new execution action sequence for the task to be processed through a pre-trained language model.

[0074] Exemplarily, if the execution action to be corrected is "carve and arrange potatoes on a plate", since this execution action is not pre-stored in the knowledge graph and the comparison is unsuccessful, the agent outputs a prompt message containing error or mismatch information.

[0075] After the agent generates the prompt message, through the pre-trained language model, a new execution action sequence for the task to be processed is regenerated according to the prompt message, and each execution action in the new execution action sequence is compared with the execution actions pre-stored in the knowledge graph to determine new execution actions to be corrected from the new execution action sequence. The specific implementation process is similar to the above embodiment. Please refer to the above embodiment and will not be repeated. If the new target execution action corresponding to the newly corrected execution action still cannot be determined in the pre-stored knowledge graph according to the newly corrected execution action, a request message is generated, and the knowledge graph is updated according to the request message. The request message is used to request feedback from the user for the purpose of enabling the user to update the knowledge graph according to the request message.

[0076] Next, through Figure 4 illustrate the process of updating the knowledge graph according to the request message. Figure 4 FIG. is a schematic flowchart of a method for updating a knowledge graph provided by an embodiment of the present application. The method may include the following steps:

[0077] S401. In response to an input operation by the user according to the request message, obtain a new target execution action corresponding to the task to be processed.

[0078] After the agent generates the request information, it can remind the user in the form of a warning. The warning methods include but are not limited to: the buzzer emits a beeping sound, the audible and visual alarm emits flashing lights and sounds, the voice alarm broadcasts voice, or it is displayed through a visual display screen, etc.

[0079] According to the request information, the user inputs the target execution actions corresponding to the task to be processed in the agent. The target execution actions include but are not limited to: detailed action descriptions, required tools, and operation steps, etc.

[0080] S402. According to the new target execution actions, identify the target object instance information and target object attribute information corresponding to the new target execution actions.

[0081] Based on the user feedback, that is, the new target execution actions, the agent identifies the target object instance information, such as the tool location, etc., and identifies the target object attribute information, such as, "potato, can be carved, true", relevant tools or operation steps, etc.

[0082] S403. Update the knowledge graph according to the target object instance information and target object attribute information.

[0083] In the Resource Description Framework (RDF) format, based on the preset text format, update the target object instance information to the state graph of the knowledge graph, and update the target object attribute information to the attribute graph of the knowledge graph.

[0084] Based on the updated knowledge graph, the agent corrects the execution action sequence that is different from the execution actions pre-stored in the knowledge graph, so as to obtain the corrected execution action sequence, and then can process the task to be processed according to the corrected execution action sequence.

[0085] In the above embodiments of the present application, in response to the input operation of the user according to the request information, obtain the new target execution actions corresponding to the task to be processed, according to the new target execution actions, identify the target object instance information and target object attribute information corresponding to the new target execution actions, and update the knowledge graph according to the target object instance information and target object attribute information. The method of this embodiment, when the agent still cannot determine the new target execution actions pre-stored in the knowledge graph corresponding to the new execution actions to be corrected according to the new execution actions to be corrected, that is, when the agent encounters an unsolvable problem during the task execution process, by generating the request information, it can timely request the user to input, and update the knowledge graph according to the user input, thereby improving the incremental accumulation of the agent's knowledge and improving the adaptability and flexibility of the agent in a complex environment.

[0086] To facilitate a better understanding of the present application, the following uses specific examples to illustrate the method of the present application. Please refer toFigure 5 , Figure 5 This is an exemplary schematic diagram of a method for processing a task provided by an embodiment of the present application. Herein, the intelligent agent is taken as an example of a robot.

[0087] Suppose the execution instruction obtained by the robot is "Prepare an apple omelette and cake toast". Based on the pre-trained language model, semantic decomposition is performed on this execution instruction to obtain an execution action sequence, and the execution action sequence is respectively "Pick up an apple", "Put the apple on the steamer", "Cook the apple", "Cut the cake", "Pick up a slice of cake", "Put the slice of cake into the toaster", "Toast the slice of cake", "Put the slice of cake on the plate", "Put the omelette on the plate", "Serve the omelette" and "Serve the toast".

[0088] The robot, based on the knowledge graph, compares the above execution action sequence with the execution actions pre-stored in the knowledge graph, and finds that "steamer" and "toaster" do not match. Then, it corrects the execution actions in the above execution action sequence that are different from the execution actions pre-stored in the knowledge graph. For example, it corrects "steamer" to "frying pan". However, since there is no oven-like device with a similar function to the "toaster" pre-stored in the knowledge graph, therefore, the execution actions pre-stored in the knowledge graph cannot be compared. So the robot re-outputs a new execution action sequence based on the pre-trained language model, and compares the new execution action sequence with the target execution actions pre-stored in the knowledge graph. If the comparison still fails, a prompt message is output.

[0089] The user updates the indication graph according to the prompt message, and the user updates and supplements the information related to the toaster in the knowledge graph. After updating the knowledge graph, the robot can then execute the following execution action sequence, that is, "Pick up an apple", "Put the apple on the frying pan", "Cook the apple", "Cut the cake", "Pick up a slice of cake", "Put the slice of cake into the toaster", "Toast the slice of cake", "Put the slice of cake on the plate", "Put the omelette on the plate", "Serve the omelette" and "Serve the toast", so as to complete the instruction "Prepare an apple omelette and cake toast".

[0090] For specific implementation details, please refer to the above-mentioned multiple embodiments. To avoid redundancy, no further description will be repeated.

[0091] Next, through specific example data, the technical effects achieved by the present application will be described.

[0092] Compared with directly processing the task to be processed only according to the execution action sequence decomposed by the pre-trained language model, the success rate of task execution in the method of this application is significantly improved. When only using version A of the pre-trained language model in the related technology, the success rate is increased from 45.2% to 56.95%. When only using version B of the pre-trained language model in the related technology, the success rate is increased from 25.41% to 33.95%. Taking the example of an agent performing a cleaning task, compared with the related technology, the success rate of the method of this application is increased from 32.63% to 98.75%.

[0093] Compared with directly processing the task to be processed only according to the execution action sequence decomposed by the pre-trained language model, the method of this application reduces the number of generated prompt messages, thereby reducing the average data (token) usage and the resource consumption of the agent. Taking the example of an agent performing a cooking task, compared with the related technology, when only using version A of the pre-trained language model in the related technology, the average data usage is reduced from 8316 to 6459. When only using version B of the pre-trained language model in the related technology, the average data usage is reduced from 8402 to 4353.

[0094] Compared with directly processing the task to be processed only according to the execution action sequence decomposed by the pre-trained language model, the method of this application improves the knowledge adaptability and accumulation ability of the agent by introducing user feedback and updating the knowledge graph according to the user's feedback, enabling the agent to better adapt to new tasks and new scenarios, thereby improving the execution ability of the agent.

[0095] The method of this application is more intuitive and easy to understand by directly matching the knowledge graph and considering the relationship between the replacement content and its context to obtain the corrected execution action sequence, which helps users better trust and use the agent, thereby enhancing the user experience.

[0096] This application can be widely applied in multiple fields, including but not limited to: industrial manufacturing field, medical care field, education field, and intelligent transportation field, etc.

[0097] Industrial manufacturing field: On the industrial production line, robots need to perform various complex assembly and processing tasks. By using the method of this application, the production task can be decomposed into a series of subtasks, and the domain knowledge about components, tools, and production processes can be provided by the knowledge graph to correct the execution action sequence of the robot. When encountering new components or production processes, the knowledge graph can be updated in a timely manner by obtaining the feedback from technicians, enabling the robot to quickly adapt to the changes in production tasks and improve production efficiency and product quality.

[0098] Medical care field: In the medical care scenario, intelligent care robots can utilize this application to, according to the care needs of patients, such as assisting patients in rehabilitation training, providing tasks like medicine management, decompose tasks through a pre-trained language model, and a knowledge graph provides professional knowledge in the medical field, such as medicine attributes. When encountering special situations or new care needs, caregivers can provide guidance to the robot through a feedback mechanism to update the knowledge graph and enhance the robot's care capabilities and adaptability.

[0099] Educational field: In an intelligent education assistance system, this application can achieve effective decomposition of learning tasks and knowledge learning. According to students' learning goals and course content, use a pre-trained language model to generate a sequence of learning tasks, and the knowledge graph provides the subject knowledge system and learning resource information. Feedback from teachers or students can be used to adjust learning tasks and update the knowledge graph to help students better understand and master knowledge and improve learning effects.

[0100] Intelligent transportation field: In an autonomous driving system, a vehicle needs to execute driving tasks based on information such as road conditions, traffic rules, and destinations. Using the method of this application, driving tasks can be decomposed into multiple subtasks, encode domain knowledge such as traffic rules and road information through a knowledge graph, and correct the vehicle's driving action sequence. When encountering special road conditions or new traffic rules, by obtaining feedback from traffic management departments or drivers, update the knowledge graph so that the autonomous driving system can better adapt to complex and changeable traffic environments and improve driving safety and efficiency.

[0101] It can be understood that the above fields are only for illustrative purposes and do not limit this application.

[0102] Figure 6 This is a schematic structural diagram of a task processing device provided by an embodiment of this application. As Figure 6 shown, an embodiment of this application also provides a task processing device, including:

[0103] An acquisition module 601, configured to acquire an execution instruction of a task to be processed.

[0104] A processing module 602, configured to perform semantic decomposition on the execution instruction based on a pre-trained language model to obtain an execution action sequence of the task to be processed.

[0105] A correction module 603, configured to correct the execution actions in the execution action sequence that are different from the execution actions pre-stored in the knowledge graph based on the knowledge graph to obtain a corrected execution action sequence.

[0106] The processing module 602 is further configured to process the task to be processed according to the corrected execution action sequence.

[0107] One possible implementation is that the correction module 603 is specifically used for:

[0108] Compare each execution action in the execution action sequence with the execution actions pre-stored in the knowledge graph, and determine the execution actions to be corrected from the execution action sequence.

[0109] According to the execution actions to be corrected, determine the pre-stored target execution actions corresponding to the execution actions to be corrected in the knowledge graph.

[0110] According to the target execution actions, correct the execution actions to be corrected to obtain a corrected execution action sequence.

[0111] One possible implementation is that the correction module 603 is specifically used for:

[0112] Correct the execution actions to be corrected to the target execution actions.

[0113] According to the positional relationship of the execution actions to be corrected in the execution action sequence, determine the target position where the target execution actions are located.

[0114] Update the target execution actions at the target position to obtain a corrected execution action sequence.

[0115] One possible implementation is that the processing module 602 is further used for:

[0116] If, according to the execution actions to be corrected, the pre-stored target execution actions corresponding to the execution actions to be corrected in the knowledge graph cannot be determined, then generate a prompt message, and the prompt message is used to request to regenerate a new execution action sequence for the task to be processed through a pre-trained language model.

[0117] One possible implementation is that after generating the prompt message, the processing module 602 is further used for:

[0118] Through the pre-trained language model, regenerate a new execution action sequence for the task to be processed according to the prompt message.

[0119] Compare each execution action in the new execution action sequence with the execution actions pre-stored in the knowledge graph, and determine the new execution actions to be corrected from the new execution action sequence.

[0120] If, according to the new execution actions to be corrected, the pre-stored new target execution actions corresponding to the new execution actions to be corrected in the knowledge graph still cannot be determined, then generate a request message.

[0121] In response to the input operation by the user according to the request message, obtain the new target execution actions corresponding to the task to be processed.

[0122] Execute actions according to the new goal, and identify the target object instance information and target object attribute information corresponding to the new goal execution action.

[0123] Update the knowledge graph according to the target object instance information and target object attribute information.

[0124] A possible implementation is that the processing module 602 is specifically further configured to:

[0125] Update the target object instance information to the state graph of the knowledge graph, and update the target object attribute information to the attribute graph of the knowledge graph.

[0126] A possible implementation is that the processing module 602 is specifically further configured to:

[0127] Under the Resource Description Framework format, based on a preset text format, update the target object instance information to the state graph of the knowledge graph.

[0128] Under the Resource Description Framework format, based on a preset text format, update the target object attribute information to the attribute graph of the knowledge graph.

[0129] For the description of the features in the corresponding embodiments of the processing device of the task, reference can be made to the relevant descriptions in the corresponding embodiments of the processing method of the task, which will not be elaborated here one by one.

[0130] Figure 7 This is a schematic structural diagram of the intelligent agent provided by the present application. As Figure 7 shown, the intelligent agent provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, a communication component 703 is further included. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus 704.

[0131] In a specific implementation process, at least one processor 701 executes the computer execution instructions stored in the memory 702, so that at least one processor 701 executes the above-mentioned embodiment of the processing method of the task.

[0132] For the specific implementation process of the processor 701, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0133] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.

[0134] The memory may include high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.

[0135] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0136] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in the method embodiment for processing any one of the above tasks when running.

[0137] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.

[0138] The embodiments of the present application also provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the method embodiment for processing any one of the above tasks.

[0139] Embodiments of the present application further provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in the embodiments of the processing method for any of the above tasks.

[0140] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0141] The above has introduced in detail a processing method, device, intelligent agent, and storage medium for a task provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A task processing method, characterized in that: include: Get the execution instructions of the tasks to be processed; Based on the pre-trained language model, the execution instruction is semantically decomposed to obtain the execution action sequence of the task to be processed; Based on the knowledge graph, modifying the execution actions in the execution action sequence that are different from those pre-stored in the knowledge graph to obtain a modified execution action sequence; The to-be-processed task is processed according to the modified execution action sequence.

2. The method according to claim 1, characterized in that The step of modifying the execution action sequence that is different from the execution action pre-stored in the knowledge graph based on the knowledge graph to obtain a modified execution action sequence includes: Compare each execution action in the execution action sequence with the execution actions pre-stored in the knowledge graph, and determine the execution action to be corrected from the execution action sequence; According to the execution action to be modified, determining a pre-stored target execution action corresponding to the execution action to be modified in the knowledge graph; According to the target execution action, the execution action to be corrected is corrected to obtain a corrected execution action sequence.

3. The method according to claim 2, characterized in that The step of modifying the to-be-modified execution action according to the target execution action to obtain a modified execution action sequence includes: Correcting the execution action to be corrected to the target execution action; Determining a target position of the target execution action according to the position relationship of the execution action to be corrected in the execution action sequence; The target execution action is updated at the target position to obtain a revised execution action sequence.

4. The method according to claim 2, characterized in that: Also includes: If, based on the execution action to be corrected, it is impossible to determine the pre-stored target execution action corresponding to the execution action to be corrected in the knowledge graph, a prompt message is generated, and the prompt message is used to request to regenerate a new execution action sequence for the task to be processed through the pre-trained language model.

5. The method according to claim 4, characterized in that After the prompt information is generated, the method further includes: Regenerate a new execution action sequence of the task to be processed according to the prompt information through the pre-trained language model; Compare each execution action in the new execution action sequence with the execution actions pre-stored in the knowledge graph, and determine a new execution action to be corrected from the new execution action sequence; If, according to the new execution action to be revised, it is still impossible to determine the new target execution action pre-stored in the knowledge graph corresponding to the new execution action to be revised, then generate request information; In response to an input operation performed by a user according to the request information, obtaining the new target execution action corresponding to the task to be processed; According to the new target execution action, identifying the target object instance information and the target object attribute information corresponding to the new target execution action; Update the knowledge graph according to the target object instance information and the target object attribute information.

6. The method according to claim 5, characterized in that The updating of the knowledge graph according to the target object instance information and the target object attribute information includes: The target object instance information is updated into the state graph of the knowledge graph, and the target object attribute information is updated into the attribute graph of the knowledge graph.

7. The method according to claim 6, characterized in that The updating of the target object instance information into the state graph of the knowledge graph, and the updating of the target object attribute information into the attribute graph of the knowledge graph, comprises: In a resource description framework format, based on a preset text format, updating the target object instance information into a state graph of the knowledge graph; In the resource description framework format, based on the preset text format, the target object attribute information is updated into the attribute graph of the knowledge graph.

8. A task processing device, characterized in that: include: An acquisition module is used to obtain execution instructions of tasks to be processed; A processing module, used for performing semantic decomposition on the execution instruction based on a pre-trained language model to obtain an execution action sequence of the task to be processed; A correction module, used for correcting the execution actions in the execution action sequence that are different from the execution actions pre-stored in the knowledge graph based on the knowledge graph to obtain a corrected execution action sequence; The processing module is further used to process the to-be-processed task according to the modified execution action sequence.

9. An intelligent agent, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for processing a task as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for processing the task as claimed in any one of claims 1 to 7.