Task processing method driven by intelligent agent and related equipment
The task processing method driven by the agent solves the poor performance of the video IoT data processing solution by integrating user problems, visual information and video knowledge bases, and uses tool agents trained by large language models, and achieves efficient and intelligent video understanding and task processing.
Patent Information
- Application Number
- CN202510196974.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video IoT data processing solutions have poor performance, domain-specific perception models are fragmented and rely on professional knowledge, and the general vision models lack understanding of specific domain knowledge, resulting in poor processing of fine-grained biometric recognition.
Adopt the task processing method driven by the agent, and integrate user problems, visual information, tool collections and video knowledge bases in multi-step inference, dynamic decision-making actions are made to answer user questions, and use the large language model (LLM) to train the tool agent to build a tool scheduling database with knowledge of the video Internet of Things field.
It improves the accuracy and efficiency of video IoT task processing, enhances the adaptability and flexibility of the system, reduces invalid calls and wrong decisions, and improves user experience and system performance.
Smart Images

Figure CN120297401A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to an agent-driven task processing method and related devices. Background Art
[0002] With the rapid development of the Internet of Things technology, Video Internet of Things (VIoT), as an important branch thereof, has realized real-time monitoring and data collection of the physical world by deploying a large number of visual sensors on the cloud, edge, and terminal sides. These visual sensors can collect a vast amount of video data, providing a rich information basis for applications such as environmental monitoring, target recognition, and behavior analysis. However, how to efficiently and intelligently analyze and process these video data to meet the complex requirements in different scenarios remains a technical challenge to be solved urgently.
[0003] In related technologies, the analysis and processing of Video Internet of Things mainly rely on domain-specific perception models, general vision models, etc. Domain-specific perception models, such as face recognition, gait recognition, vehicle re-identification, etc., although having high accuracy in their respective fields, these models are usually fragmented and need to be called one by one manually according to different video analysis requirements, which requires high professional knowledge of users. General vision models aim to obtain a unified visual representation to achieve multimodal video understanding, but such models may lack in-depth understanding of domain-specific knowledge and perform less satisfactorily especially when dealing with tasks such as fine-grained biometric recognition.
[0004] That is to say, the current video Internet of Things data processing solutions still have the problem of poor performance. Summary of the Invention
[0005] In view of this, the purpose of one or more embodiments of the present disclosure is to propose an agent-driven task processing method to solve the problems raised in the background art.
[0006] Based on the above purpose, one or more embodiments of the present disclosure provide an agent-driven task processing method, including:
[0007] Obtain a user question and visual information, where the visual information includes a video clip or a video image;
[0008] According to the user question and a preset agent framework, construct a user question prompt, where the user question prompt represents an input data structure after integrating the user question, the visual information, a tool set, and a video knowledge base, and the tool set includes a plurality of dedicated visual perception models;
[0009] Input the user question prompt into the tool agent to obtain an answer to the user question; wherein, in multi-step reasoning, the tool agent takes the user question prompt and the context prompt as inputs, dynamically makes decisions on actions to obtain an action result, the context prompt includes all actions and action results in the historical reasoning steps, and the action result indicates the answer to the user question;
[0010] Send the answer to the user question to the user.
[0011] Optionally, the calculation process of the tool agent in multi-step reasoning includes:
[0012] In response to determining that the calculation round is not the first time, the tool agent obtains the user question prompt and the context prompt;
[0013] Input the user question prompt and the context prompt into the tool agent, and dynamically determine whether to call a tool in the tool set in this round through the tool agent;
[0014] In response to determining that a tool in the tool set needs to be called, determine the target tool to be called through the tool agent;
[0015] According to the selection of the target tool, dynamically determine the input content of the target tool through the tool agent, and the input content includes videos and / or video images in the video knowledge base;
[0016] Input the input content into the target tool to obtain the action result of this round;
[0017] Update the context prompt template according to the action and action result of this round, and the action includes the selection of whether to call a tool in the tool set, the called tool, and the input content of the tool.
[0018] Optionally, the calculation process of the tool agent in multi-step reasoning further includes:
[0019] In response to determining that a tool in the tool set does not need to be called, generate an answer to the user question according to the action result of the historical iteration round;
[0020] Update the context prompt template according to the action and action result of this round and end the calculation, where the action includes the selection of whether to call a tool in the tool set and the answer to the user question, and the action result represents the answer to the user question.
[0021] Optionally, the training steps of the tool agent include:
[0022] Obtain multiple sets of training data, where each set of training data includes a training user question, training visual information, and an artificially annotated tool call instruction with a corresponding relationship;
[0023] Convert the training user question and the training visual information into a natural language instruction format, and construct an instruction fine-tuning sample set with the artificially annotated tool call instruction;
[0024] According to the training data and the instruction fine-tuning sample set, iteratively optimize the parameters of a preset large language model by minimizing the cross-entropy loss between the model output and the annotated tool instruction to obtain the tool agent.
[0025] Optionally, the training steps of the tool agent include:
[0026] Obtain multiple sets of training data, where each set of training data includes a training user question, training visual information, tool prompt words, and an artificially annotated tool call instruction with a corresponding relationship, and the tool prompt words include one or more of video knowledge base metadata, description documents of a tool set, and tool call examples;
[0027] Obtain an enhanced input according to the tool prompt words and the training data;
[0028] According to the enhanced input and the artificially annotated tool call instruction, iteratively optimize the parameters of a preset large language model by minimizing the cross-entropy loss between the model output and the annotated tool instruction to obtain the tool agent.
[0029] Optionally, determining the target tool to be called by the tool agent includes:
[0030] Fuse the user question prompt and the context prompt to generate a decision basis;
[0031] Determine candidate tool types according to the decision basis and the tool set, and the candidate tool types include one or more of a human-centered perception algorithm, a vehicle-centered perception algorithm, and an event-related perception algorithm;
[0032] Rank the candidate tools corresponding to the candidate tool types according to the decision basis, the candidate tool types, and the tool agent, and select the candidate tool with the highest priority as the target tool.
[0033] Optionally, it further includes:
[0034] Generate a visual feature representation according to the visual information and a preset visual encoder;
[0035] Construct a user question prompt according to the visual feature representation, the user question, and the preset agent framework, where the user question prompt represents an input data structure after integrating the user question, the visual feature representation, a set of tools, and a video knowledge base.
[0036] Based on the same inventive concept, one or more embodiments of the present disclosure further provide an agent-driven task processing device, including:
[0037] An acquisition module, configured to acquire a user question and visual information, where the visual information includes a video segment or a video image;
[0038] A first calculation module, configured to construct a user question prompt according to the user question and a preset agent framework, where the user question prompt represents an input data structure after integrating the user question, the visual information, a set of tools, and a video knowledge base, and the set of tools includes multiple dedicated visual perception models;
[0039] A second calculation module, configured to input the user question prompt into a tool agent to obtain an answer to the user question; wherein, in multi-step reasoning, the tool agent uses the user question prompt and a context prompt as inputs, dynamically makes decisions on actions to obtain an action result, the context prompt includes all actions and action results in historical reasoning steps, and the action result indicates the answer to the user question;
[0040] A sending module, configured to send the answer to the user question to the user.
[0041] Based on the same inventive concept, one or more embodiments of the present disclosure further provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the agent-driven task processing method as described in any one of the above.
[0042] Based on the same inventive concept, one or more embodiments of the present disclosure further provide a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the agent-driven task processing method as described in any one of the above.
[0043] As can be seen from the above, the technical solution provided by one or more embodiments of the present disclosure constructs a user question prompt by obtaining the user question and visual information and according to the user question and the preset agent framework. This prompt integrates the user question, visual information, tool set (including multiple dedicated visual perception models), and video knowledge base to form an input data structure. After inputting this prompt into the tool agent, during the multi-step reasoning process, the tool agent takes the user question prompt and context prompt as inputs, dynamically makes decisions on actions to obtain action results, finally gets the answer to the user question, and sends it to the user.
[0044] The technical solution of the present disclosure has the following advantages:
[0045] 1) Efficiently integrate information, integrate the user question, visual information, tool set, and video knowledge base into a unified input data structure, provide comprehensive information support for the tool agent, and improve the accuracy and efficiency of problem processing;
[0046] 2) The tool agent dynamically makes decisions on actions during the multi-step reasoning process, continuously adjusts the strategy according to the actions and action results in the historical reasoning steps, effectively processes complex problems, and improves the adaptability and flexibility of the system;
[0047] 3) Improve the success rate of tool calls. Integrating visual information and video knowledge base enables the tool agent to more accurately understand the user question, thereby improving the success rate of tool calls, reducing ineffective calls and incorrect decisions, and enhancing the system performance;
[0048] 4) Enhance the user experience, quickly and accurately answer the user question and timely feedback the result, enhance the user experience, enable the user to more conveniently obtain the required information, and enhance the practicality of the system and user satisfaction.
[0049] An agent-driven task processing device, an electronic device, and a computer-readable storage medium provided by the present disclosure can all implement the steps of the above-mentioned agent-driven task processing method, and thus also have the beneficial effects of the above-mentioned agent-driven task processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only one or more embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a flowchart of the agent-driven task processing method for one or more embodiments of the present disclosure;
[0052] Figure 2 Structural schematic diagram of an agent-driven task processing device according to one or more embodiments of the present disclosure;
[0053] Figure 3 Schematic diagram of the multi-step reasoning process of a tool agent according to one or more embodiments of the present disclosure;
[0054] Figure 4 Process schematic diagram of an agent-driven task processing method according to one or more embodiments of the present disclosure;
[0055] Figure 5 Scenario schematic diagram of an agent-driven task processing method according to one or more embodiments of the present disclosure;
[0056] Figure 6 Hardware structural schematic diagram of an electronic device according to one or more embodiments of the present disclosure. Detailed implementation manners
[0057] To make the objectives, technical solutions, and advantages of the present disclosure more clear and understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.
[0058] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in one or more embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0059] As described in the background art section, the widespread visual sensors, the communication technologies developed by Feishu, and the large-capacity network have made the wide application of VIoT possible. The interconnected network of a large number of visual sensors in the video Internet of Things has collected an unprecedented amount of video data on the cloud, edge, and terminal sides; the perception technology analyzes and processes the perceived and transmitted video data, providing comprehensive environmental monitoring and realizing the intelligence of monitoring. Among them, the perception technology has diversity and extensiveness in terms of analysis targets, understanding granularity, and practical applications, including from biometric recognition, human performance analysis to general scene understanding, etc. For example, the face recognition algorithm identifies the identity of a given image from a candidate image or video library based on facial features; the gait recognition algorithm identifies the identity by analyzing people's walking postures and comparing them with candidate videos; the person re-identification algorithm involves identifying pedestrians in different cameras or scenes; the vehicle re-identification identifies vehicles in different cameras or scenes based on appearance information; the license plate recognition identifies vehicles based on finer-grained license plate information; the crowd counting calculates the number of people from candidate images or video frames of a dense crowd; the fire and smoke detection is an algorithm for identifying and detecting dangerous fires and smoke in videos to eliminate fire hazards in a timely manner; the video violence detection extracts frame features and uses them to identify violent behaviors, etc.
[0060] In the related art, there are mainly two types of VIoT task processing solutions, including the video Internet of Things task processing solution based on a specific domain perception model and the video Internet of Things task processing solution based on a general vision model.
[0061] However, specific domain perception models, such as face recognition, gait recognition, vehicle re-identification, etc., although having high accuracy in their respective fields, these models are usually fragmented and need to be called one by one manually according to different video analysis requirements, which requires relatively high professional knowledge of users. The general vision model aims to obtain a unified visual representation to achieve multimodal video understanding, but such models may lack in-depth understanding of specific domain knowledge, especially when dealing with tasks such as fine-grained biometric recognition, and the performance is not ideal.
[0062] Therefore, the applicant proposes an agent-driven task processing method and related devices to solve the above technical problems.
[0063] The applicant found in the process of implementing the present disclosure that the large language model (LLM), as an agent, has the potential ability to use tools. If this ability can be utilized to construct and call various lightweight specific domain visual models to analyze various videos (or video images) collected by the video Internet of Things, the task processing performance of the video Internet of Things can be greatly improved.
[0064] Considering the characteristics of the LLM, this disclosure mainly puts forward requirements for tool agents from two aspects: 1) having the ability to distinguish fine-grained visual tools. For example, when identifying a person in a video, it is necessary to judge whether to call a face recognition algorithm, a person re-identification algorithm, or a gait recognition algorithm. When identifying a vehicle in a video, it is necessary to judge whether to call a vehicle re-identification algorithm or a license plate recognition algorithm; 2) being able to judge whether to perform multi-step reasoning according to the user's question and having the ability to decide which visual algorithm to select in the next round of reasoning.
[0065] Compared with fragmented perception algorithms, the technical solution of this disclosure can more intelligently understand video understanding requirements, uniformly, efficiently, and intelligently schedule and use existing fragmented perception algorithms, realize fine-grained and multi-step fragmented perception algorithm calls, and complete complex video understanding requirements.
[0066] In addition, the technical solution of this disclosure also needs to build a tool scheduling database that conforms to the knowledge in the field of video Internet of Things to realize the core of a learnable tool agent based on a local LLM, without relying on a commercial open-source model that only provides commercial interfaces, and build a localized tool agent system with independent intellectual property rights.
[0067] Reference Figure 1 , the agent-driven task processing method according to one or more embodiments of this disclosure includes the following steps:
[0068] Step S101: Obtain the user's question and visual information, where the visual information includes a video segment or a video image;
[0069] Step S102: According to the user's question and a preset agent framework, construct a user question prompt, where the user question prompt represents an input data structure after integrating the user's question, the visual information, a tool set, and a video knowledge base, and the tool set includes multiple dedicated visual perception models;
[0070] Step S103: Input the user question prompt into the tool agent to obtain an answer to the user's question; where, in multi-step reasoning, the tool agent takes the user question prompt and a context prompt as inputs, dynamically makes decisions on actions to obtain an action result, the context prompt includes all actions and action results in the historical reasoning steps, and the action result indicates the answer to the user's question;
[0071] Step S104: Send the answer to the user's question to the user.
[0072] In an embodiment of this disclosure, the visual information in step S101 can be obtained from the user or from the video knowledge base.
[0073] User questions can be input through a Graphical User Interface (GUI), and visual information can be input or selected.
[0074] The implementation of this disclosure is based on a preset agent framework, which includes a video knowledge base, a tool set, and a tool agent.
[0075] Among them, the video knowledge base is used to store video images or video clips collected by various intelligent terminal devices in VIoT. For example, the road condition data collected by road monitoring cameras can be represented as K = {K1, K2, …, K n}. The video knowledge base is the data basis of the agent framework of this disclosure, providing spatio-temporal information, environmental information, etc. for the task processing of VIoT.
[0076] The tool set includes various visual perception algorithms, and this visual perception algorithm can be a specific visual perception algorithm. In the embodiments of this disclosure, the visual perception algorithms can be classified into three types: human-centered visual perception algorithms, vehicle-centered visual perception algorithms, and event-related perception algorithms. Under these three types of perception algorithms, there are various specific visual perception algorithms. For example, human-centered visual perception algorithms can include face recognition algorithms, pedestrian re-identification algorithms, gait recognition algorithms, and crowd counting algorithms. Vehicle-centered visual perception algorithms can include vehicle re-identification algorithms and license plate recognition algorithms. Event-related perception algorithms can include smoke and fire detection algorithms, scene recognition algorithms, anomaly recognition algorithms, pose estimation algorithms, and action recognition algorithms. The tool set can be represented as T = {T1, T2, …, T m}.
[0077] In subsequent steps, when using the tool agent to determine the target tool to be called, first, the above user question prompt and the above context prompt are fused to generate a decision basis; then, according to the above decision basis and the above tool set, the candidate tool types are determined. The above candidate tool types include one or more of human-centered perception algorithms, vehicle-centered perception algorithms, and event-related perception algorithms; finally, according to the above decision basis, the above candidate tool types, and the above tool agent, the candidate tools corresponding to the above candidate tool types are sorted by priority, and the candidate tool with the highest priority is selected as the target tool.
[0078] The tool agent can be trained based on a Large Language Model (LLM). In the implementation of this disclosure, the tool agent is used to summarize the user questions and other information input by the user into a predefined template, and with the assistance of the knowledge base and the tool set, it conducts planning, observation, and reasoning, and finally provides integrated processing information for the user in the reply.
[0079] Specifically, as described in step S102, after obtaining the user's question and the corresponding visual information, a user question prompt is constructed based on a preset video knowledge base and a set of tools, and then the user question prompt is input into the tool agent to obtain the final result.
[0080] It can be understood that the tool agent is the core of the technical solution of the present disclosure. This tool agent not only needs to interact with users, tools, video data, and the environment, but also needs to consider the output of historical tool calls according to the context and observe the actions in multi-step reasoning.
[0081] Specifically, step S103 can be expressed as: Given a human query q i with potential visual information v i , use a predefined template to summarize and format the visual information v i , the human query q i , and the overall framework information (video knowledge base and tool set ) into a new prompt According to the input prompt p i , the tool agent determines the action A i,t in each step, records the observation result of the tool o i,t , and updates the context C i,t = (A i,1 , o i,1 , …, A i,t , o i,t ). Through repeated reasoning, until the final answer f i is obtained.
[0082] Among them, the action A i,t includes whether to select a tool, the name of the selected tool, and the input information of the selected tool. That is, the work of the tool agent in the t-th round can be expressed as:
[0083]
[0084] n i,t represents the decision on whether to use the tool , a i,t represents the name of the selected tool, k i,t represents the input information of the selected tool a i,t (usually queried from ), f i represents the final feedback of the tool agent, which can be the output content presented to the user.
[0085] In other words, in the embodiments of the present disclosure, in the first round of reasoning, the user question prompt is used as the input to obtain the action A of the first round of reasoning. i,1 This action indicates selecting a tool and indicates the name of the selected tool and the input of the tool; or indicates not selecting a tool and indicates the final feedback.
[0086] In the second round and each subsequent round of reasoning, not only the user question prompt is used as the input, but also the context prompt is synchronously used as the input of the tool agent. This context prompt is generated based on the actions and action results in the historical reasoning steps. That is, in the t-th round of reasoning, A i,t is determined by the tool agent according to the prompt p i and the context C of the previous step i,t-1 =(A i,1 , o i,1 , …, A i,t-1 , o i,t-1, ). This process can be expressed as P θ (A i,t ) = P θ (A i,t |p i , C i,t-1 ).
[0087] As Figure 3 shown, in one embodiment of the present disclosure, in the first round of reasoning, the user question prompt is used as the input to obtain the response of the first round of reasoning, and this response includes actions and action records. Among them, the user prompt includes a prompt word prefix (system information), conversation history, user question ("In the uploaded video, is there an abnormal event?"), and visual information (pictures or videos obtained from the user input). The actions obtained in the first round of reasoning include thinking (need to use tools), action (select the target visual model of "identifying the scene in the video"), action input (the input video or image of the target visual model), and observation (the output of the target visual model).
[0088] In the second round of reasoning, the user question prompt and the context prompt (the actions and action records of the first round of reasoning) are jointly used as the input of the tool agent to obtain the response of the second round of reasoning. The actions obtained in the second round of reasoning include thinking (need to use tools), action (select the target visual model of "detecting abnormal events based on the video scene"), action input (the input video or image of the target visual model), and observation (the output of the target visual model).
[0089] In the third round of reasoning, the user question prompt and context prompts (actions and action records from the first and second rounds of reasoning) are jointly used as the input to the tool agent to obtain the response for the third round of reasoning. The actions obtained from the third round of reasoning include thinking (no tool is required) and final answer (final feedback that can be shown to the user).
[0090] The language used during the reasoning process can be various languages such as Chinese, English, etc. according to the actual situation. This disclosure does not limit the specific language selection.
[0091] In the embodiments of this disclosure, the above tool agent can be fine-tuned and trained according to the ReAct instruction, or can be trained according to the prompt words including the video knowledge base, tool set description, and tool usage examples.
[0092] Specifically, in the first training scheme, the training steps may include: obtaining multiple sets of training data, each set of training data including a training user question, training visual information, and an artificially annotated tool call instruction with a corresponding relationship; converting the above training user question and the above training visual information into a natural language instruction format, and constructing an instruction fine-tuning sample set with the above artificially annotated tool call instruction; according to the above training data and the above instruction fine-tuning sample set, by minimizing the cross-entropy loss between the model output and the annotated tool instruction, iteratively optimizing the parameters of the preset large language model to obtain the above tool agent.
[0093] The above training process can be expressed as: L sft =-∑ i ∑ t log P θ (A i,t |p i ,C i,t-1 ). Through supervised instruction fine-tuning, the tool agent can call tools and query the necessary knowledge of VIoT to execute targeted instructions, especially in fine-grained tool usage and multi-step reasoning of related tools.
[0094] In the second training scheme, the training steps may include: obtaining multiple sets of training data, each set of training data including a training user question, training visual information, tool prompt words, and an artificially annotated tool call instruction, and the above tool prompt words include one or more of video knowledge base metadata, description documents of the tool set, and tool call examples; obtaining enhanced input according to the above tool prompt words and the training data; according to the above enhanced input and the artificially annotated tool call instruction, by minimizing the cross-entropy loss between the model output and the annotated tool instruction, iteratively optimizing the parameters of the preset large language model to obtain the above tool agent.
[0095] In the implementation of the present disclosure, before constructing a user question prompt based on the user's question and visual information, a visual encoder can also be used to initially understand the above visual information or the data in the video knowledge base to obtain a video feature vector, so as to retrieve the optimal tool from the tool set. It should be noted that after obtaining the video feature vector, the user question prompt can be obtained according to the user's question, the video feature vector of the visual information, the tool set, and the video knowledge base. The data in the video knowledge base can also be represented by the video feature vector.
[0096] In the implementation of the present disclosure, in order to improve the accuracy of tool retrieval results, breadth-first search or depth-first search can also be combined.
[0097] Hereinafter, taking one embodiment of the present disclosure as an example, a further description will be given.
[0098] As Figure 4 shown, the intelligent agent framework of the present disclosure includes a video knowledge base, a tool set, and a tool agent. Among them, the tool agent is trained based on a large language model and ReAct instruction data. During the application process, the tool agent gives a preset video data knowledge base and a perception algorithm tool set according to the visual information (an image of a bus) input by the user and the user's question ("Can you confirm whether the vehicle in the photo has been sighted in Beijing", which can be input in various languages such as Chinese and English), and obtains the answer to the user's question through multi-step reasoning. The answer content includes text ("We found this vehicle", which can be input in various languages such as Chinese and English) and two video images containing buses.
[0099] The above video data knowledge base includes videos collected by end devices, and these videos can be stored in the cloud storage space through data transmission. The above perception algorithm tool set includes a human-centered perception algorithm, a vehicle-centered perception algorithm, and an event-related perception algorithm.
[0100] In the above embodiment, only 11 tools are introduced for experiments, so a direct tool search function is used. If more tools are introduced, the perception algorithm hierarchical classification method can be combined to first retrieve the algorithm category according to the instruction and context information, and then retrieve the algorithm subclass to improve the tool call success rate. The present disclosure does not limit the specific retrieval scheme.
[0101] To verify the technical effects of the technical solution of the present disclosure, the applicant has conducted experiments in various scenarios. As Figure 5 shown, the technical solution of the present disclosure can be used in various application scenarios such as face recognition, person re-identification, gait recognition, crowd counting, license plate recognition, vehicle re-identification, smoke detection, anomaly detection, and behavior analysis.
[0102] It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0103] It should be noted that the method of one or more embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of one or more embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0104] It should be noted that the above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an agent-driven task processing device. As Figure 2 shown, it includes:
[0106] An acquisition module 11, configured to acquire a user question and visual information, where the visual information includes video segments or video images;
[0107] A first calculation module 12, configured to construct a user question prompt according to the user question and a preset agent framework, where the user question prompt represents an input data structure after integrating the user question, the visual information, a tool set, and a video knowledge base, and the tool set includes multiple dedicated visual perception models;
[0108] A second calculation module 13, configured to input the user question prompt into a tool agent to obtain an answer to the user question; where, in multi-step reasoning, the tool agent uses the user question prompt and a context prompt as inputs, dynamically makes decisions on actions to obtain an action result, the context prompt includes all actions and action results in the historical reasoning steps, and the action result indicates the answer to the user question;
[0109] A sending module 14, configured to send the answer to the user question to the user.
[0110] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0111] The device in the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0112] Figure 6 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0113] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure.
[0114] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present disclosure through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0115] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0116] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0117] The bus 1050 includes a path for transmitting information among various components of the device, such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040.
[0118] It should be noted that although only the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050 are shown in the above device, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0119] The electronic device of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.
[0120] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0121] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0122] In addition, for simplicity of explanation and discussion, and so as not to make one or more embodiments of the present disclosure difficult to understand, well-known power / ground connections of integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making one or more embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details of the implementation of such block diagram devices are highly dependent on the platform on which one or more embodiments of the present disclosure are to be implemented (i.e., these details should be entirely within the understanding of those of ordinary skill in the art). In cases where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those of ordinary skill in the art that one or more embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.
[0123] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0124] One or more embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. An agent-driven task processing method, characterized in that, including: obtaining a user question and visual information, where the visual information includes a video clip or a video image; constructing a user question prompt according to the user question and a preset agent framework, where the user question prompt represents an input data structure after integrating the user question, the visual information, a tool set, and a video knowledge base, and the tool set includes multiple dedicated visual perception models; inputting the user question prompt into a tool agent to obtain an answer to the user question; wherein, in multi-step reasoning, the tool agent takes the user question prompt and a context prompt as inputs, dynamically makes decisions on actions to obtain an action result, the context prompt includes all actions and action results in historical reasoning steps, and the action result indicates the answer to the user question; sending the answer to the user question to the user.
2. The method according to claim 1, wherein The calculation process of the tool agent in multi-step reasoning includes: in response to determining that the calculation round is not the first time, the tool agent obtains the user question prompt and the context prompt; inputting the user question prompt and the context prompt into the tool agent, and dynamically determining by the tool agent whether it is necessary to call a tool in the tool set in this round; in response to determining that it is necessary to call a tool in the tool set, determining the target tool to be called by the tool agent; dynamically determining the input content of the target tool according to the selection of the target tool, where the input content includes videos and / or video images in the video knowledge base; inputting the input content into the target tool to obtain the action result of this round; updating the context prompt template according to the action and action result of this round, where the action includes the selection of whether to call a tool in the tool set, the called tool, and the input content of the tool.
3. The method according to claim 2, wherein The calculation process of the tool agent in multi-step reasoning further includes: in response to determining that it is not necessary to call a tool in the tool set, generating an answer to the user question according to the action result of the historical iteration round; updating the context prompt template according to the action and action result of this round and ending the calculation, where the action includes the selection of whether to call a tool in the tool set and the answer to the user question, and the action result represents the answer to the user question.
4. The method according to claim 2, characterized in that, The training steps of the tool agent include: obtaining multiple groups of training data, where each group of training data includes a training user question, training visual information, and an artificially annotated tool call instruction with a corresponding relationship; converting the training user question and the training visual information into a natural language instruction format, and constructing an instruction fine-tuning sample set with the artificially annotated tool call instruction; iteratively optimizing the parameters of a preset large language model according to the training data and the instruction fine-tuning sample set by minimizing the cross-entropy loss between the model output and the annotated tool instruction to obtain the tool agent.
5. The method according to claim 2, wherein The training steps of the tool agent include: Obtain multiple sets of training data, where each set of training data includes a training user question, training visual information, tool prompt words, and manually annotated tool call instructions with corresponding relationships. The tool prompt words include one or more of video knowledge base metadata, description documents of the tool set, and tool call examples; Obtain an enhanced input based on the tool prompt words and the training data; According to the enhanced input and the manually annotated tool call instructions, iteratively optimize the parameters of a preset large language model by minimizing the cross-entropy loss between the model output and the annotated tool instructions to obtain the tool agent; 6. The method according to claim 2, wherein The determining of the target tool to be called by the tool agent includes: Perform a fusion process on the user question prompt and the context prompt to generate a decision basis; Determine candidate tool types according to the decision basis and the tool set. The candidate tool types include one or more of human-centered perception algorithms, vehicle-centered perception algorithms, and event-related perception algorithms; According to the decision basis, the candidate tool types, and the tool agent, perform a priority ranking on the candidate tools corresponding to the candidate tool types, and select the candidate tool with the highest priority as the target tool.
7. The method according to claim 1, characterized in that It further includes: Generate a visual feature representation according to the visual information and a preset visual encoder; Construct a user question prompt according to the visual feature representation, the user question, and the preset agent framework. The user question prompt represents the input data structure after integrating the user question, the visual feature representation, the tool set, and the video knowledge base.
8. An agent-driven task processing device, characterized in that It includes: An acquisition module configured to acquire a user question and visual information, where the visual information includes a video segment or a video image; A first calculation module configured to construct a user question prompt according to the user question and a preset agent framework. The user question prompt represents the input data structure after integrating the user question, the visual information, the tool set, and the video knowledge base. The tool set includes multiple dedicated visual perception models; A second calculation module configured to input the user question prompt into the tool agent to obtain an answer to the user question. Among them, in multi-step reasoning, the tool agent uses the user question prompt and the context prompt as inputs, dynamically makes decisions on actions to obtain an action result, and the action result indicates the answer to the user question; A sending module configured to send the answer to the user question to the user.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.