Image processing method, electronic device, and computer-readable storage medium

CN120726653BActive Publication Date: 2026-09-08ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410383763.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2026-09-08
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种图像处理方法、电子设备以及计算机可读存储介质,以至少解决相关技术中对图表图像的处理准确度较低的技术问题

Benefits of technology

[0018] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726653B_ABST
    Figure CN120726653B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, an electronic device and a computer readable storage medium, and relates to the fields of large model technology, intelligent agents and graph processing. The method comprises the following steps: acquiring a graph image and text data, wherein the text data is used to represent a processing task to be performed on the graph image; dividing the processing task based on the text data to obtain a plurality of subtasks having a correlation relationship; and performing the plurality of subtasks on the graph image to obtain a processing result of the graph image. The application solves the technical problem of low processing accuracy of the graph image in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to large model technology, intelligent agent field, and image processing field, and more specifically, to an image processing method, electronic device, and computer-readable storage medium. Background Technology

[0002] Currently, some powerful multimodal large models have emerged, which have shown good performance in answering natural images and can perform end-to-end reasoning. However, due to the complexity of graph and image problems, which require models to perform visual perception and symbolic reasoning simultaneously, current models do not perform well when processing graph and image problems, especially when dealing with manually written problems in real-world scenarios.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides an image processing method, an electronic device, and a computer-readable storage medium to at least solve the technical problem of low accuracy in processing chart images in related technologies.

[0005] According to one aspect of the embodiments of this application, an image processing method is provided, comprising: acquiring a chart image and text data, wherein the text data is used to characterize a processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with related relationships; and performing the multiple sub-tasks on the chart image to obtain a processing result of the chart image.

[0006] According to one aspect of the embodiments of this application, an image processing method is also provided, comprising: acquiring a chart image and query information, wherein the query information is used to characterize a processing task to be performed on the chart image; dividing the processing task based on the query information to obtain multiple sub-tasks with related relationships; using execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain a processing result of the chart image; and generating response information corresponding to the query information based on the processing result.

[0007] According to one aspect of the embodiments of this application, an image processing method is also provided, comprising: responding to an input command applied to an operation interface, displaying a chart image and text data on the operation interface, wherein the text data is used to characterize a processing task to be performed on the chart image; and responding to a processing command applied to the operation interface, displaying a processing result of the chart image on the operation interface, wherein the processing result is obtained by using execution agents corresponding to the plurality of sub-tasks to perform a plurality of related sub-tasks on the chart image, and the plurality of sub-tasks are obtained by dividing the processing task based on the text data.

[0008] According to one aspect of the embodiments of this application, an image processing method is also provided, comprising: acquiring a chart image and text data by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter including the chart image and text data, and the text data being used to characterize a processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with an association relationship; using execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain a processing result of the chart image; and outputting the processing result by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter including the processing result.

[0009] According to one aspect of the embodiments of this application, an image processing platform is also provided, comprising: a planning node, configured to acquire a chart image and text data, and to divide the processing tasks to be performed on the chart image based on the text data to obtain multiple sub-tasks with related relationships, wherein the text data is used to characterize the processing tasks; and multiple execution nodes, wherein execution agents corresponding to the multiple sub-tasks are deployed on the multiple execution nodes, and the multiple execution nodes are configured to use the execution agents corresponding to the multiple sub-tasks to perform the sub-tasks corresponding to the execution agents on the chart image respectively, to obtain the processing results of the chart image.

[0010] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: an acquisition module for acquiring a chart image and text data, wherein the text data is used to characterize a processing task to be performed on the chart image; a division module for dividing the processing task based on the text data to obtain multiple sub-tasks with an association relationship; and an execution module for using execution agents corresponding to the multiple sub-tasks to execute the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image.

[0011] According to one aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: an acquisition module, configured to acquire a chart image and query information, wherein the query information is used to characterize a processing task to be performed on the chart image; a division module, configured to divide the processing task based on the query information to obtain multiple sub-tasks with related relationships; an execution module, configured to use execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain a processing result of the chart image; and a generation module, configured to generate response information corresponding to the query information based on the processing result.

[0012] According to one aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: a first display module, configured to respond to an input command applied to an operation interface and display a chart image and text data on the operation interface, wherein the text data is used to characterize a processing task to be performed on the chart image; and a second display module, configured to respond to a processing command applied to the operation interface and display a processing result of the chart image on the operation interface, wherein the processing result is obtained by using execution agents corresponding to multiple sub-tasks to perform multiple related sub-tasks on the chart image, and the multiple sub-tasks are obtained by dividing the processing task based on text data.

[0013] According to one aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: an acquisition module, configured to acquire a chart image and text data by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter including the chart image and text data, and the text data being used to characterize a processing task to be performed on the chart image; a division module, configured to divide the processing task based on the text data to obtain multiple sub-tasks with an association relationship; an execution module, configured to use execution agents corresponding to the multiple sub-tasks to execute the multiple sub-tasks on the chart image respectively to obtain a processing result of the chart image; and an output module, configured to output the processing result by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter including the processing result.

[0014] According to another aspect of the embodiments of this application, a computer terminal is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the methods in various embodiments of this application.

[0018] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0019] In this embodiment, a chart image and text data are acquired, with the text data representing the processing task to be performed on the chart image. The processing task is divided based on the text data, resulting in multiple related sub-tasks. The execution agents corresponding to these sub-tasks then perform the sub-tasks on the chart image, obtaining the processing result. This achieves the technical effect of improving the accuracy of chart image processing. It is noteworthy that the chart image processing task can be divided into multiple related sub-tasks, allowing different agents to be used for different sub-tasks. Compared to the limitations of using a single agent, this application combines the advantages of multiple agents to improve the processing accuracy of executing multiple sub-tasks on the chart image, resulting in a more accurate chart image processing result. This solves the technical problem of low accuracy in chart image processing in related technologies.

[0020] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 This is a schematic diagram illustrating an application scenario of an image processing method according to an embodiment of this application;

[0023] Figure 2 This is a flowchart of the image processing method according to Embodiment 1 of this application;

[0024] Figure 3 This is a schematic diagram of a smart agent cooperative optimization structure according to an embodiment of this application;

[0025] Figure 4 This is a flowchart of an image processing method according to Embodiment 2 of this application;

[0026] Figure 5 This is a flowchart of an image processing method according to Embodiment 3 of this application;

[0027] Figure 6 This is a flowchart of an image processing method according to Embodiment 4 of this application;

[0028] Figure 7 This is a schematic diagram of an image processing apparatus according to Embodiment 5 of this application;

[0029] Figure 8 This is a schematic diagram of an image processing apparatus according to Embodiment 6 of this application;

[0030] Figure 9 This is a schematic diagram of an image processing apparatus according to Embodiment 7 of this application;

[0031] Figure 10 This is a schematic diagram of an image processing apparatus according to Embodiment 8 of this application;

[0032] Figure 11 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] The technical solution provided in this application is mainly implemented using large-scale model technology. Here, "large-scale model" refers to a deep learning model with a massive number of parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of parameters. Large-scale models can also be called foundational models. They are pre-trained using large-scale unlabeled corpora to produce pre-trained models with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multimodal pre-training models.

[0036] It's important to note that in practical applications, large models can be fine-tuned using a small number of samples after pre-training, allowing them to be applied to various tasks. For example, large models can be widely used in Natural Language Processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios for large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0037] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0038] A multi-agent system is a computing system composed of multiple interacting agents. Multi-agent systems can solve some problems that are difficult to solve by a single multi-agent system or a single system.

[0039] Autonomous agents based on large models, or those based on large model languages, utilize large model languages ​​as their core cognitive components or controllers. They broaden their perception and action capabilities through methods such as multimodal perception and tool usage, thereby completing complex reasoning tasks.

[0040] The system automatically optimizes natural language prompts by leveraging the powerful text editing capabilities of large model languages, thereby achieving better task performance.

[0041] Graph symbol reasoning refers to question answering based on graph images. The model first needs to extract relevant information from the image, organize it in a reasonable way, and then perform mathematical operations or logical reasoning on the extracted items to obtain the final answer.

[0042] Currently, the main methods for chart and image reasoning are as follows:

[0043] Using an end-to-end multimodal large model, this method inputs the graph image and the question together into the multimodal large model for end-to-end reasoning, and uses the output of the large model as the final answer. This method is relatively simple and efficient. However, since the graph questions are quite complex, the model needs to perform visual perception and symbolic reasoning simultaneously, and the current multimodal large model does not perform well in this regard.

[0044] By calling external tools for graph and image reasoning, this method can transform multimodal graph and image problem-solving tasks into pure language table problem-solving tasks. Then, a large model can be used to reason based on the transformed table. This method fully combines the advantages of multiple models to complete the task. However, the performance of this method depends to a large extent on the quality of the prompts used to guide the model's behavior, and effective prompts require tedious artificial natural language prompting engineering.

[0045] The main problems faced by chart and image reasoning methods are as follows:

[0046] Although a number of powerful multimodal large models have emerged, which have shown good performance in answering natural images, this end-to-end reasoning approach still performs poorly in graph and image reasoning, especially when faced with manually written questions in real-world scenarios, where it often fails to provide satisfactory answers.

[0047] To address some of the problems existing in the above-mentioned chart and image reasoning, this application proposes a multi-agent prompting collaborative optimization method for chart and image reasoning. This method can automatically collect multi-faceted error feedback from the multi-agent system during execution to make reliable prompting improvements, and effectively explore the vast space of multi-agent prompts by utilizing a collaborative reward mechanism. This frees humans from tedious natural language prompting engineering. The automatically optimized prompts can help the multi-agent system to better cooperate between models and greatly improve the chart and image reasoning ability.

[0048] Example 1

[0049] According to an embodiment of this application, an image processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0050] Considering the large number of model parameters in large models and the limited computing resources of mobile terminals, the image processing method provided in this application embodiment can be applied to, for example, Figure 1 The application scenarios shown are not limited to these. Figure 1 This is a schematic diagram illustrating an application scenario of an image processing method according to an embodiment of this application. Figure 1In the application scenario shown, the large model is deployed on server 10. Server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), internet connection, or other types of data network. These client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to access the large model, thereby implementing the method provided in this embodiment.

[0051] In this embodiment, the system consisting of a client device and a server can perform the following steps: The client device performs the step of generating chart images and text data. The server performs the step of acquiring chart images and text data; based on the text data, the processing task is divided into multiple sub-tasks with related relationships; the execution agents corresponding to the multiple sub-tasks are used to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image. It should be noted that, provided that the operating resources of the client device can meet the deployment and operation conditions of the large model, this embodiment can be performed on the client device.

[0052] Under the aforementioned operating environment, this application provides the following: Figure 2 The image processing method shown. Figure 2 This is a flowchart of the image processing method according to Embodiment 1 of this application. Figure 2 As shown, the method may include the following steps:

[0053] Step S202: Obtain chart image and text data, wherein the text data is used to characterize the processing task to be performed on the chart image;

[0054] The aforementioned charts and graphs refer to images that use graphics, tables, or other formats to express data and information. For example, they can take the form of bar charts, line charts, pie charts, etc., used to visually display data trends and compare relationships between different data points. Charts and graphs help people understand data more quickly, discover patterns and trends, and thus make more accurate decisions. They have wide applications in business, scientific research, education, and other fields.

[0055] The aforementioned text data can be a pre-set processing task to be performed on the chart image. Alternatively, the text data can be a processing task generated according to project requirements. The specific method of generating the text data is not limited here; it can be determined based on actual needs. For example, the text data can be used to analyze visual elements in the chart image to obtain the data represented by those elements; it can also be used to extract numerical or text elements from the chart image; and it can also be other processing tasks that can be performed on the chart image.

[0056] In one optional embodiment, a chart image can be acquired first, and the processing tasks to be performed on the chart image can be automatically determined by the model, thereby generating text data based on the processing tasks. Alternatively, the chart image can be displayed on the user's client so that after viewing the chart image, the user can determine the processing tasks to be performed on the chart image according to project requirements and generate text data based on the processing tasks. Alternatively, multiple processing tasks can be pre-set, and text data can be generated by selecting the processing task to be performed on the chart image from multiple processing tasks.

[0057] Step S204: Divide the processing task based on the text data to obtain multiple sub-tasks with related relationships;

[0058] The aforementioned related subtasks can refer to subtasks whose processing procedures are related; for example, the input data of a later subtask needs to be the output data of a previous subtask. Alternatively, these related subtasks can refer to subtasks that require interaction with each other.

[0059] In one optional embodiment, the processing task can be decomposed based on text data to obtain multiple sub-tasks corresponding to different models. These sub-tasks can be interconnected, which facilitates the collaborative processing of multiple sub-tasks by multiple different models, improves the processing efficiency of multiple sub-tasks, and enables better analysis of chart data and reasoning about chart symbols in the chart data.

[0060] Step S206: The execution agents corresponding to the multiple subtasks perform multiple subtasks on the chart image respectively to obtain the processing result of the chart image.

[0061] The execution agents corresponding to the above multiple subtasks can be the execution agents corresponding to the multiple subtasks respectively, and any multiple subtasks can also use the same execution agent. Here, the relationship between multiple subtasks and execution agents is not specifically limited.

[0062] The aforementioned execution agent is mainly responsible for performing tasks. This execution agent is primarily responsible for performing multiple sub-tasks on the chart image to obtain the processing result of the chart image.

[0063] The intelligent agent mentioned in this application can refer to a subject with intelligence and autonomous decision-making capabilities. An intelligent agent can be a model, an actual robot, program, or system. It can perceive its environment, make decisions and take actions, and possess learning capabilities. Therefore, an intelligent agent can be more than just a model; it can also be a subject with intelligent behavior.

[0064] In one optional embodiment, multiple subtasks can be executed on the chart image based on a model of multiple subtasks. The multiple subtasks can be executed in the order of execution, so as to obtain the processing result of the chart image based on the collaborative execution of multiple subtasks.

[0065] If multiple subtasks are dependent on each other, they can be executed sequentially on the chart image according to these dependencies to obtain the processed chart image. For example, the first subtask could be identifying information about visual elements in the chart data. Based on this information, a second subtask could be executed, which could extract text or numerical information from the visual elements. Subsequent subtasks could then be executed based on this text or numerical information to obtain the processed chart image.

[0066] If there is no dependency between multiple subtasks, multiple subtasks can be executed in parallel on the chart image to obtain the processing result of the chart image; if some subtasks have a dependency between multiple subtasks, and other subtasks can be processed in parallel, then some subtasks can be executed according to the dependency of some subtasks, and other subtasks can be processed in parallel to obtain the processing result of the chart image.

[0067] Through the above steps, chart images and text data are obtained, whereby the text data is used to represent the processing tasks to be performed on the chart images. The processing tasks are divided based on the text data, resulting in multiple related sub-tasks. The execution agents corresponding to these sub-tasks then perform the sub-tasks on the chart images, obtaining the processing results. This achieves the technical effect of improving the accuracy of chart image processing. It is noteworthy that the chart image processing task can be divided into multiple related sub-tasks, allowing different agents to be used for different sub-tasks. Compared to the limitations of using a single agent to process tasks, this application can combine the advantages of multiple agents to improve the processing accuracy of multiple sub-tasks on the chart images, resulting in more accurate chart image processing results. This solves the technical problem of low accuracy in chart image processing in related technologies.

[0068] In the above embodiments of this application, the processing task is divided based on text data to obtain multiple sub-tasks with related relationships, including: inputting text data into a planning agent to divide the processing task using the planning agent and obtaining multiple sub-tasks output by the planning agent.

[0069] The aforementioned planning agent is primarily responsible for formulating a plan. This agent is mainly responsible for decomposing the processing tasks in the chart / image into sub-tasks. The planning agent can organize multiple execution agents corresponding to these sub-tasks to extract relevant information from the chart data and perform symbolic reasoning based on the context.

[0070] In one alternative embodiment, the planning agent can analyze and process the text data, divide the processing task into multiple sub-tasks, and output them. This allows the processing task to be more effectively allocated to different agents, thereby improving the execution efficiency and completion quality of the processing task.

[0071] In another alternative embodiment, text data can be input into a planning agent. The process of dividing the processing task by the planning agent can help determine the specific steps and execution order of the processing task. The planning agent can decompose the processing task into multiple smaller sub-tasks based on factors such as the complexity of the processing task, resource availability, and time constraints, and arrange the execution order and schedule of these sub-tasks.

[0072] By planning multiple sub-tasks output by the intelligent agent, a better understanding of the specific execution flow and timeline of the tasks can be achieved, enabling better organization and allocation of resources. This helps teams collaborate and execute tasks more effectively, improving the speed and quality of task completion.

[0073] The aforementioned planning agent plays a crucial role in the process of task division and execution, enabling better task planning and execution, and improving work efficiency and completion quality.

[0074] In the above embodiments of this application, multiple subtasks are executed on a chart image by execution agents corresponding to multiple subtasks to obtain the processing result of the chart image. This includes: determining the execution order of the execution agents corresponding to multiple subtasks based on the association relationship of multiple subtasks; and executing the corresponding subtask on the chart image by one of the execution agents corresponding to multiple subtasks in sequence according to the execution order to obtain the processing result.

[0075] The aforementioned executive agents can be action agents. These executive agents include, but are not limited to, data retrieval agents and program agents.

[0076] The relationships between the multiple subtasks mentioned above can be either interdependent or independent. If there are dependencies between multiple subtasks, their execution order will be affected and must be determined according to the dependencies. For example, if subtask A needs to be executed before subtask B, then there is a dependency between them. However, if the multiple subtasks are independent, their execution order can be arbitrary, and they can be executed in parallel.

[0077] When determining the execution order of agents corresponding to multiple subtasks, it is necessary to consider the relationships between the subtasks and the cooperation methods between the agents. Optionally, the execution order of the agents can be determined based on the dependencies between the subtasks to ensure that multiple subtasks can obtain the correct processing results.

[0078] When executing a chart image using one of the multiple subtask agents in sequence, it's also necessary to consider the relationships between the subtasks. After each subtask is executed, the processing result can be used as input for the next agent, thus enabling collaborative processing of multiple subtasks to achieve the aforementioned result.

[0079] In one optional embodiment, multiple sub-tasks to be performed can be identified first, such as data retrieval, image processing, and text generation. Then, based on the capabilities and characteristics required by the multiple sub-tasks, the execution agents to be used are determined, such as data retrieval agents and program agents. Based on the relationships between the multiple sub-tasks, the execution order of the execution agents corresponding to the multiple sub-tasks is determined: the execution order of the execution agents corresponding to the multiple sub-tasks is determined according to the dependencies and execution order between the multiple sub-tasks to ensure that the sub-tasks are processed by the execution agents in the correct order. Following the execution order, one execution agent corresponding to each of the multiple sub-tasks is used sequentially to perform the corresponding sub-task on the chart image to obtain the processing result. For example, first, the data retrieval agent is used to obtain the data to be processed, and then the program agent is used to process the image to finally obtain the processing result.

[0080] In the above embodiments of this application, the execution agents corresponding to the multiple sub-tasks include: a data retrieval agent and a program agent; according to the execution order, one of the execution agents corresponding to the multiple sub-tasks is used to perform the corresponding sub-tasks on the chart image in turn to obtain the processing result, including: using the data retrieval agent to extract chart data from the chart image; using the program agent to generate a task program based on the chart data, and executing the task program to obtain the processing result.

[0081] The aforementioned data retrieval agent is used to extract data attributes from chart images, that is, chart data in chart images.

[0082] The aforementioned program agent can utilize other agents to obtain context, so as to write and execute task programs based on the context information, thereby effectively solving the graph query problem.

[0083] The above-mentioned task program can be a data analysis program, which can be used to perform statistical analysis on chart data.

[0084] The results of the above processing can be presented as analysis reports, visualization charts, visualization reports, etc.

[0085] In one optional embodiment, the data retrieval agent can be used as a program to extract data attributes from a chart image. This data retrieval agent can analyze the chart image, identify the data within it, and then extract this data for subsequent processing and analysis. The extracted data attributes may include, but are not limited to, various data points, labels, values, and other information from the chart. The extracted data attributes can be converted into a data format that can be processed by a computer program to obtain the aforementioned chart data.

[0086] In another alternative embodiment, the program agent can generate a task program based on the chart data. The chart data can be input into the program agent to obtain the task program, which can then be run to obtain the processing result of the chart data.

[0087] In the above embodiments of this application, the execution agents corresponding to the multiple sub-tasks include: a data retrieval agent, a visual retrieval agent, and a program agent. Following the execution order, one of the execution agents corresponding to the multiple sub-tasks is used sequentially to perform the corresponding sub-tasks on the chart image to obtain processing results. This includes: using the data retrieval agent to extract chart data from the chart image; using the visual retrieval agent to extract visual information from the chart image based on text data; and using the program agent to generate a task program based on the chart data and visual information, and executing the task program to obtain processing results.

[0088] The aforementioned visual retrieval agent can retrieve visual knowledge that is important for solving complex problems.

[0089] In one optional embodiment, visual information can be incorporated when generating the task program to obtain a more accurate program. Chart images contain a large amount of image data, such as bar charts, lines, etc., which needs to be retrieved by a visual retrieval agent. Optionally, the visual retrieval agent can retrieve visual information from the chart image through a dialogue between a larger and a smaller model. For example, the larger model can generate a retrieval question such as "What is the label of the red bar in the chart image?", which can then be answered by a smaller model to obtain the visual information from the chart image.

[0090] The programmatic agent can be used to perform more accurate analysis and processing of chart images based on chart data and visual information, thereby obtaining a task program with a higher degree of matching with the chart image, thus improving the accuracy of the generated task program. Then, the task program is executed to obtain a more accurate processing result for the chart image.

[0091] In the above embodiments of this application, the method further includes: outputting a processing result; receiving a first feedback result corresponding to the processing result, wherein the first feedback result is used to characterize whether an error has occurred in the processing result; if the first feedback result characterizes that an error has occurred in the processing result, obtaining an intermediate execution result output by the agent set, wherein the agent set includes a planning agent and execution agents corresponding to multiple subtasks; and optimizing the prompt information input to the agent set based on the intermediate execution result.

[0092] The above-mentioned prompts can be the prompts that the intelligent agent set needs to process. The prompts of the intelligent agent set can be optimized by using intermediate execution results to improve the processing accuracy of the intelligent agent set.

[0093] In one optional embodiment, the processing result can be sent to the user's client so that the user can check whether there is an error in the processing result through the client. A first feedback message can be generated based on the user's judgment of the processing result. If the processing result is correct, the first feedback message can be used to indicate that there is no error in the processing result, and no subsequent adjustment steps need to be performed. If the processing result has an error, the first feedback message can be used to indicate that there is an error in the processing result, and the intelligent agent set needs to be optimized based on the intermediate execution results output by the intelligent agent set to obtain an intelligent agent set with higher accuracy.

[0094] By analyzing intermediate execution results, specific problematic steps can be identified, allowing for targeted corrections and optimizations. The agent set can be updated using intermediate execution results to adjust and improve the planning and execution agents, thereby enhancing the accuracy and efficiency of the overall processing. It's important to note that receiving and feeding back processing results, as well as updating the agent set based on intermediate execution results, is an iterative process. Through continuous analysis and updates, the overall performance of the agent set can be continuously improved, enabling it to better handle various complex tasks and scenarios.

[0095] In the above embodiments of this application, optimizing the agent set based on intermediate execution results includes: analyzing the intermediate execution results using an optimization model to generate an optimization strategy for the agent set; and optimizing the agent set based on the optimization strategy.

[0096] The above-mentioned prompts can be the prompts that the ensemble of agents needs to process. The prompts input to the ensemble of agents can be optimized through optimization strategies to improve the processing accuracy of the ensemble of agents.

[0097] The process of analyzing intermediate execution results using the optimization model described above can be achieved through a reliable prompting mechanism.

[0098] In an optional embodiment, to accurately pinpoint the specific cause of errors in the execution results of the multi-agent system, when an error occurs in the agent ensemble on a specific graph problem, the intermediate execution results of multiple agents in the agent ensemble at the error instance can be collected to form a multi-faceted error feedback, i.e., f tFor example, we can collect the execution strategies of the execution agents corresponding to multiple sub-tasks generated by the planning agent, the tabular data generated by the data retrieval agent, and the program code (Python code) generated by the program agent for specific graph problems. After collecting error feedback from various sources, we can leverage the self-reflection capabilities of the large language model to summarize the reasons for the incorrect answers generated by the current multi-agent system, thereby extracting insightful domain knowledge and generating fine-grained suggestions to improve the multi-agent prompts.

[0099] In the above embodiments of this application, the intermediate execution results are analyzed using an optimization model to generate an optimization strategy for the intelligent agent set, including: inputting the intermediate execution results into the optimization model and obtaining optimization suggestion information output by the optimization model; inputting the optimization suggestion information into the optimization model and obtaining the optimization strategy output by the optimization model.

[0100] In an alternative embodiment, intermediate execution results can be input into an optimization model to implement a reliable suggestion mechanism through the optimization suggestion information output by the optimization model. Optimization model O can be used to generate insightful suggestions, i.e., optimization suggestion information a. t , represented as:

[0101] a t =O(f t );

[0102] Among them, f t a is an intermediate execution result. t To optimize the suggested information.

[0103] The optimization model can be used to edit the prompts to the agent based on the generated optimization suggestions, thus obtaining the final optimization strategy, as shown in the formula below:

[0104] P t+1 =O(a) t ,P t );

[0105] Among them, a t To optimize the suggested information, P t Given a set of agents, the suggested information can be optimized. By optimizing the set of agents, P can be obtained. t+1 .

[0106] In one optional embodiment, intermediate execution results can be input into the optimization model to obtain optimization suggestions. These intermediate execution results can be various data from the production process, such as production efficiency, cost, and resource utilization. The optimization model analyzes this data and provides corresponding optimization suggestions, such as adjusting data retrieval methods, adjusting task program generation methods, and adjusting visual retrieval methods. Inputting optimization suggestions into the optimization model yields the optimization strategy output by the model. The optimization strategy refers to the specific action plan or improvement plan formulated based on the optimization suggestions. A reliable prompting mechanism can output optimization strategies with relatively good results.

[0107] In the above embodiments of this application, optimizing the prompt information input to the intelligent agent set based on the optimization strategy includes: inputting the optimization strategy into the optimization model and obtaining the monitoring results output by the optimization model, wherein the monitoring results are used to characterize whether the intelligent agent set has completed the corresponding task; generating reward information for optimizing the prompt information based on the monitoring results; and determining the optimized prompt information from the optimization space of the prompt information using a search algorithm based on the reward information.

[0108] The reward information of the aforementioned agent set can be used to reward the coordination among multiple agents in the set. Positive rewards are given for good coordination, and negative rewards are given for poor coordination. Coordination can represent the ability of multiple agents to work together. For example, if a program agent automatically assumes that the red tones in a chart image are in the second row and first column without a visual retrieval agent, exceeding its own assigned role, it indicates a lack of coordination among the agents.

[0109] In one optional embodiment, an optimization strategy can be input into an optimization model to check whether multiple agents accurately receive information transmitted by other agents and complete their respective tasks, thereby determining the coordination of execution. The optimization model can be prompted to reflect on whether multiple agents accurately receive information transmitted by other agents and complete their respective tasks. Based on the monitoring results, it can be determined whether multiple agents in the agent set have completed the corresponding tasks. If the monitoring results show that the corresponding tasks can be completed, a corresponding reward can be given; if the monitoring results show that the corresponding tasks cannot be completed, no reward can be given or a corresponding penalty can be given.

[0110] The aforementioned search algorithms can include beam search, linear search, interpolation search, etc. Among these, beam search can be an optimization algorithm used to find a better solution. Beam search simulates the beam search process, approximating a better solution by continuously adjusting the search direction and range. A key feature of beam search is its ability to maintain multiple search directions during the search process and its rapid convergence to a better solution. Therefore, beam search has high efficiency and accuracy in practical applications.

[0111] The target agent mentioned above can be an agent with a high reward level, for example, an agent with a high reward score.

[0112] In an alternative embodiment, a collaborative reward mechanism can be combined with a conventional search algorithm to reduce the trade-off between exploration and development. The selection mechanism in the beam search algorithm can be replaced by a reward mechanism to select better prompts. The rewarded prompts can be input into the search algorithm, which then identifies multiple prompts with higher reward levels. This allows for efficient exploration of the optimization space of natural language prompts, thereby finding the prompt combination that maximizes the chart image.

[0113] In another alternative embodiment, an optimization strategy can be input into an optimization model, and the monitoring results output by the optimization model can be used to characterize whether the agent ensemble has completed the corresponding task. These monitoring results may include data on task completion progress, quality, and efficiency. Based on these monitoring results, reward information for the agent ensemble can be generated to incentivize it to perform tasks better. The reward information can be determined based on task completion, resource utilization, and other factors. Then, based on the reward information, a search algorithm can be used to determine optimized prompt information. The search algorithm can optimize the configuration of the agent ensemble based on the reward information to achieve better task execution. Through the optimization of the search algorithm, the optimized prompt information required to perform the task can be determined to achieve collaborative and optimized task execution. Through this series of optimization processes, the agent ensemble can achieve high efficiency, collaboration, and optimization when performing tasks, thereby improving the effectiveness and efficiency of task execution.

[0114] In the above embodiments of this application, the method further includes: outputting multiple subtasks; receiving multiple new subtasks obtained by adjusting the multiple subtasks; and performing multiple new subtasks on the chart image to obtain a processing result.

[0115] In one optional embodiment, multiple subtasks can be output to the user's client. If the user believes that the multiple subtasks are biased or that the multiple subtasks need further adjustment, the multiple subtasks can be adjusted to obtain multiple new subtasks. Multiple new subtasks can be executed on the chart image to obtain the processing result of the chart image.

[0116] In the above embodiments of this application, the method further includes: outputting execution agents and execution order corresponding to multiple sub-tasks; receiving a second feedback result, wherein the second feedback result is obtained by adjusting the execution agents corresponding to the multiple sub-tasks and / or adjusting the execution order; when the second feedback result is obtained by adjusting the execution agents corresponding to the multiple sub-tasks, executing corresponding sub-tasks on the chart image sequentially using one agent from the second feedback result according to the execution order to obtain a processing result; when the second feedback result is obtained by adjusting the execution order, executing corresponding sub-tasks on the chart image sequentially using one agent from the multiple sub-tasks according to the second feedback result to obtain a processing result; when the second feedback result is obtained by adjusting the execution agents and execution order corresponding to the multiple sub-tasks, executing corresponding sub-tasks on the chart image sequentially using one agent from the second feedback result according to the second feedback result to obtain the processing result.

[0117] In one optional embodiment, the execution agents and execution order corresponding to multiple subtasks can be output to the user's client so that the user can check whether the execution agents corresponding to the multiple subtasks need to be adjusted, or whether the execution order needs to be adjusted. If the second feedback result is obtained by adjusting the execution agents corresponding to the multiple subtasks, the corresponding subtasks of the chart image can be executed sequentially using one of the agents in the second feedback result according to the execution order, thereby obtaining a processing result with higher accuracy. If the second feedback result is obtained by adjusting the execution order, the corresponding subtasks of the chart image can be executed one by one using one of the execution agents corresponding to the multiple subtasks according to the adjusted execution order in the second feedback result, thereby obtaining a processing result with higher accuracy.

[0118] By adding interactive steps to adjust the execution agents corresponding to multiple subtasks and to adjust the execution order, users can easily intervene in the processing in a timely manner, so as to make the execution agents and execution order corresponding to multiple subtasks more accurate, thereby obtaining a more accurate processing result.

[0119] Figure 3 This is a schematic diagram of a structure for intelligent agent cooperative optimization according to an embodiment of this application, such as... Figure 3As shown, the current prompt can contain prompts from the following agents: planning agent, data retrieval agent, visual retrieval agent, and program agent. Text data for the processing task can be input into the planning agent, which then divides the task into sub-tasks, resulting in multiple sub-tasks output by the planning agent. The execution agents corresponding to these sub-tasks can then be used to complete these sub-tasks on the chart / image. During task execution, intermediate execution results from the execution agents corresponding to these sub-tasks can be obtained, such as... Figure 3 The planning agent can output the execution results of two agents: the data retrieval agent and the program agent. The data retrieval agent can output data 1, data 2, and data 3 from the chart data. The program agent outputs the red bar, which is in the second row and first column, exceeding the program agent's own division of labor, resulting in an erroneous feedback. The feedback from this erroneous program agent can be collected. The feedback content is: the red bar was assumed when the planning agent did not plan the visual retrieval agent. The suggestion is to introduce logic to identify when visual context is needed. Multiple agents in the agent set can be collaboratively rewarded to efficiently optimize multiple agents.

[0120] Current end-to-end multimodal large model methods suffer from insufficient capabilities in graph symbol reasoning. In the multi-agent prompting collaboration optimization method for graph symbol reasoning proposed in this application, a multi-agent collaborative system architecture is designed, combining the advantages of multiple different models to complete different sub-tasks, thereby enabling better graph symbol reasoning.

[0121] Current solutions that rely on external tools create a huge human design burden because the model itself is overly dependent on the manual design of natural language prompts. In this application, the text editing and self-reflection capabilities of the large model language can be utilized to automatically optimize prompts through two steps: a reliable prompting mechanism and efficient prompt optimization based on collaborative rewards. This eliminates the need for manual prompt design and greatly reduces labor costs.

[0122] This application designs a multi-agent collaborative system architecture for graph symbol reasoning. The system can introduce a planning agent to organize multiple execution agents corresponding to sub-tasks to complete the process of extracting information from graph data and performing symbol reasoning based on the extracted context. By decomposing the processing task according to the text data, multiple sub-tasks are obtained, and different models are used to complete their respective sub-tasks, which can better perform graph symbol reasoning.

[0123] This application also proposes a prompting mechanism method that can collect intermediate execution results of the agents corresponding to multiple subtasks in error instances, forming a set of multifaceted error feedback. It can leverage the self-reflection capability of the Large Language Model (LLM) to summarize the reasons why the agents corresponding to the multiple subtasks produced incorrect answers, and extract insightful domain knowledge. That is, visual problems involving charts and images generally ask for parts including the recognition of colors, patterns, or text annotations, so as to generate fine-grained suggestions to improve the prompting information of the agent set, thereby making targeted improvements to the agent set.

[0124] This application also proposes an efficient prompt optimization framework based on collaborative rewards. Under this framework, rewards can be given based on the collaborative cooperation between multiple agents, and better prompt combinations can be found efficiently in an infinite prompt optimization space, thereby greatly reducing training costs.

[0125] The proposed solution can automatically optimize prompts for multi-agent systems, freeing humans from the tedious engineering of natural language prompts and thus significantly reducing manual costs. The innovation of this application lies in designing a multi-agent collaborative system architecture for graph symbol reasoning, thereby improving graph data reasoning. A reliable prompting mechanism is proposed to accurately control the direction of optimization, thereby improving the accuracy of optimization. An efficient prompting optimization framework based on collaborative rewards is proposed. Under this framework, based on the collaborative rewards among multiple agents, better prompt combinations are efficiently found in an infinitely large prompting optimization space, greatly reducing training costs.

[0126] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0127] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0129] Example 2

[0130] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order.

[0131] Figure 4 This is a flowchart of an image processing method according to Embodiment 2 of this application, as follows: Figure 4 As shown, the method includes the following steps:

[0132] Step S402: Obtain the chart image and query information. The query information is used to characterize the processing task to be performed on the chart image.

[0133] The aforementioned query information can be input by the user on the client side, indicating the processing task to be performed on the chart image. The chart image and query information can be input via a dialog-based generative dialog interface, so that the processing result of the chart image can be output through the generative dialog interface.

[0134] Step S404: Divide the processing task based on the query information to obtain multiple sub-tasks with related relationships;

[0135] Step S406: Utilize the execution agents corresponding to multiple subtasks to perform multiple subtasks on the chart image to obtain the processing result of the chart image;

[0136] Step S408: Based on the processing result, generate the response information corresponding to the query information.

[0137] Through the above steps, a chart image and query information are obtained. The query information is used to characterize the processing task to be performed on the chart image. Based on the query information, the processing task is divided into multiple sub-tasks with related relationships. The execution agents corresponding to the multiple sub-tasks are used to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image. Based on the processing result, the response information corresponding to the query information is generated, thereby achieving the technical effect of improving the accuracy of chart image processing. It is easy to note that the processing task of chart image can be divided into multiple sub-tasks with related relationships, so that different agents can be used to process different sub-tasks. Compared with the limitations of using a single agent to process the task, this application can combine the advantages of multiple agents to improve the processing accuracy of multiple sub-tasks on the chart image, obtain a more accurate chart image processing result, and thus solve the technical problem of low processing accuracy of chart images in related technologies.

[0138] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0139] Example 3

[0140] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order.

[0141] Figure 5 This is a flowchart of an image processing method according to Embodiment 3 of this application, as follows: Figure 5 As shown, the method includes the following steps:

[0142] Step S502: In response to the input command applied to the operation interface, display the chart image and text data on the operation interface. The text data is used to characterize the processing task to be performed on the chart image.

[0143] The aforementioned user interface can be used to display charts, images, and text data. The interface includes various controls for user operation, thereby displaying charts, images, and text data.

[0144] Step S504: In response to the processing instructions applied to the operation interface, the processing result of the chart image is displayed on the operation interface. The processing result is obtained by using the execution agents corresponding to multiple sub-tasks to perform multiple related sub-tasks on the chart image. The multiple sub-tasks are obtained by dividing the processing tasks based on text data.

[0145] The above processing instructions can be generated by the user interacting with controls on the interface when they need to process chart images.

[0146] Through the above steps, in response to input commands applied to the operation interface, chart images and text data are displayed on the operation interface. The text data represents the processing task to be performed on the chart image. In response to processing commands applied to the operation interface, the processing result of the chart image is displayed on the operation interface. The processing result is obtained by using multiple sub-tasks corresponding to execution agents to perform multiple related sub-tasks on the chart image. These multiple sub-tasks are obtained by dividing the processing task based on text data, thus achieving the technical effect of improving the accuracy of chart image processing. It is noteworthy that the chart image processing task can be divided into multiple related sub-tasks, so that different agents can be used to process different sub-tasks. Compared to the limitations of using a single agent to process the task, this application can combine the advantages of multiple agents to improve the processing accuracy of executing multiple sub-tasks on the chart image, obtaining a more accurate chart image processing result, thereby solving the technical problem of low accuracy in chart image processing in related technologies.

[0147] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0148] Example 4

[0149] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order.

[0150] Figure 6 This is a flowchart of an image processing method according to Embodiment 4 of this application, as follows: Figure 6 As shown, the method includes the following steps:

[0151] Step S602: Obtain chart images and text data by calling the first interface;

[0152] The first interface includes a first parameter, the value of which includes a chart image and text data. The text data is used to characterize the processing task to be performed on the chart image.

[0153] The aforementioned first interface can be an interface for data interaction between the cloud server and the client. The client can pass chart images and text data into the interface function as the first parameter of the interface function to achieve the purpose of uploading chart images and text data to the cloud server.

[0154] Step S604: Divide the processing task based on the text data to obtain multiple sub-tasks with related relationships;

[0155] Step S606: Utilize the execution agents corresponding to the multiple sub-tasks to perform multiple sub-tasks on the chart image respectively, and obtain the processing result of the chart image;

[0156] Step S608: Output the processing result by calling the second interface.

[0157] The second interface includes a second parameter, and the value of the second parameter includes the processing result.

[0158] The aforementioned second interface can be an interface for data interaction between the cloud server and the client. The cloud server can pass the processing result into the interface function as the second parameter of the interface function, thereby achieving the purpose of sending the processing result to the client.

[0159] Through the above steps, a chart image and text data are obtained by calling a first interface, where the first interface includes a first parameter whose value includes the chart image and text data. The text data represents the processing task to be performed on the chart image. The processing task is divided based on the text data to obtain multiple sub-tasks with related relationships. The execution agents corresponding to these sub-tasks are used to execute the sub-tasks on the chart image, respectively, to obtain the processing result. The processing result is output by calling a second interface, where the second interface includes a second parameter whose value includes the processing result. This achieves the technical effect of improving the accuracy of chart image processing. It is noteworthy that the chart image processing task can be divided into multiple related sub-tasks, allowing different agents to be used for different sub-tasks. Compared to the limitations of using a single agent to process a task, this application can combine the advantages of multiple agents to improve the processing accuracy of executing multiple sub-tasks on the chart image, resulting in a more accurate chart image processing result, thereby solving the technical problem of low accuracy in chart image processing in related technologies.

[0160] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0161] Example 5

[0162] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 7 This is a schematic diagram of an image processing apparatus according to Embodiment 5 of this application, as shown below. Figure 7 As shown, the device 700 includes: an acquisition module 702, a division module 704, and an execution module 706.

[0163] The acquisition module is used to acquire chart images and text data, wherein the text data is used to represent the processing tasks to be performed on the chart image; the partitioning module is used to partition the processing tasks based on the text data to obtain multiple sub-tasks with related relationships; the execution module is used to use the execution agents corresponding to the multiple sub-tasks to execute the multiple sub-tasks on the chart image respectively to obtain the processing results of the chart image.

[0164] It should be noted that the acquisition module 702, the partitioning module 704, and the execution module 706 mentioned above correspond to steps S202 to S206 in Embodiment 1. The instances and application scenarios implemented by the two modules and their corresponding steps are the same, but they are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the server 10 provided in Embodiment 1.

[0165] In the above embodiments of this application, the partitioning module is further configured to input text data into the planning agent, so as to use the planning agent to partition the processing task and obtain multiple sub-tasks output by the planning agent.

[0166] In the above embodiments of this application, the execution module is further configured to determine the execution agents corresponding to multiple sub-tasks; determine the execution order of the execution agents corresponding to multiple sub-tasks based on the association relationship of multiple sub-tasks; and, according to the execution order, use one of the execution agents corresponding to multiple sub-tasks to execute the corresponding sub-tasks on the chart image in sequence to obtain the processing result.

[0167] In the above embodiments of this application, the execution module is further configured to extract chart data from the chart image using a data retrieval agent; generate a task program based on the chart data using a program agent; and execute the task program to obtain the processing result.

[0168] In the above embodiments of this application, the execution module is further configured to: extract chart data from the chart image using a data retrieval agent; extract visual information from the chart image based on text data using a visual retrieval agent; generate a task program based on the chart data and visual information using a program agent; and execute the task program to obtain the processing result.

[0169] In the above embodiments of this application, the device further includes: an output module and an optimization module.

[0170] The output module is used to output the processing result; receive the first feedback result corresponding to the processing result, wherein the first feedback result is used to indicate whether an error has occurred in the processing result; the acquisition module is also used to acquire the intermediate execution result output by the agent set when the first feedback result indicates that an error has occurred in the processing result, wherein the agent set includes the planning agent and the execution agents corresponding to multiple sub-tasks; the optimization module is used to optimize the prompt information input to the agent set based on the intermediate execution result.

[0171] In the above embodiments of this application, the optimization module is further configured to analyze the intermediate execution results using an optimization model to generate an optimization strategy for the intelligent agent set; and to optimize the prompt information input to the intelligent agent set based on the optimization strategy.

[0172] In the above embodiments of this application, the optimization module is further configured to input intermediate execution results into the optimization model and obtain optimization suggestion information output by the optimization model; input optimization suggestion information into the optimization model and obtain optimization strategy output by the optimization model.

[0173] In the above embodiments of this application, the optimization module is further configured to input the optimization strategy into the optimization model and obtain the monitoring results output by the optimization model, wherein the monitoring results are used to characterize whether the set of intelligent agents has completed the corresponding task; based on the monitoring results, reward information for optimizing the prompt information is generated; based on the reward information, the optimized prompt information is determined from the optimization space of the prompt information using a search algorithm.

[0174] In the above embodiments of this application, the device further includes a receiving module.

[0175] The output module is used to output multiple subtasks; the receiving module is used to receive multiple new subtasks obtained by adjusting the multiple subtasks; and the execution module is used to execute multiple new subtasks on the chart image to obtain the processing results.

[0176] In the above embodiments of this application, the output module is further configured to output the execution agents and execution order corresponding to multiple sub-tasks; the receiving module is further configured to receive a second feedback result, wherein the second feedback result is obtained by adjusting the execution agents corresponding to multiple sub-tasks and / or adjusting the execution order; the execution module is further configured to, when the second feedback result is obtained by adjusting the execution agents corresponding to multiple sub-tasks, sequentially use one agent from the second feedback result to execute the corresponding sub-tasks on the chart image according to the execution order, and obtain the processing result; the execution module is further configured to, when the second feedback result is obtained by adjusting the execution order, sequentially use one agent from the execution agents corresponding to multiple sub-tasks to execute the corresponding sub-tasks on the chart image according to the second feedback result, and obtain the processing result; the execution module is further configured to, when the second feedback result is obtained by adjusting the execution agents corresponding to the multiple sub-tasks and the execution order, sequentially use one agent from the second feedback result to execute the corresponding sub-tasks on the chart image according to the second feedback result, and obtain the processing result.

[0177] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0178] Example 6

[0179] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 8 This is a schematic diagram of an image processing apparatus according to Embodiment 6 of this application, as shown below. Figure 8 As shown, the device 800 includes: an acquisition module 802, a division module 804, an execution module 806, and a generation module 808.

[0180] The acquisition module is used to acquire chart images and query information, whereby the query information represents the processing task to be performed on the chart image. The partitioning module is used to partition the processing task based on the query information to obtain multiple sub-tasks with related relationships. The execution module is used to use the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image. The generation module is used to generate response information corresponding to the query information based on the processing result.

[0181] It should be noted that the acquisition module 802, division module 804, execution module 806, and generation module 808 mentioned above correspond to steps S402 to S408 in Embodiment 2. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the server 10 provided in Embodiment 1.

[0182] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0183] Example 7

[0184] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 9 This is a schematic diagram of an image processing apparatus according to Embodiment 7 of this application, as shown below. Figure 9 As shown, the device 900 includes: a first display module 902 and a second display module 904.

[0185] The first display module is used to respond to input commands applied to the operation interface and display chart images and text data on the operation interface. The text data is used to represent the processing task to be performed on the chart image. The second display module is used to respond to processing commands applied to the operation interface and display the processing result of the chart image on the operation interface. The processing result is the result obtained by using the execution agents corresponding to the multiple sub-tasks to perform multiple related sub-tasks on the chart image. The multiple sub-tasks are the result obtained by dividing the processing task based on text data.

[0186] It should be noted that the first display module 902 and the second display module 904 mentioned above correspond to steps S502 to S504 in Embodiment 3. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the server 10 provided in Embodiment 1.

[0187] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0188] Example 8

[0189] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided. Figure 10 This is a schematic diagram of an image processing apparatus according to Embodiment 8 of this application, as shown below. Figure 10 As shown, the device 1000 includes: an acquisition module 1002, a division module 1004, an execution module 1006, and an output module 1008.

[0190] The acquisition module is used to acquire chart images and text data by calling a first interface, wherein the first interface includes a first parameter, the value of which includes the chart image and text data, and the text data is used to represent the processing task to be performed on the chart image; the partitioning module is used to partition the processing task based on the text data to obtain multiple sub-tasks with related relationships; the execution module is used to use the execution agents corresponding to the multiple sub-tasks to execute the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; the output module is used to output the processing result by calling a second interface, wherein the second interface includes a second parameter, the value of which includes the processing result.

[0191] It should be noted that the acquisition module 1002, partitioning module 1004, execution module 1006, and output module 1008 mentioned above correspond to steps S602 to S608 in Embodiment 4. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of a device and run in the server 10 provided in Embodiment 1.

[0192] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0193] Example 9

[0194] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic devices. Optionally, in this embodiment, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.

[0195] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0196] In this embodiment, the computer terminal described above can execute the program code in the method.

[0197] Optionally, Figure 11This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 11 As shown, the electronic device A may include: one or more (only one is shown in the figure) processors 102, memory 104, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to the radio frequency module, the audio module, and the display.

[0198] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0199] The processor can invoke information and applications stored in the memory through a transmission device to perform the following steps: acquiring chart images and text data, wherein the text data is used to characterize the processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with correlation; and using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image.

[0200] Optionally, the processor may also execute program code that performs the following steps: inputting text data into the planning agent to divide the processing task using the planning agent and obtaining multiple sub-tasks output by the planning agent.

[0201] Optionally, the processor may also execute program code that performs the following steps: determining the execution agents corresponding to multiple subtasks; determining the execution order of the execution agents corresponding to the multiple subtasks based on the relationship between the multiple subtasks; and, according to the execution order, sequentially using one of the execution agents corresponding to the multiple subtasks to perform the corresponding subtasks on the chart image to obtain the processing result.

[0202] Optionally, the processor may also execute program code that performs the following steps: extracting chart data from the chart image using a data retrieval agent; generating a task program based on the chart data using a program agent; executing the task program; and obtaining the processing result.

[0203] Optionally, the processor may also execute program code for the following steps: extracting chart data from the chart image using a data retrieval agent; extracting visual information from the chart image based on text data using a visual retrieval agent; generating a task program based on the chart data and visual information using a program agent, and executing the task program to obtain the processing result.

[0204] Optionally, the processor may also execute program code that performs the following steps: outputting the processing result; receiving a first feedback result corresponding to the processing result, wherein the first feedback result is used to characterize whether an error has occurred in the processing result; if the first feedback result characterizes that an error has occurred in the processing result, obtaining the intermediate execution result output by the agent set, wherein the agent set includes a planning agent and execution agents corresponding to multiple subtasks; and optimizing the prompt information input to the agent set based on the intermediate execution result.

[0205] Optionally, the processor may also execute program code that performs the following steps: analyzes intermediate execution results using an optimization model to generate an optimization strategy for the agent set; and optimizes the prompts input to the agent set based on the optimization strategy.

[0206] Optionally, the processor may also execute program code that performs the following steps: inputting intermediate execution results into the optimization model and obtaining optimization suggestion information output by the optimization model; inputting optimization suggestion information into the optimization model and obtaining optimization strategy output by the optimization model.

[0207] Optionally, the processor may also execute program code for the following steps: inputting the optimization strategy into the optimization model and obtaining the monitoring results output by the optimization model, wherein the monitoring results are used to characterize whether the agent set has completed the corresponding task; generating reward information to optimize the prompt information based on the monitoring results; and determining the optimized prompt information from the optimization space of the prompt information using a search algorithm based on the reward information.

[0208] Optionally, the processor may also execute program code that performs the following steps: outputs multiple subtasks; receives multiple new subtasks obtained by adjusting the multiple subtasks; performs multiple new subtasks on the chart image to obtain the processing result.

[0209] Optionally, the processor may also execute program code with the following steps: outputting the execution agents and execution order corresponding to multiple subtasks; receiving a second feedback result, wherein the second feedback result is obtained by adjusting the execution agents corresponding to the multiple subtasks and / or adjusting the execution order; if the second feedback result is obtained by adjusting the execution agents corresponding to the multiple subtasks, then, according to the execution order, sequentially using one agent from the second feedback result to execute the corresponding subtasks on the chart image to obtain a processing result; if the second feedback result is obtained by adjusting the execution order, then, according to the second feedback result, sequentially using one agent from the execution agents corresponding to the multiple subtasks to execute the corresponding subtasks on the chart image to obtain a processing result; if the second feedback result is obtained by adjusting the execution agents and execution order corresponding to the multiple subtasks, then, according to the second feedback result, sequentially using one agent from the second feedback result to execute the corresponding subtasks on the chart image to obtain the processing result.

[0210] The processor can invoke information and application programs stored in the memory through a transmission device to perform the following steps: acquiring a chart image and query information, wherein the query information is used to characterize the processing task to be performed on the chart image; dividing the processing task based on the query information to obtain multiple sub-tasks with correlation; using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; and generating response information corresponding to the query information based on the processing result.

[0211] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: in response to input instructions applied to the operation interface, displaying chart images and text data on the operation interface, wherein the text data is used to characterize the processing task to be performed on the chart image; in response to processing instructions applied to the operation interface, displaying the processing result of the chart image on the operation interface, wherein the processing result is obtained by using the execution agents corresponding to the multiple sub-tasks to perform multiple related sub-tasks on the chart image, and the multiple sub-tasks are obtained by dividing the processing task based on the text data.

[0212] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring chart images and text data by calling a first interface, wherein the first interface includes a first parameter, the value of which includes the chart image and text data, and the text data characterizes the processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with related relationships; using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; and outputting the processing result by calling a second interface, wherein the second interface includes a second parameter, the value of which includes the processing result.

[0213] By employing the embodiments of this application, chart images and text data are acquired, with the text data representing the processing tasks to be performed on the chart image. The processing tasks are divided based on the text data to obtain multiple sub-tasks with interrelationships. The execution agents corresponding to these sub-tasks are then used to execute the sub-tasks on the chart image, obtaining the processing result of the chart image. This achieves the technical effect of improving the accuracy of chart image processing. It is noteworthy that the processing task of the chart image can be divided into multiple interrelationships, allowing different agents to be used for different sub-tasks. Compared to the limitations of using a single agent to process the task, this application can combine the advantages of multiple agents to improve the processing accuracy of executing multiple sub-tasks on the chart image, obtaining a more accurate chart image processing result, thereby solving the technical problem of low accuracy in chart image processing in related technologies.

[0214] Those skilled in the art will understand that, Figure 11 The structure shown is for illustrative purposes only; the electronic device can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, PDA, mobile internet device (MID), PAD, and other terminal devices. Figure 11 This does not limit the structure of the aforementioned electronic device. For example, electronic device A may include more or fewer components (such as network interfaces, display devices, etc.) than shown in the figure, or may have a different configuration than shown in the figure.

[0215] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0216] Example 10

[0217] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.

[0218] Optionally, in this embodiment, the storage medium may be located in any one of the electronic devices in the group of electronic devices in the computer network, or in any one of the mobile terminals in the group of mobile terminals.

[0219] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring a chart image and text data, wherein the text data is used to characterize the processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with related relationships; and using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image.

[0220] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: inputting text data into the planning agent to divide the processing task using the planning agent, and obtaining multiple subtasks output by the planning agent.

[0221] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: determining execution agents corresponding to multiple subtasks; determining the execution order of the execution agents corresponding to the multiple subtasks based on the association relationship between the multiple subtasks; and, in accordance with the execution order, sequentially using one of the execution agents corresponding to the multiple subtasks to perform the corresponding subtasks on the chart image to obtain the processing result.

[0222] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: extracting chart data from a chart image using a data retrieval agent; generating a task program based on the chart data using a program agent, and executing the task program to obtain processing results.

[0223] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: extracting chart data from a chart image using a data retrieval agent; extracting visual information from the chart image based on text data using a visual retrieval agent; generating a task program based on the chart data and visual information using a program agent, and executing the task program to obtain processing results.

[0224] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: outputting a processing result; receiving a first feedback result corresponding to the processing result, wherein the first feedback result is used to characterize whether an error has occurred in the processing result; if the first feedback result characterizes that an error has occurred in the processing result, obtaining an intermediate execution result output by a set of agents, wherein the set of agents includes a planning agent and execution agents corresponding to multiple subtasks; and optimizing the prompt information input to the set of agents based on the intermediate execution result.

[0225] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: analyzing intermediate execution results using an optimization model to generate an optimization strategy for the agent set; and optimizing the prompt information input to the agent set based on the optimization strategy.

[0226] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: inputting intermediate execution results into the optimization model and obtaining optimization suggestion information output by the optimization model; inputting optimization suggestion information into the optimization model and obtaining optimization strategy output by the optimization model.

[0227] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: inputting an optimization strategy into an optimization model and obtaining monitoring results output by the optimization model, wherein the monitoring results are used to characterize whether the agent set has completed the corresponding task; generating reward information to optimize the prompt information based on the monitoring results; and determining the optimized prompt information from the optimization space of the prompt information using a search algorithm based on the reward information.

[0228] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: outputting multiple subtasks; receiving multiple new subtasks obtained by adjusting the multiple subtasks; performing multiple new subtasks on the chart image to obtain processing results.

[0229] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: outputting execution agents and execution order corresponding to multiple subtasks; receiving a second feedback result, wherein the second feedback result is obtained by adjusting the execution agents corresponding to the multiple subtasks and / or adjusting the execution order; if the second feedback result is obtained by adjusting the execution agents corresponding to the multiple subtasks, sequentially using one agent from the second feedback result to perform the corresponding subtasks on the chart image according to the execution order, to obtain a processing result; if the second feedback result is obtained by adjusting the execution order, sequentially using one agent from the multiple subtasks to perform the corresponding subtasks on the chart image according to the second feedback result, to obtain a processing result; if the second feedback result is obtained by adjusting the execution agents and execution order corresponding to the multiple subtasks, sequentially using one agent from the second feedback result to perform the corresponding subtasks on the chart image according to the second feedback result, to obtain the processing result.

[0230] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring a chart image and query information, wherein the query information is used to characterize the processing task to be performed on the chart image; dividing the processing task based on the query information to obtain multiple sub-tasks with related relationships; using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; and generating response information corresponding to the query information based on the processing result.

[0231] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in response to an input instruction applied to the operation interface, displaying a chart image and text data on the operation interface, wherein the text data is used to characterize the processing task to be performed on the chart image; in response to a processing instruction applied to the operation interface, displaying the processing result of the chart image on the operation interface, wherein the processing result is obtained by using the execution agents corresponding to the multiple sub-tasks to perform multiple related sub-tasks on the chart image respectively, and the multiple sub-tasks are the result of dividing the processing task based on the text data.

[0232] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining a chart image and text data by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the chart image and text data, and the text data is used to characterize the processing task to be performed on the chart image; dividing the processing task based on the text data to obtain multiple sub-tasks with related relationships; using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; and outputting the processing result by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter includes the processing result.

[0233] By employing the embodiments of this application, chart images and text data are acquired, with the text data representing the processing tasks to be performed on the chart image. The processing tasks are divided based on the text data to obtain multiple sub-tasks with interrelationships. The execution agents corresponding to these sub-tasks are then used to execute the sub-tasks on the chart image, obtaining the processing result of the chart image. This achieves the technical effect of improving the accuracy of chart image processing. It is noteworthy that the processing task of the chart image can be divided into multiple interrelationships, allowing different agents to be used for different sub-tasks. Compared to the limitations of using a single agent to process the task, this application can combine the advantages of multiple agents to improve the processing accuracy of executing multiple sub-tasks on the chart image, obtaining a more accurate chart image processing result, thereby solving the technical problem of low accuracy in chart image processing in related technologies.

[0234] Example 11

[0235] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.

[0236] Example 12

[0237] Embodiments of this application also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.

[0238] Example 13

[0239] Embodiments of this application also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.

[0240] Example 14

[0241] Embodiments of this application also provide an image processing platform, comprising: a planning node, configured to acquire chart images and text data, and to divide the processing tasks to be performed on the chart images based on the text data to obtain multiple sub-tasks with related relationships, wherein the text data is used to characterize the processing tasks; and multiple execution nodes, wherein execution agents corresponding to the multiple sub-tasks are deployed on the multiple execution nodes, and the multiple execution nodes are configured to use the execution agents corresponding to the multiple sub-tasks to perform the sub-tasks corresponding to the execution agents on the chart images respectively, thereby obtaining the processing results of the chart images.

[0242] The aforementioned planning node can be a computing node in an image processing platform used for planning, and a planning agent can be placed in the planning node.

[0243] In the above embodiments of this application, a planning agent is deployed on the planning node. The planning node is used to input the text data into the planning agent so as to divide the processing task using the planning agent and obtain the multiple sub-tasks output by the planning agent.

[0244] In the above embodiments of this application, the plurality of execution nodes are used to determine the execution order of the plurality of execution nodes based on the association relationship of the plurality of sub-tasks, and according to the execution order, use one of the execution agents corresponding to the plurality of sub-tasks to execute the corresponding sub-tasks on the chart image in sequence to obtain the processing result.

[0245] The aforementioned execution node can be a computing node in an image processing platform used to perform tasks, and an execution agent can be installed in the execution node.

[0246] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0247] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0248] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0249] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0250] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0251] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0252] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An image processing method, characterized in that, include: Acquire a chart image and text data, wherein the text data is used to characterize the processing task to be performed on the chart image; The processing task is divided based on the text data to obtain multiple sub-tasks with related relationships. The execution agents corresponding to the multiple sub-tasks are used to execute the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; If the first feedback result corresponding to the processing result indicates that the processing result has an error, the intermediate execution result output by the agent set is obtained, wherein the agent set includes a planning agent and execution agents corresponding to multiple subtasks; The prompts input to the agent set are optimized based on the intermediate execution results.

2. The method according to claim 1, characterized in that, The process of dividing the processing task based on the text data to obtain multiple related sub-tasks includes: The text data is input into the planning agent to divide the processing task and obtain the multiple sub-tasks output by the planning agent.

3. The method according to claim 1, characterized in that, The process of using the execution agents corresponding to the multiple sub-tasks to perform the multiple sub-tasks on the chart image to obtain the processing result of the chart image includes: Based on the relationships between the multiple sub-tasks, the execution order of the execution agents corresponding to the multiple sub-tasks is determined; According to the execution order, one of the execution agents corresponding to the multiple sub-tasks is used to perform the corresponding sub-tasks on the chart image in turn to obtain the processing result.

4. The method according to claim 3, characterized in that, The execution agents corresponding to the multiple sub-tasks include: a data retrieval agent and a program agent; the step of sequentially using one of the execution agents corresponding to the multiple sub-tasks to perform the corresponding sub-tasks on the chart image according to the execution order, and obtaining the processing result, includes: The data retrieval agent is used to extract chart data from the chart image; The program agent generates a task program based on the chart data and executes the task program to obtain the processing result.

5. The method according to claim 3, characterized in that, The execution agents corresponding to the multiple sub-tasks include: a data retrieval agent, a visual retrieval agent, and a program agent; the step of sequentially using one of the execution agents corresponding to the multiple sub-tasks to perform the corresponding sub-task on the chart image according to the execution order, and obtaining the processing result, includes: The data retrieval agent is used to extract chart data from the chart image; The visual retrieval agent is used to extract visual information from the chart image based on the text data. The program agent generates a task program based on the chart data and the visual information, and executes the task program to obtain the processing result.

6. The method according to any one of claims 2 to 5, characterized in that, The optimization of the agent set based on the intermediate execution results includes: The intermediate execution results are analyzed using an optimization model to generate an optimization strategy for the set of intelligent agents; The prompt information input to the set of intelligent agents is optimized based on the optimization strategy.

7. The method according to claim 6, characterized in that, The step of analyzing the intermediate execution results using an optimization model to generate an optimization strategy for the agent set includes: The intermediate execution results are input into the optimization model, and the optimization suggestion information output by the optimization model is obtained; The optimization suggestion information is input into the optimization model, and the optimization strategy output by the optimization model is obtained.

8. The method according to claim 6, characterized in that, The optimization of the prompt information input to the set of agents based on the optimization strategy includes: The optimization strategy is input into the optimization model, and the monitoring results output by the optimization model are obtained, wherein the monitoring results are used to characterize whether the set of agents has completed the corresponding task; Based on the monitoring results, reward information is generated to optimize the prompt information; Based on the reward information, an optimized prompt information is determined from the optimization space of the prompt information using a search algorithm.

9. The method according to claim 1, characterized in that, The method further includes: Output the multiple subtasks; Receive multiple new subtasks obtained by adjusting the multiple subtasks; The multiple new subtasks are performed on the chart image to obtain the processing result.

10. The method according to claim 3, characterized in that, The method further includes: Output the execution agents and execution order corresponding to the multiple subtasks; Receive a second feedback result, wherein the second feedback result is obtained by adjusting the executing agents corresponding to the plurality of subtasks and / or adjusting the execution order; When the second feedback result is obtained by adjusting the execution agents corresponding to the multiple sub-tasks, the corresponding sub-tasks are executed on the chart image by one of the agents in the second feedback result in the order of execution to obtain the processing result. If the second feedback result is obtained by adjusting the execution order, then according to the second feedback result, one of the execution agents corresponding to the multiple sub-tasks is used to execute the corresponding sub-tasks on the chart image in sequence to obtain the processing result; If the second feedback result is obtained by adjusting the execution agents and execution order of the multiple sub-tasks, then according to the second feedback result, one of the agents in the second feedback result is used to execute the corresponding sub-tasks on the chart image in sequence to obtain the processing result.

11. An image processing method, characterized in that, include: Acquire a chart image and query information, wherein the query information is used to characterize the processing task to be performed on the chart image; The processing task is divided based on the query information to obtain multiple sub-tasks with related relationships. The execution agents corresponding to the multiple sub-tasks are used to execute the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; Based on the processing result, a response message corresponding to the query message is generated; If the first feedback result corresponding to the processing result indicates that the processing result has an error, the intermediate execution result output by the agent set is obtained, wherein the agent set includes a planning agent and execution agents corresponding to multiple subtasks; The prompts input to the agent set are optimized based on the intermediate execution results.

12. An image processing method, characterized in that, include: The chart image and text data are obtained by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter includes the chart image and the text data, and the text data is used to characterize the processing task to be performed on the chart image; The processing task is divided based on the text data to obtain multiple sub-tasks with related relationships. The execution agents corresponding to the multiple sub-tasks are used to execute the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image; The processing result is output by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the processing result; If the first feedback result corresponding to the processing result indicates that the processing result has an error, the intermediate execution result output by the agent set is obtained, wherein the agent set includes a planning agent and execution agents corresponding to multiple subtasks; The prompts input to the agent set are optimized based on the intermediate execution results.

13. An image processing platform, characterized in that, include: A planning node is used to acquire chart images and text data, and to divide the processing tasks to be performed on the chart images based on the text data to obtain multiple sub-tasks with related relationships, wherein the text data is used to characterize the processing tasks; Multiple execution nodes are configured with execution agents corresponding to the multiple sub-tasks. The multiple execution nodes are used to execute the sub-tasks corresponding to the execution agents of the multiple sub-tasks on the chart image respectively to obtain the processing result of the chart image. The plurality of execution nodes are further configured to, when the first feedback result corresponding to the processing result indicates that the processing result has an error, obtain intermediate execution results output by the agent set, wherein the agent set includes planning agents and execution agents corresponding to the plurality of subtasks; and optimize the prompt information input to the agent set based on the intermediate execution results.

14. The platform according to claim 13, characterized in that, A planning agent is deployed on the planning node. The planning node is used to input the text data into the planning agent so that the planning agent can divide the processing task and obtain the multiple sub-tasks output by the planning agent.

15. The platform according to claim 13, characterized in that, The plurality of execution nodes are used to determine the execution order of the plurality of execution nodes based on the association relationship of the plurality of sub-tasks, and according to the execution order, use one of the execution agents corresponding to the plurality of sub-tasks to execute the corresponding sub-tasks on the chart image in turn to obtain the processing result.

16. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 15.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 15.

18. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Cross-modal agent implementation method

    CN117521710A

  • Data generation method, electronic equipment and storage medium

    CN117556026A