Business processing method and device based on computing power of intelligent computing center

By introducing task planning models and executors into the intelligent computing center, the problem that the agent cannot meet the changing business needs is solved, and more efficient and reliable business processing is achieved.

CN120259852APending Publication Date: 2025-07-04DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510322466.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing agents cannot meet the changing business needs in actual applications, and the processing is difficult and inefficient.

Method used

By introducing a task planning model and a task executor in the intelligent computing center, it receives business requirements information, performs dynamic planning and execution, uses the task planning model to generate task planning information, and executes image data by the task executioner to obtain execution results, and supports temporary writing of adapter functions to cover the basic function vacancy.

Benefits of technology

It reduces the difficulty of the agent in processing business demand information, improves the probability of successful execution of task planning information, and improves the efficiency of reasoning for business demand information and the reliability of execution results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259852A_ABST
    Figure CN120259852A_ABST
Patent Text Reader

Abstract

The invention provides a computing power business processing method and device based on an intelligent computing center, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the steps: S1, receiving business demand information based on an intelligent agent; s2, based on the task planning large model, performing task planning by taking target data in the business demand information as a planning target to obtain task planning information; s3, based on a task executor, taking the image data in the business demand information as an executed object, taking the task planning information as an execution process, performing planning execution, and obtaining an execution result; and S4, outputting the execution result based on the intelligent agent. According to the method and the system, variable business requirements in practical application can be met by utilizing a dynamic planning measure, and meanwhile, task planning operation and planning execution operation are split, so that the processing difficulty of an intelligent agent on business requirement information can be reduced, and the probability that the task planning information is successfully executed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and in particular, to a service processing method and device based on the computing power of an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] Currently, the intelligent agents constructed based on the existing technology cannot meet the changing business requirements in actual applications. Summary of the Invention

[0008] The purpose of the present disclosure is to provide a service processing method and device based on the computing power of an intelligent computing center, which is used to solve the problem that the intelligent agents constructed based on the existing technology cannot meet the changing business requirements in actual applications.

[0009] To solve the above technical problems, the present invention is implemented as follows:

[0010] In a first aspect, the present invention provides a service processing method based on the computing power of an intelligent computing center, including:

[0011] Step S1: Based on the agent, receive the business requirement information input by the user. Among them, the agent is a computer vision agent, and the business requirement information includes: image data and target data. The image data is used to represent the image to be processed, and the target data is used to represent the processing target for processing the image to be processed;

[0012] Step S2: Based on the task planning large model, use the target data as the planning target for task planning to obtain task planning information. Among them, the task planning information includes multiple step data bodies, and the multiple step data bodies correspond to multiple business processing steps one by one. The multiple business processing steps are used to achieve the processing target, and the task planning large model is the planning component of the agent;

[0013] Step S3: Based on the task executor, use the image data as the object to be executed and the task planning information as the execution process for planning execution to obtain an execution result. Among them, the task executor is the execution component of the agent;

[0014] Step S4: Output the execution result based on the agent.

[0015] In one embodiment, step S3 includes:

[0016] Step S31: Based on the task executor, use the image data as the object to be executed, and sequentially call multiple functions using the multiple step data bodies for planning execution operations to obtain the execution result;

[0017] Among them, during the planning execution operation, when the preset multiple basic functions do not include the target function, the task executor performs function encoding based on the step data body corresponding to the target function to obtain an adapter function, and uses the adapter function as the target function for calling. The target function is one of the multiple functions, and the multiple basic functions are used to achieve the general functions in the visual processing service corresponding to the agent;

[0018] The multiple basic functions correspond to multiple basic function modes one by one. The basic function modes are used to train the large model to use the corresponding basic functions, and the task planning large model is trained based on the multiple basic function modes.

[0019] In one embodiment, after step S1, the method further includes:

[0020] Step S5: In the case where the task planning in step S2 fails and / or the planning execution in step S3 fails, determine that the agent has not successfully processed the business requirement information, and output abnormal information based on the agent.

[0021] In one embodiment, after the step S5, the method further includes:

[0022] Step S6, obtaining anomaly optimization information, where the anomaly optimization information is used to repair the anomaly corresponding to the anomaly information;

[0023] Step S7, optimizing the agent based on the anomaly optimization information and the anomaly information;

[0024] Wherein, the anomaly information includes at least one of the following: anomaly type, anomaly code, anomaly error reporting data, service requirement information corresponding to the anomaly, and task planning information corresponding to the anomaly.

[0025] In one embodiment, the step S1 includes:

[0026] Step S11, receiving the original requirement information input by the user;

[0027] Step S12, when the original requirement information is the requirement information of the computer vision service, determining the original requirement information as the service requirement information.

[0028] In one embodiment, the service requirement information further includes: service supplementary information, where the service supplementary information is used to supplement the description of the processing target; the service supplementary information is based on the user identity information of the user and / or the historical operation information of the user.

[0029] In a second aspect, the present invention further provides a service processing device based on the computing power of an intelligent computing center, including:

[0030] A receiving module, configured to receive the service requirement information input by the user based on an agent, where the agent is a computer vision agent, and the service requirement information includes: image data and target data, the image data is used to represent the image to be processed, and the target data is used to represent: the processing target for processing the image to be processed;

[0031] A task planning module, configured to perform task planning based on a task planning large model with the target data as the planning target to obtain task planning information, where the task planning information includes a plurality of step data bodies, and the plurality of step data bodies correspond one-to-one to a plurality of service processing steps, and the plurality of service processing steps are used to achieve the processing target, and the task planning large model is a planning component of the agent;

[0032] A planning execution module, configured to perform planning execution based on a task executor with the image data as the object to be executed and the task planning information as the execution process to obtain an execution result, where the task executor is an execution component of the agent;

[0033] A result output module for outputting the execution result based on the agent.

[0034] In a third aspect, the present invention further provides a server, including: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps in the service processing method based on the computing power of the intelligent computing center as described in the first aspect above.

[0035] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the service processing method based on the computing power of the intelligent computing center as described in the first aspect above.

[0036] In a fifth aspect, the present invention provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps in the service processing method based on the computing power of the intelligent computing center as described in the first aspect above.

[0037] In the present invention, after the agent receives the service requirement information input by the user, the task planning large model is used as the planning component of the agent to implement the dynamic planning of the service requirement information, and the task planning information matching the service requirement information is obtained. Then, the task executor is used as the execution component of the agent to plan and execute based on the foregoing task planning information to obtain the execution result and output it. Among them, splitting the task planning operation and the planning execution operation can reduce the processing difficulty of the agent for the service requirement information, increase the probability of successful execution of the task planning information, improve the efficiency of the reasoning work for the service requirement information, so as to meet the changing service requirements in actual applications while using dynamic planning measures and ensure the reliability of the execution result output for the service requirements. Description of the Drawings

[0038] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0039] Figure 1 is a schematic flowchart of a service processing method using the computing power of an intelligent computing center provided by the present invention;

[0040] Figure 2 is an architecture diagram of a two-stage agent automation workflow provided by the present invention;

[0041] Figure 3 It is a schematic diagram of the results of an automated workflow of the Qwen2.5-72B-Instruct experimental agent provided by the present invention;

[0042] Figure 4 It is a schematic diagram of the structure of a service processing device that utilizes the computing power of an intelligent computing center provided by the present invention;

[0043] Figure 5 It is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the scope of protection of the present invention.

[0045] First, the technical terms related to the present invention will be briefly described below.

[0046] The "computing power" referred to in the present invention is the ability of a computer device or a computing / data center to process information, which is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to process information data and output a target result. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0047] The "computational power" (Computational Power, CP) referred to in the present invention is a kind of ability of a data center server to process data and output results, which is a comprehensive index to measure the computing ability of a data center and includes general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18 FLOPS). The larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0048] The "Network Power (NP)" described in the present invention is an indication of the data transmission capacity of computing power facilities, and is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. Network Power involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability. In the present invention, the video memory bandwidth is adopted for Network Power.

[0049] The "Storage Power (SP)" described in the present invention is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, and includes external storage devices such as storage arrays and internal storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (Input / Output Operations Per Second / TB, IOPS / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0050] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.

[0051] The "new type of information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0052] The "computing power" described in the present invention includes general computing power, intelligent computing power, and super computing power.

[0053] The "general computing power" described in the present invention is the computing power provided by servers based on central processing unit (CPU) chips, and is used to support basic general computing such as cloud computing and edge computing.

[0054] The "intelligent computing power" described in the present invention is for various artificial intelligence innovation applications, and is a computing platform deployed on a large scale based on dedicated chips such as graphics processing unit (GPU), field programmable gate array (FPGA), and application specific integrated circuit (ASIC), such as natural language processing and machine vision.

[0055] The "super computing power" described in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0056] The "intelligent computing center" described in the present invention refers to a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0057] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".

[0058] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0059] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0060] The "supercomputing center" described in the present invention refers to, namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0061] The "computing power resources" described in the present invention refer to technologies and facilities required for the development of the digital society and having information computing, transmission, storage, and application capabilities, including, but not limited to, computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0062] The "models" and "large models" described in the present invention include, but are not limited to, "large language models" and "multi-modal large models".

[0063] The "large language model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0064] The "multimodal large models" described in the present invention refer to models that jointly train multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0065] The "agent" described in the present invention refers to an entity that can perceive the environment and take actions to achieve specific goals. It can be software, hardware, or a system, with autonomy, adaptability, and interaction capabilities. An agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then executes actions to affect the environment or achieve a predetermined goal. Agents are widely used in the field of artificial intelligence, commonly found in automation systems, robots, virtual assistants, and game characters, etc. The core lies in their ability to learn autonomously and evolve continuously to better complete tasks and adapt to complex environments.

[0066] The "base function" described in the present invention refers to a function pre-constructed by developers according to the general / high-frequency business operations of the business scenarios faced by the agent. For example, when the agent is specifically a computer vision agent, the general / high-frequency business operations of the business scenarios faced by the agent may include: image preprocessing operations, image segmentation operations, image noise filtering operations, image target detection operations, etc.

[0067] The "base function schema" described in the present invention refers to the "base function user manual" formed by developers based on the functions and usage of the base function. The base function schema may include items such as name, description, type, etc. Among them, the name item may include the function name, the input parameter names of the function, the output parameter names of the function, etc. The description item can be understood as the function description, the semantic description of the input parameters of the function, and the semantic description of the output parameters of the function; the type item includes the data format description of the input parameters of the function and the data format description of the output parameters of the function. The large model can master the uses and usages of the base function through the "base function schema".

[0068] The "task planning big model" (planner) described in the present invention refers to a big model for task planning. After the task rule big model masters the purpose and usage of the basic function through the basic function mode, it can be used to generate a static plan in one go. The static plan includes all business processing steps to complete business needs. Each business processing step includes its explanation (also called step description text), the name of the called method, and input parameter information (each input parameter and its corresponding assignment).

[0069] The "autocoder" described in the present invention refers to a large model for code writing, which is used to write an adapter function.

[0070] The "static planning information" mentioned in the present invention refers to the data representation of the static plan, also known as task planning information, which at least includes multiple step data bodies arranged in order, and the multiple step data bodies correspond one-to-one to multiple business processing steps. The multiple business processing steps are executed in sequence to complete the user's business needs.

[0071] The "step data body" described in the present invention is the data representation of a corresponding business processing step, including but not limited to: a step description text of a corresponding business processing step (used to describe the business processing step to indicate the function, purpose and / or execution method of the step, so as to facilitate the user or developer to understand and implement it), the step sequence number of the corresponding business processing step in the multiple business processing steps, the input data of the corresponding business processing step (such as the data name of the input data, data type requirements, data source, etc.), and the data processing method of the corresponding business processing step (calling the specific method through the method name).

[0072] The "pipeline" described in the present invention refers to the process of executing a static plan.

[0073] The "planning component" mentioned in the present invention refers to the part of the intelligent body that is responsible for task planning operations, wherein the task planning operations are specifically: based on the processing objectives reflected in the business demand information, a series of continuous business processing steps are generated, and each business processing step exists in the form of a step data body to facilitate the execution component to perform subsequent planning and execution operations.

[0074] The "execution component" described in the present invention refers to the part of the intelligent agent that is responsible for planning and executing operations, wherein the planning and execution operations are specifically: based on the task planning information, the corresponding functions are scheduled in sequence to implement the aforementioned continuous business processing steps and finally form the execution result; the corresponding function can be a pre-generated function (that is, a basic function) or an adapter function temporarily written based on the automatic encoder during the planning and execution operation.

[0075] Please refer to Figure 1 , Figure 1 which is a business processing method based on the computing power of an intelligent computing center provided by the present invention. As Figure 1 shown, it includes the following steps:

[0076] Step S1: Receive business requirement information input by the user based on the intelligent agent.

[0077] Among them, the intelligent agent is a computer vision intelligent agent, and the business requirement information includes: image data and target data. The image data is used to represent the image to be processed, and the target data is used to represent: the processing target for processing the image to be processed.

[0078] The computer vision intelligent agent can be used to process vision processing services, such as: image recognition service, target detection service, image segmentation service, pose estimation service, face recognition service, image generation service, video analysis service, augmented reality service, autonomous driving service, etc.

[0079] Among them, the image data can be the image to be processed, or the storage address or download address of the image to be processed.

[0080] It should also be noted that the image to be processed can be one or more.

[0081] The processing target can be understood as: the information that the user expects to obtain from the image to be processed.

[0082] For example, the processing target can indicate whether the image to be processed includes a person / animal / specific building (target detection scenario), the processing target can also indicate whether the image to be processed includes a user with entered identity information (access control scenario), and the processing target can also indicate a fused image or fused video (AR scenario) formed after fusing the image to be processed (virtual image) with a set image or set video (real image / video).

[0083] Step S2: Based on the task planning large model, use the target data as the planning target for task planning to obtain task planning information.

[0084] Among them, the task planning information includes multiple step data bodies, and the multiple step data bodies correspond one-to-one to multiple business processing steps. The multiple business processing steps are used to achieve the processing target, and the task planning large model is the planning component of the intelligent agent.

[0085] Step S3: Based on the task executor, use the image data as the object to be executed and the task planning information as the execution process for planned execution to obtain an execution result.

[0086] Among them, the task executor is the execution component of the intelligent agent.

[0087] Exemplarily, if the set image data is image A and the processing target is to detect oil stains on image A, the processing flow indicated by the task planning information is successively image preprocessing, object detection, segmenting and detecting oil stains, and outputting the detection result; the image preprocessing operation is implemented based on function 1, the object detection operation is implemented based on function 2, the segmenting and detecting oil stains operation is implemented based on function 3, and the outputting the detection result operation is implemented based on function 4. Then the execution process of step S3 is as follows:

[0088] Based on the task executor, functions 1 to 4 are successively scheduled to obtain an execution result. Among them, the input of function 1 is image A, the output of function 1 is the input of function 2, the output of function 2 is the input of function 3, the output of function 3 is the input of function 4, and the output of function 4 forms the execution result.

[0089] Step S4: Output the execution result based on the intelligent agent.

[0090] The execution result is used to respond to the service demand information. For example, when the service demand information is used to detect whether the image to be processed includes a set target, the execution result is used to indicate that the image to be processed includes the set target (or does not include the set target); when the service demand information indicates that the image to be processed is to be fused into a set image, the execution result is the fused image formed after the image to be processed is fused into the set image.

[0091] Exemplarily, the output of the execution result can be implemented by one or more of information pop-up windows, voice broadcasts, text message pushes, etc.

[0092] In the present invention, after the intelligent agent receives the service demand information input by the user, the task planning large model is used as the planning component of the intelligent agent to realize the dynamic planning of the service demand information, obtain the task planning information matching the service demand information, and then the task executor is used as the execution component of the intelligent agent to plan and execute based on the foregoing task planning information to obtain an execution result and output it; among them, splitting the task planning operation and the planning execution operation can reduce the processing difficulty of the intelligent agent for the service demand information, increase the probability that the task planning information is successfully executed, improve the efficiency of the reasoning work for the service demand information, so as to meet the changing service demands in actual applications while using dynamic planning measures and ensure the reliability of the execution result output for the service demand.

[0093] Moreover, both the core information transfer between various business processing steps in the task planning information and the core information transfer between the task planning large model and the task executor support manual handling. On the one hand, it can increase the transparency of the agent in the process of processing business requirement information, helping developers better discover potential problems in the process of processing business requirement information and make targeted improvements; on the other hand, it can effectively constrain the transferred core information, thereby improving the training efficiency and training effect of the task planning large model and the task executor.

[0094] In one embodiment, step S3 includes:

[0095] Step S31: Based on the task executor, using the image data as the execution object, sequentially call a plurality of functions by using the plurality of step data bodies to perform a planning execution operation to obtain the execution result.

[0096] Wherein, during the planning execution operation, when a plurality of preset basic functions do not include the target function, the task executor performs function encoding based on the step data body corresponding to the target function to obtain an adapter function, and calls the adapter function as the target function. The target function is one of the plurality of functions, and the plurality of basic functions are used to implement: general functions in the visual processing service corresponding to the agent;

[0097] The plurality of basic functions correspond to a plurality of basic function modes one by one. The basic function mode is used to train the large model to use the corresponding basic function, and the task planning large model is trained based on the plurality of basic function modes.

[0098] Exemplarily, the general functions in the visual processing service may include: image preprocessing (image size adjustment, image enhancement, denoising, color adjustment), object detection (identifying and locating specific objects such as faces, vehicles, goods, etc.), image segmentation (separating foreground and background, extracting feature regions), object tracking (identifying and tracking moving targets), generating reports (data statistics and visualization, generating analysis reports).

[0099] Among them, the general functions in the visual processing service can be set manually based on experience or obtained through data analysis. The process of obtaining the general functions in the visual processing service through data analysis is as follows:

[0100] Grab datasets corresponding to the visual processing service from different data sources;

[0101] Determine the data type corresponding to each data in the dataset (such as plain text, text with serial numbers or specific formats, flowcharts, audio, etc.), and adopt an analysis scheme matching the determined data type for data analysis to obtain the business processing flow information corresponding to each data. Each business processing flow information is formed by arranging multiple functions in an orderly manner;

[0102] Determine the functions that appear in the business processing flow information as candidate functions, and count the occurrence frequency of each candidate function in multiple business processing flow information (the number of occurrences divided by the total number of multiple business processing flow information). Determine several candidate functions with the highest occurrence frequency as the general functions in the visual processing business, or determine several candidate functions with an occurrence frequency higher than the set frequency threshold as the general functions in the visual processing business.

[0103] In this embodiment, based on the setting of multiple determined basic functions, it helps the task planning large model and the task executor to better cooperate and connect, and improves the probability that the task planning information output by the task planning large model can be successfully executed by the task executor.

[0104] Among them, build multiple basic functions based on the general functions in the visual processing business to reduce the probability of temporary coding adapter functions, thereby reducing the risk of function execution exceptions caused by adapter functions, and improving the probability that the task planning information output by the task planning large model can be successfully executed by the task executor.

[0105] Aiming at the problem that the basic functions may not cover all business requirements in the actual scenario, support temporary code writing to generate adapter functions that can fill the gaps in the basic functions to ensure comprehensive coverage of business requirements.

[0106] It should be noted that the task planning large model is obtained by pre-training and fine-tuning the large model.

[0107] Among them, the large model after pre-training can be called the initial model. The initial model has a certain task planning ability. The initial model can output corresponding planning data (plan) for the input business requirement information (also called user requirements). In the application, use the available planning data and the corresponding business requirement information as fine-tuning data to further fine-tune the model parameters and model output format of the initial model, which can standardize the data format of the planning data output by the model and improve the probability that the planning data output by the model can be successfully executed by the task executor. Among them, the available planning data refers to the planning data that can be successfully executed by the task executor.

[0108] In this embodiment, it can also be set that a task executor is introduced to participate in the fine-tuning of the task planning large model. In this case, the fine-tuning data used not only includes available planning data and corresponding business requirement information, but also includes the first execution result obtained by the task executor executing the available planning data; the model loss in the fine-tuning stage is used to represent: the difference between the planning data output by the initial model in the fine-tuning stage and the corresponding available planning data, and the difference between the second execution result obtained by the task executor executing the planning data output by the initial model in the fine-tuning stage and the aforementioned first execution result. Based on this setting, it helps the task planning large model and the task executor to perform more efficient information transmission, increase the probability of obtaining the execution result, as well as the accuracy and reliability of the obtained execution result.

[0109] In addition, in this embodiment, it can also be set that in the aforementioned fine-tuning stage, the parameter names and method names of the available planning data included in the fine-tuning data are masked, and / or a summary of the available planning data is added to the fine-tuning data as a hint to increase the probability that the task planning information output by the task planning large model is successfully executed.

[0110] In one embodiment, after the step S1, the method further includes:

[0111] Step S5: In the case where the task planning in the step S2 fails, and / or the planning execution in the step S3 fails, it is determined that the agent has not successfully processed the business requirement information, and abnormal information is output based on the agent.

[0112] Among them, in the case where the task planning large model fails to successfully output task planning information, or the task planning information output by the task planning large model fails the pre-check, it can be determined that the task planning in the step S2 fails.

[0113] In the case where the task executor cannot obtain the task planning information, or the task executor cannot correctly parse the task planning information, or the task executor cannot normally call the basic function, or the task executor cannot correctly perform function encoding to obtain the adapter function, it can be determined that the planning execution in the step S3 fails.

[0114] Among them, the abnormal information at least includes the abnormal type, and different abnormal types indicate different abnormal problems. For example: abnormal output of task planning information, abnormal pre-check of task planning information, abnormal acquisition of task planning information, abnormal parsing of task planning information, abnormal call of basic function, abnormal encoding of adapter function, etc.

[0115] In this embodiment, based on the output of the exception information, in the case that the agent fails to successfully process the service requirement information, it helps the user to quickly locate the cause of the exception and the location where the exception occurs, thereby accelerating the exception resolution efficiency and enhancing the user experience of using the agent.

[0116] In one embodiment, after the step S5, the method further includes:

[0117] Step S6, obtaining exception optimization information, where the exception optimization information is used to repair the exception corresponding to the exception information;

[0118] Step S7, optimizing the agent based on the exception optimization information and the exception information;

[0119] Wherein, the exception information includes at least one of the following: exception type, exception code, exception error data, service requirement information corresponding to the exception, and task planning information corresponding to the exception.

[0120] Exemplarily, the exception optimization information can be used as a positive example, and the exception information can be used as a negative example to fine-tune the task planning large model or the task executor, so as to teach the task planning large model to better output task planning information, or to teach the task executor to better complete the planned execution operation.

[0121] The exception code can be understood as: several codes within a set number of lines centered on / bounded by the exception location corresponding to the exception problem; or, the code part where the exception location corresponding to the exception problem is located.

[0122] The exception error data can be understood as: the data output based on the set exception throwing logic after the exception problem occurs.

[0123] When an exception problem occurs during the agent's processing of a certain service requirement information, the service requirement information is determined as the service requirement information corresponding to the exception problem.

[0124] When an exception problem occurs during the task executor's execution of a certain task planning information, the task planning information is determined as the task planning information corresponding to the exception.

[0125] In this embodiment, by supporting various types of data that may cause exceptions to be included in the exception information, it helps the user to more flexibly and quickly locate the cause of the exception, thereby accelerating the exception resolution efficiency.

[0126] In one embodiment, the step S1 includes:

[0127] Step S11, receiving the original requirement information input by the user;

[0128] Step S12, when the original requirement information is requirement information for a computer vision service, determine the original requirement information as the service requirement information.

[0129] In this embodiment, based on the above settings, it is to prevent the agent from receiving requirement information other than computer vision services, avoid responding to unnecessary requirement information, reduce the risk of the agent being contaminated by requirement information other than computer vision services, and at the same time avoid unnecessary waste of the agent's computing resources.

[0130] Among them, step S1 further includes:

[0131] Step S12, when the original requirement information is not requirement information for a computer vision service, output a prompt message to inform the user through the prompt message that the agent does not support processing / responding to requirement information for non-computer vision services.

[0132] Exemplarily, semantic feature extraction can be performed on the original requirement information to determine the first semantic feature corresponding to the original requirement information, and then calculate the feature similarity between the first semantic feature and the second semantic feature. When the feature similarity is lower than the set similarity threshold, it is determined that the original requirement information is not requirement information for a computer vision service; when the feature similarity is higher than or equal to the set similarity threshold, it is determined that the original requirement information is requirement information for a computer vision service; among them, the second semantic feature is used to represent the semantic feature of requirement information for a computer vision service.

[0133] In one embodiment, the service requirement information further includes: service supplementary information, and the service supplementary information is used to supplement and describe the processing target; the service supplementary information is based on the user's identity information and / or the user's historical operation information.

[0134] Exemplarily, the user identity information may include: user gender, user's affiliated industry, user's affiliated enterprise, the core product or core business of the user's affiliated enterprise, etc. The historical operation information may include: the service requirement information input by the user to the agent in the past and the corresponding task planning information / execution results, etc.

[0135] In this embodiment, based on the addition of the service supplementary information, the content of the service requirement information is enriched to help the task planning large model better understand the processing target indicated by the service requirement information, so as to generate more accurate and reliable task planning information, and while ensuring the successful execution of the task planning information, improve the reliability and accuracy of the finally obtained execution results.

[0136] It should be noted that the above service supplementary information can be automatically analyzed by the agent without manual input by the user.

[0137] For easy understanding, the examples are described as follows:

[0138] The present invention also provides a solution for implementing the aforementioned intelligent agent automated workflow by utilizing the powerful computing power and flexible resource allocation of the intelligent computing center. The corresponding architecture of the two-stage auto-pipeline can be as Figure 2 shown Figure 2 described in the relevant diagrams as follows:

[0139] User query: Users input their requests or requirement content.

[0140] Planner: Used to parse user queries and formulate an execution plan (plan). If the parsing fails, it may output an error that cannot be converted to the JSON format.

[0141] Plan list: A list of specific steps generated according to user queries.

[0142] Retrieve all-step function dependencies: Retrieve the function dependencies related to each step to ensure that all required functions are available.

[0143] Get executable input values:

[0144] Given values: Values directly provided to the function.

[0145] User inputs: Dynamic input values provided by users in the query.

[0146] Previous function outputs: Function output values from the previous step.

[0147] Media download: This link can also support the download of additional required media content.

[0148] Execute one step: Refers to the demonstration of the execution details of a specific step.

[0149] Execute function: Used to actually call the function to perform data processing work according to the obtained input values.

[0150] Base function: It performs specific basic functions and enables the agent to have the set basic business processing capabilities.

[0151] Adapter function (autoencoder): It serves as a supplement to the base function.

[0152] Retrieve depended function schemas: It obtains the structures or schemas of other dependent functions.

[0153] Autoencoder coding: It automatically generates the code involved in the above processes.

[0154] Exception code: If an error or exception occurs, the system will return a specific exception code for subsequent analysis and debugging. For example:

[0155] Exceptions can be classified into the following six categories:

[0156] 0: No exception, the pipeline execution is successful

[0157] 1: Other pipeline exceptions

[0158] 2: Autoencoder execution fails

[0159] 3: Base function execution fails

[0160] 4: Plan acquisition fails (usually due to the failure of converting plan to json)

[0161] 5: The variable placeholder in the plan is expressed incorrectly and the parsing fails.

[0162] It should be noted that applying the two-stage auto-pipeline of the agent has at least the following advantages:

[0163] 1. Higher accuracy: There are more cases where the given plan can be successfully executed;

[0164] 2. By decomposing tasks, the total inference time of the LLM can be shortened;

[0165] 3. By manually handling the core information transfer between steps, the scope of effective information can be clarified and the transparency of the system can be increased.

[0166] The results of experimenting with the two-stage auto-pipeline of the agent using the Qwen2.5-72B-Instruct tool for different data sources are as Figure 3 shown.

[0167] It is experimentally found that errors mainly occur in the basic function execution stage and the variable placeholder expression stage. Among them, the errors in the basic function execution stage include: incorrect parameter passing, confusion in using the image list and video path, some parameters unrelated to the input (such as thresholds), etc.; the errors in the variable placeholder expression stage include: importing external packages (lacking package dependencies), incorrect understanding and processing of complex parameters, etc. In addition, some errors are caused by poor compliance with the json format output by the plan.

[0168] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a service processing device based on the computing power of an intelligent computing center provided by the present invention. As Figure 4 shown, the service processing device 400 based on the computing power of the intelligent computing center includes:

[0169] A receiving module 401, configured to receive service requirement information input by a user based on an intelligent agent. Among them, the intelligent agent is a computer vision intelligent agent, and the service requirement information includes: image data and target data. The image data is used to represent the image to be processed, and the target data is used to represent: the processing target for processing the image to be processed;

[0170] A task planning module 402, configured to perform task planning based on a task planning large model, using the target data as the planning target to obtain task planning information. Among them, the task planning information includes multiple step data bodies, and the multiple step data bodies correspond to multiple service processing steps one by one. The multiple service processing steps are used to achieve the processing target, and the task planning large model is the planning component of the intelligent agent;

[0171] A planning execution module 403, configured to perform planning execution based on a task executor, using the image data as the execution object and the task planning information as the execution process to obtain an execution result. Among them, the task executor is the execution component of the intelligent agent;

[0172] A result output module 404, configured to output the execution result based on the intelligent agent.

[0173] In one embodiment, the planning execution module 403 includes:

[0174] An execution unit, configured to perform a planning execution operation based on the task executor, using the image data as the execution object, and sequentially calling multiple functions by using the multiple step data bodies to obtain the execution result;

[0175] Among them, during the process of executing the plan, when the preset multiple basic functions do not include the target function, the task executor performs function encoding based on the step data body corresponding to the target function to obtain an adapter function, and calls the adapter function as the target function. The target function is one of the multiple functions, and the multiple basic functions are used to implement the general functions in the visual processing service corresponding to the intelligent agent;

[0176] The multiple basic functions correspond to multiple basic function patterns one by one. The basic function patterns are used to train the large model to use the corresponding basic functions, and the task planning large model is trained based on the multiple basic function patterns.

[0177] In one embodiment, the service processing device 400 based on the computing power of the intelligent computing center further includes:

[0178] An exception output module, configured to determine that the intelligent agent fails to successfully process the service requirement information and output exception information based on the intelligent agent when the task planning module 402 fails in task planning and / or the plan execution module 403 fails in plan execution.

[0179] In one embodiment, the service processing device 400 based on the computing power of the intelligent computing center further includes:

[0180] An optimization information acquisition module, configured to acquire exception optimization information, where the exception optimization information is used to repair the exception corresponding to the exception information;

[0181] An intelligent agent optimization module, configured to optimize the intelligent agent based on the exception optimization information and the exception information;

[0182] Among them, the exception information includes at least one of the following: exception type, exception code, exception error data, service requirement information corresponding to the exception, and task planning information corresponding to the exception.

[0183] In one embodiment, the receiving module 401 includes:

[0184] A receiving unit, configured to receive the original requirement information input by the user;

[0185] A determining unit, configured to determine the original requirement information as the service requirement information when the original requirement information is the requirement information of the computer vision service.

[0186] In one embodiment, the service requirement information further includes: service supplementary information, where the service supplementary information is used to supplement the description of the processing target; the service supplementary information is based on the user identity information of the user and / or the historical operation information of the user.

[0187] The business processing device based on the computing power of the intelligent computing center provided by the present invention can implement each process of the above-mentioned business processing method based on the computing power of the intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0188] It should be noted that the business processing device based on the computing power of the intelligent computing center in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0189] The present invention also provides an electronic device. Refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction running on the memory 501. When the program or instruction is executed by the processor 502, it can implement Figure 1 any step in the corresponding embodiment of the business processing method based on the computing power of the intelligent computing center and achieve the same beneficial effects. They will not be elaborated here.

[0190] Among them, the processor 502 can be a CPU, an ASIC, an FPGA, or a GPU.

[0191] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned embodiment of the business processing method based on the computing power of the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0192] The present invention also provides a readable storage medium. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1 any step in the corresponding embodiment of the business processing method based on the computing power of the intelligent computing center and can achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0193] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In addition, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, the use of "and / or" in the present invention means at least one of the connected objects. For example, A and / or B and / or C means including seven cases: A alone, B alone, C alone, A and B both present, B and C both present, A and C both present, and A, B and C all present.

[0194] It should be noted that in this text, the term "comprising", "including" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of the present invention.

[0196] The above describes the embodiments of the present invention in conjunction with the accompanying drawings, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them fall within the protection scope of the present invention.

Claims

1. A business processing method based on the computing power of an intelligent computing center, characterized in that, Including: Step S1: Based on an agent, receive business requirement information input by a user. Among them, the agent is a computer vision agent, and the business requirement information includes: image data and target data. The image data is used to represent an image to be processed, and the target data is used to represent: a processing target for processing the image to be processed; Step S2: Based on a task planning large model, use the target data as a planning target for task planning to obtain task planning information. Among them, the task planning information includes multiple step data bodies, and the multiple step data bodies correspond to multiple business processing steps one by one. The multiple business processing steps are used to achieve the processing target, and the task planning large model is a planning component of the agent; Step S3: Based on a task executor, use the image data as an execution object and the task planning information as an execution process for planning execution to obtain an execution result. Among them, the task executor is an execution component of the agent; Step S4: Output the execution result based on the agent.

2. The method according to claim 1, wherein The step S3 includes: Step S31: Based on the task executor, use the image data as an execution object, and sequentially call multiple functions using the multiple step data bodies to perform a planning execution operation to obtain the execution result; Among them, during the planning execution operation, when a plurality of preset basic functions do not include a target function, the task executor performs function encoding based on the step data body corresponding to the target function to obtain an adapter function, and uses the adapter function as the target function for calling. The target function is one of the multiple functions, and the multiple basic functions are used to achieve: general functions in the vision processing business corresponding to the agent; The multiple basic functions correspond to multiple basic function modes one by one. The basic function modes are used to train the large model to use the corresponding basic functions, and the task planning large model is trained based on the multiple basic function modes.

3. The method according to claim 2, wherein After the step S1, the method further includes: Step S5: In the case where the task planning in the step S2 fails, and / or the planning execution in the step S3 fails, determine that the agent has not successfully processed the business requirement information, and output exception information based on the agent.

4. The method according to claim 3, characterized in that, After the step S5, the method further includes: Step S6: Obtain exception optimization information, where the exception optimization information is used to repair the exception corresponding to the exception information; Step S7: Optimize the agent based on the exception optimization information and the exception information; Among them, the exception information includes at least one of the following: exception type, exception code, exception error data, business requirement information corresponding to the exception, and task planning information corresponding to the exception.

5. The method according to claim 1, wherein The step S1 includes: Step S11: Receive the original requirement information input by the user; Step S12: In the case where the original requirement information is the requirement information of the computer vision business, determine the original requirement information as the business requirement information.

6. The method according to claim 1, wherein The business requirement information further includes: business supplementary information, which is used to supplement the description of the processing target; the business supplementary information is based on the user identity information of the user and / or the historical operation information of the user.

7. A service processing device based on the computing power of an intelligent computing center, characterized in that, including: a receiving module, configured to receive business requirement information input by a user based on an intelligent agent, where the intelligent agent is a computer vision intelligent agent, and the business requirement information includes: image data and target data, the image data is used to represent an image to be processed, and the target data is used to represent: a processing target for processing the image to be processed; a task planning module, configured to perform task planning based on a task planning large model with the target data as the planning target to obtain task planning information, where the task planning information includes a plurality of step data bodies, and the plurality of step data bodies correspond to a plurality of business processing steps one by one, and the plurality of business processing steps are used to implement the processing target, and the task planning large model is a planning component of the intelligent agent; a planning execution module, configured to perform planning execution based on a task executor with the image data as the execution object and the task planning information as the execution process to obtain an execution result, where the task executor is an execution component of the intelligent agent; a result output module, configured to output the execution result based on the intelligent agent.

8. A server, characterized in that, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, the steps of the business processing method based on the computing power of the intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the business processing method based on the computing power of the intelligent computing center as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that, including computer instructions, and when the computer instructions are executed by a processor, the steps of the business processing method based on the computing power generated by the intelligent computing center as described in any one of claims 1 to 6 are implemented.