Visual task processing method, device, equipment, medium and program product
By decoupling local computing power, calling external visual services and combining them with local rendering services, the high cost and low efficiency problems caused by GPU resource dependence in existing technologies are solved, and efficient processing of visual tasks and multi-scenario adaptability are achieved.
Patent Information
- Application Number
- CN202510657828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing technologies are highly dependent on GPU resources in a series of image or video generation tasks, resulting in high computing costs and low efficiency. In addition, they lack standardized interfaces and service support, making it difficult to operate efficiently in large-scale production environments.
By decoupling local computing power, calling external visual services to perform visual tasks, using task distribution services to dynamically schedule computing resources, and combining local rendering services to achieve efficient rendering and display of visual results.
It improves resource utilization efficiency and system scalability, ensures real-time rendering and display of visual task results, and adapts to dynamic needs in multiple scenarios.
Smart Images

Figure CN120256062B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a method, apparatus, device, medium, and program product for processing visual tasks. Background Art
[0002] In the current field of artificial intelligence generated content (AIGC), workflow-based architectures have become a key technology for automating complex tasks. Through this architecture, users can combine different functional modules to form customized image generation processes, greatly improving flexibility and efficiency.
[0003] However, since the workflow architecture needs to run the workflow and generate artificial intelligence tasks simultaneously, it is highly dependent on GPU resources. In practical applications, this not only leads to high computing costs, but also affects its efficiency when faced with a series of processing requirements of the workflow, especially in a series of image or video generation tasks.
[0004] Therefore, there is an urgent need for a visual task processing method that can liberate computing power. Summary of the Invention
[0005] In view of this, embodiments of this specification provide a method for processing a visual task. One or more embodiments of this specification also relate to a visual task processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0006] According to a first aspect of an embodiment of this specification, a method for processing a visual task is provided, comprising:
[0007] Responding to a call request for a target workflow, obtaining the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate a corresponding visual task, and the call request carries task parameters of the visual task;
[0008] Run the target workflow based on the task parameters, and when running to the visual task node in the target workflow, call the external visual service to execute the visual task corresponding to the visual task node, and obtain the visual result corresponding to the visual task node;
[0009] Render the visual results corresponding to each visual task node.
[0010] According to a second aspect of the embodiments of this specification, a visual task processing apparatus is provided, comprising:
[0011] an acquisition module configured to acquire the target workflow in response to a call request of the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate a corresponding visual task, and the call request carries task parameters of the visual task;
[0012] The calling module is configured to run the target workflow based on the task parameters, and when running to the visual task node in the target workflow, call the external visual service to execute the visual task corresponding to the visual task node, and obtain the visual result corresponding to the visual task node;
[0013] The rendering module is configured to render the visual results corresponding to each visual task node.
[0014] According to a third aspect of the embodiments of this specification, a computing device is provided, including:
[0015] memory and processor;
[0016] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned visual task processing method are implemented.
[0017] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which implements the steps of the above-mentioned visual task processing method when executed by a processor.
[0018] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above-mentioned visual task processing method when executed by a processor.
[0019] One embodiment of the present specification implements a method of obtaining a target workflow in response to a call request of the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate the corresponding visual task, and the call request carries the task parameters of the visual task; running the target workflow based on the task parameters, and when running to the visual task node in the target workflow, calling the external visual service to execute the visual task corresponding to the visual task node, and obtaining the visual result corresponding to the visual task node; and rendering the visual result corresponding to each visual task node. Flexible arrangement and execution of visual tasks are achieved through workflow scenario services. When the call request triggers the target workflow to run, the visual task node relies on the external visual service to complete the specific task, avoiding direct dependence on local computing power, thereby achieving decoupling of rendering and computing power, effectively improving resource utilization efficiency and system scalability, while ensuring that each visual task result can be rendered and displayed in real time to meet dynamic needs in multiple scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of a visual task processing method provided by one embodiment of this specification;
[0021] Figure 2 This is an architectural diagram of a visual task workflow management platform provided by one embodiment of this specification;
[0022] Figure 3 is a flowchart of a processing process of a visual task processing method provided by one embodiment of this specification;
[0023] Figure 4 This is a schematic diagram of the front-end interface of the first visual task workflow management platform provided by an embodiment of this specification;
[0024] Figure 5 This is a schematic diagram of the front-end interface of the second visual task workflow management platform provided by an embodiment of this specification;
[0025] Figure 6 This is a schematic diagram of the front-end interface of the third visual task workflow management platform provided in one embodiment of this specification;
[0026] Figure 7 This is a schematic diagram of the front-end interface of the fourth visual task workflow management platform provided by an embodiment of this specification;
[0027] Figure 8 This is a schematic diagram of the front-end interface of the fifth visual task workflow management platform provided by an embodiment of this specification;
[0028] Figure 9 This is a schematic diagram of the front-end interface of the sixth visual task workflow management platform provided by an embodiment of this specification;
[0029] Figure 10 This is a schematic diagram of the front-end interface of the seventh visual task workflow management platform provided by an embodiment of this specification;
[0030] Figure 11 This is a schematic diagram of the front-end interface of an eighth visual task workflow management platform provided by an embodiment of this specification;
[0031] Figure 12 This is a schematic diagram of a front-end interface of a visual task application management platform provided by an embodiment of this specification;
[0032] Figure 13 This is a schematic diagram of the structure of a visual task processing device provided by one embodiment of this specification;
[0033] Figure 14This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0034] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0035] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0036] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0037] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0038] First, the terms involved in one or more embodiments of this specification are explained.
[0039] AIGC (AI Generated Content): Artificial intelligence generated content.
[0040] Node: corresponds to a specific computing task or functional module, and is used to build the basic unit of the AIGC image generation process.
[0041] Workflow: A node-based task execution description used to build and execute complex AIGC tasks.
[0042] ComfyUI: Workflow-based visual programming software that supports nodes to expand their capabilities.
[0043] GPU (Graphics Processing Unit): Graphics processing unit, used for high-performance computing.
[0044] API (Application Programming Interface): Application programming interface.
[0045] Computing resources: Computing resources include CPU (Center Processing Unit, core processor) and GPU and other computing and storage devices.
[0046] Plugin-based computing power decoupling: Computing power resources are separated from core logic such as workflow editing and task submission through plug-in methods.
[0047] Configurable API: Create and encapsulate APIs through simple configuration.
[0048] Computing power routing: Computing power scheduling and task resource allocation based on dynamic configuration.
[0049] AK / SK (Access Key / Secret Key): A pair of cryptographic keys used for authentication. AK is the access key, used to identify the user, and SK is the secret key, used to sign requests to ensure their security and integrity. This key pair enables users to securely access and manage APIs across various services, ensuring that only authorized users can perform specific operations while protecting data transmission and preventing unauthorized access and data tampering. When calling services that require authentication, users use this key pair to prove their identity and obtain the appropriate access rights.
[0050] With the rapid development of artificial intelligence (AI), AIGC (artificial intelligence-generated content) has become a key force driving innovation across multiple industries. AIGC technology demonstrates tremendous potential in areas such as image generation, natural language processing, audio synthesis, and video editing, and is widely used in industries such as media and entertainment, advertising and marketing, education, and healthcare. Through AIGC, businesses and individuals can generate high-quality content at lower costs and higher efficiency, thereby enhancing the user experience. The node-based workflow architecture has rapidly gained widespread recognition within the developer community. This architecture allows users to quickly build complex task flows with simple drag-and-drop operations, significantly lowering the development barrier and improving efficiency, making it an ideal tool for learning and research.
[0051] While these open-source solutions have played an important role in driving technological innovation, they still face significant challenges in production-level integration and resource scheduling. First, many solutions rely on resident GPU resources, which leads to high costs. This is especially true during the workflow building and debugging phases, when GPU resources are often idle, resulting in low resource utilization. Second, the lack of standardized interfaces and service-oriented support makes it complex and difficult to integrate built workflows into existing systems, and they cannot be easily published as services. In addition, most solutions are designed primarily for single-machine configurations and fail to fully consider the deployment requirements of large-scale production environments, resulting in poor performance under high load and high concurrency scenarios. Finally, static resource configuration methods limit the ability to dynamically adjust computing resources according to actual task requirements, which can easily lead to performance bottlenecks and service reliability issues. These issues collectively restrict the widespread adoption and efficient operation of these solutions in enterprise-level applications.
[0052] To address the above issues, this specification provides a method for processing visual tasks. One or more embodiments of this specification also involve a visual task processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments.
[0053] See also Figure 1 , Figure 1 A flowchart of a visual task processing method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0054] Step 102: In response to a call request of a target workflow, the target workflow is obtained, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate a corresponding visual task, and the call request carries task parameters of the visual task.
[0055] A target workflow is a collection of visual task nodes, predefined or dynamically generated in the system, that describe the task flow to be completed. For example, a target workflow could be a combination of visual tasks such as text-to-image, image-to-image, text-to-video, and image-to-video.
[0056] Vision task nodes are the basic units in the target workflow. Each node represents a specific vision task, such as a text-to-image task node for generating images, an image-to-image task for style transfer, a text-to-video task for generating dynamic video content, or an image-to-video task for converting static images into videos. Together, these nodes form the core execution component of the target workflow.
[0057] A call request is an input signal that triggers the target workflow to run. It is usually initiated by a user or an external system and carries the key information required to start the target workflow, including workflow task parameters, input and output parameters of each node, user identity, key, and other parameters.
[0058] In actual applications, the system receives a call request from a user or an external system, and parses the task parameters carried in the request, including the user ID, key, and the input and output parameter information of each node in the workflow. The system obtains the corresponding target workflow from the workflow repository based on the target workflow identification information in the call request. If the target workflow is not found, the system will return an error message to prompt the user to recheck the input. Regarding the method of obtaining the target workflow in the step, one optional method is to directly query the configuration information of the target workflow from the local database, and another optional method is to dynamically load the definition file of the target workflow through a distributed storage system. In addition, the system will also verify the user ID and key in the call request to ensure the legitimacy and security of the call request. The parsed task parameters will be mapped to each visual task node of the target workflow to ensure that each node can correctly obtain the input task parameters and generate the corresponding output results during subsequent execution.
[0059] In the embodiments of this specification, by parsing the task parameters in the call request and dynamically obtaining the target workflow, the system can flexibly adapt to various visual task requirements, improve the flexibility and execution efficiency of task scheduling, and ensure the security and accuracy of the call process.
[0060] For example, on a virtual task platform supporting multimodal visual task processing, a content creator wishes to implement a Wensheng video task through the system. The creator submits a call request, which includes the target workflow's identification information, user ID, key, and task parameters, such as the text description "a flying bird," video length "10 seconds," and resolution "1080p." Upon receiving the call request, the system first verifies the legitimacy of the user ID and key. Then, based on the target workflow's identification information, it loads the corresponding Wensheng video workflow definition file from the distributed storage system. This Wensheng video workflow definition file contains multiple visual task nodes, such as a text parsing node, an image generation node, and a video synthesis node. The system maps the task parameters in the call request to the input parameters of each node, for example, passing the text description to the text parsing node and the video length and resolution to the video synthesis node. This ensures that each node correctly obtains input data and generates the corresponding output results during subsequent task execution.
[0061] Furthermore, in response to a call request of the target workflow, the target workflow is obtained, including: determining the task type and task parameters of the visual task in response to the call request of the target workflow; obtaining an initial workflow based on the task type; and updating the initial workflow based on the task parameters to obtain the target workflow.
[0062] Task types are categories of visual tasks, such as text-to-image (generating images from text) and image-to-image (generating new images based on existing images). These types define the core processing logic and required resources of the workflow.
[0063] Task parameters are the input data or configuration information required during the execution of the target workflow, including but not limited to text descriptions, image resolution, style options, etc. These parameters guide the system on how to perform a specific task type.
[0064] The initial workflow is a basic workflow template preset according to the task type. It contains a set of basic nodes and connection relationships, but has not yet been customized according to specific task parameters.
[0065] The target workflow is the final workflow after customization and update, which includes all necessary nodes, parameter settings, and connection relationships, and can be directly used to perform specific visual tasks.
[0066] In practice, when the system receives a request for a target workflow, it first parses the request to determine the task type and parameters for the vision task. For example, a user might submit a request for a text-based image task, including task parameters such as "Text description: A tranquil lake" and "Image resolution: 1920x1080." The system then retrieves the corresponding initial workflow based on the task type. For text-based image tasks, the system loads a preset text-based image workflow template, which contains basic text parsing nodes, image generation nodes, and their default connections. The system then updates the initial workflow based on the task parameters to obtain the target workflow. For example, in the example above, the system passes the "text description" parameter to the text parsing node and sets the output resolution of the image generation node to 1920x1080. Furthermore, if the task parameters include other custom settings (such as style options), the system will also update the relevant node configurations in the workflow accordingly.
[0067] In the embodiments of this specification, by obtaining the initial workflow according to the task type and updating it based on the task parameters, the system can flexibly adapt to the requirements of different types of visual tasks, ensuring that each task can be efficiently executed according to the user's specific requirements, while reducing the workload of manual configuration and improving work efficiency.
[0068] For example, on a multimodal vision task processing platform, a user wishes to create a task to generate an image with a specific style from a text description. The user initiates a call request, specifies the task type as "Venice Image," and provides the task parameters: "Text Description: A Tranquil Lake," "Image Resolution: 1920x1080," and "Style Option: Oil Painting." Upon receiving the request, the system first determines the task type as "Venice Image" and loads the corresponding initial workflow template, which contains basic text parsing nodes and image generation nodes. Next, the system updates the initial workflow template based on the task parameters: passing the "text description" to the text parsing node, setting the output resolution of the image generation node to 1920x1080, and selecting "Oil Painting Style" as the style option for the generated image. After completing all necessary updates, the system obtains a customized target workflow that can be immediately used to execute the user-specified "Venice Image" task. This approach not only simplifies the task configuration process but also ensures that the task results meet the user's expectations.
[0069] Step 104: Run the target workflow based on the task parameters. When running to the visual task node in the target workflow, call the external visual service to execute the visual task corresponding to the visual task node and obtain the visual result corresponding to the visual task node.
[0070] External vision services are third-party services that operate independently of the target workflow's runtime environment and are designed to handle specific vision tasks, such as text-to-image, image-to-image, text-to-video, and image-to-video. These services are typically deployed on high-performance computing clusters, equipped with powerful computing power to efficiently complete complex vision tasks.
[0071] In actual applications, the system starts the running process of the target workflow based on the task parameters in the call request. When the target workflow runs to a certain visual task node, the system will not rely on local computing power or local services to execute the task of the node, but will complete the specific visual task by calling an external visual service. The system first extracts the input and output parameter information required by the visual task node from the task parameters, and encapsulates it into a request format that complies with the external visual service interface specification. Subsequently, the system calls the interface of the external visual service through the network and sends the encapsulated request to the external visual service. After receiving the request, the external visual service uses its powerful computing resources to perform the visual task and returns the generated visual results to the system. The system receives and parses the returned results and uses them as the output of the visual task node for use by subsequent nodes. For the method of calling the external visual service in the step, one optional method is to call it through the RESTful API interface, and another optional method is to asynchronously trigger the execution of the external visual service through a message queue to ensure the flexibility and scalability of the system.
[0072] In the embodiments of this specification, by calling external visual services instead of relying on local computing power, the system achieves decoupling of rendering and computing, effectively improving resource utilization efficiency and system scalability, while avoiding the impact of local computing power bottlenecks on task execution.
[0073] For example, on a virtual task platform supporting multimodal visual task processing, a user submitted a request to invoke a target workflow containing a Vincent video task. The request included task parameters, such as a text description of "sunrise in a forest," a video length of "15 seconds," and a resolution of "4K." After the system launched the target workflow and reached the Vincent video visual task node, it extracted the required input and output parameters and packaged them into a request that complied with the external visual service interface specification. For example, the system packaged parameters such as the text description, video length, and resolution into a JSON-formatted data packet and invoked an external Vincent video service deployed on a high-performance computing cluster via a RESTful API. Upon receiving the request, the external Vincent video service leveraged its powerful GPU computing power to generate the corresponding video content and returned the generated video file to the system. The system then received and parsed the returned video file, using it as the output of the visual task node for subsequent nodes to consume. Throughout this process, the system efficiently completed complex visual tasks without relying on local computing power.
[0074] Furthermore, calling an external visual service to execute the visual task corresponding to the visual task node includes: parsing the computing power requirement information corresponding to the visual task node through the task distribution service; determining the target service node among the selected service nodes based on the computing power requirement information, wherein the selected service node is used to provide external visual services; calling the target service node to execute the visual task corresponding to the visual task node.
[0075] The Task Distribution Service is a module in the system responsible for analyzing the computing power requirements of vision task nodes and allocating appropriate external vision services to ensure efficient execution of vision tasks. For example, the Task Distribution Service can assess the required computing resources for a text-based image task based on parameters such as resolution and image complexity.
[0076] Computing power requirements are the specific computing resource requirements for each visual task node in the target workflow, including the required GPU memory size, number of CPU cores, and memory capacity. This information guides the system in selecting the appropriate external visual service to perform the task.
[0077] The target service node is an external visual service provider selected from the candidate service nodes, capable of meeting the computing power requirements of the current visual task. For example, the target service node can be a high-performance GPU cluster dedicated to processing tasks with high computing power, such as text-to-image and image-to-video.
[0078] In practice, only tasks or concepts are presented without describing specific technical solutions. When the target workflow reaches a visual task node, the system first uses the task dispatch service to analyze the node's computing power requirements. This computing power requirement is typically determined by the task parameters in the call request and the characteristics of the visual task. For example, in a text-based image task, the system estimates the required GPU memory and computation time based on parameters such as the resolution and style complexity of the generated image. The task dispatch service then selects eligible target service nodes from the candidate service nodes based on this computing power requirement information. For example, the task dispatch service analyzes the computing power requirements of the visual task node, such as the resolution and style complexity of the text-based image task, and estimates the required GPU memory, computing power, and memory. Based on these requirements, the task dispatch service selects eligible target service nodes, such as a GPU node with 16GB of video memory to process a 4K complex style image, to ensure efficient task completion. This screening process may involve a comprehensive evaluation of the availability, load, and performance metrics of multiple candidate service nodes. After determining the target service node, the system sends the encapsulated task request to the target service node by calling the API provided by the target service node. After receiving the request, the target service node uses its powerful computing resources to complete the specific visual task and returns the generated visual results to the system. Regarding the method of calling the target service node in the step, one option is to complete the call using the Python SDK, and another option is to implement the call through the command line tool or RESTful API interface.
[0079] In the embodiments of this specification, the system dynamically analyzes computing power requirements and selects appropriate target service nodes through task distribution services, thereby achieving efficient scheduling and execution of visual tasks, avoiding dependence on local computing power, and improving resource utilization and service scalability.
[0080] For example, on a virtual task platform supporting the Wensheng Image task, a user wishes to have the system generate a high-definition image depicting a "sunset beach." When submitting a call request, the user provides AKSK authentication information, a workflow alias, and task parameters, such as image resolution "1920x1080" and style options "realistic." Upon receiving the request, the system first uses the task distribution service to analyze the computing requirements for the Wensheng Image task, determining, for example, that a GPU with at least 8GB of video memory and a 4-core CPU are required. The task distribution service then selects a target service node from the candidate service nodes that meets the requirements and ultimately selects a high-performance GPU cluster as the target service node. To call the target service node, the user consults the Python Access Guide and completes the configuration using the SDK. This involves replacing the AKSK information, workflow ID, workflow alias, and task parameters in the main.py file. For example, the user replaces the AKSK with their own access key, sets the workflow alias to "Wensheng Image Task," and configures the task parameters to "resolution=1920x1080,style=realistic." The system sends the task request to the target service node through the SDK. After the target service node completes image generation, it returns the result to the system for subsequent rendering and display.
[0081] Furthermore, before determining the target service node among the candidate service nodes based on the computing power demand information, it also includes: determining the task stage of the target workflow based on the task parameters; when the task stage is the production stage, parsing the call request to obtain the user identity identifier, and determining the corresponding proprietary service node based on the user identity identifier, and determining the proprietary service node as the candidate service node; when the task stage is the testing stage, determining the public service node as the candidate service node.
[0082] A user identity is information used to uniquely identify a user in a call request. It typically includes the user's access key (AKSK) or other authentication information. This identity is used to verify user permissions and allocate computing resources based on the user's identity. For example, an enterprise user may have exclusive private computing resources, while ordinary users may use shared public computing resources.
[0083] The task phase is the specific stage of the target workflow execution process, which is divided into the production phase and the test phase. The production phase refers to the stage where the task is officially run and the final results are generated, and generally has high performance and security requirements. The test phase refers to the stage where the task is used for verification or debugging, and is generally cost-sensitive and has relatively low performance requirements.
[0084] Dedicated service nodes are independent computing resources dedicated to specific users or tasks. They are typically deployed in private environments and offer high performance and security. For example, a company's production environment might use dedicated service nodes to ensure data privacy and task stability.
[0085] Public service nodes are shared computing resources open to all users. They are usually deployed in public cloud environments and are low-cost and flexible in expansion, making them suitable for tasks in the testing phase.
[0086] In actual applications, the system first determines the task phase of the target workflow based on the task parameters in the call request. If the task phase is production, the system further parses the call request to obtain the user identity. Based on the user identity, it determines the corresponding private service nodes and selects these private service nodes as candidate service nodes. For example, in a document image task, the system determines whether the user is an enterprise user based on their AKSK information and allocates a dedicated GPU cluster to ensure efficient task execution and data security. If the task phase is testing, the system selects public service nodes as candidate service nodes to reduce resource usage costs. For example, when a user is debugging a document image task, the system allocates shared GPU resources for task execution, avoiding the waste of expensive private computing resources. Regarding the method for determining the candidate service nodes in this step, one option is to directly map the user identity to a predefined list of private service nodes, while another option is to dynamically select public or private service nodes based on the task phase. This approach ensures high performance and security for production tasks while significantly reducing costs during the testing phase.
[0087] In the embodiments of this specification, by distinguishing the task requirements of the production stage and the testing stage and using dedicated service nodes and public service nodes respectively, the system achieves a reasonable allocation of resources, effectively saving the cost of the testing stage while ensuring the quality of production tasks.
[0088] For example, on a virtual task platform supporting text-based imagery, a corporate user submitted a task to generate an image of a "city nightscape." The user provided AKSK information, a workflow alias, and task parameters, such as "4K" image resolution and "sci-fi" style options. Upon receiving the request, the system first parsed the task parameters to determine that the task was in the production phase. Subsequently, based on the user's identity, the system identified the user as an enterprise user and assigned them a dedicated high-performance GPU cluster as a dedicated service node. These dedicated service nodes are located in the enterprise's private cloud environment, ensuring data privacy and performance stability during task execution. These dedicated service nodes in the private cloud environment are selected as candidate service nodes. On the other hand, when a general user submitted a test text-based imagery task, such as generating a low-resolution "cartoon-style" image, the system parsed the task parameters, determined that the task was in the testing phase, and selected shared public service nodes as candidate service nodes. These public service nodes, located in the public cloud environment, have relatively lower performance but are more cost-effective, meeting the needs of the testing phase. Through this mechanism, the system rationally allocates resources at different task stages, which not only ensures the quality of production tasks but also significantly reduces the cost of testing tasks.
[0089] Furthermore, after calling the external visual service to execute the visual task corresponding to the visual task node, it also includes: responding to a node query request containing a task node identifier, determining the task node to be queried based on the task node identifier; querying the query visual results and node output parameters corresponding to the task node to be queried, and generating query result information based on the query visual results and node output parameters; and sending the query result information to the user end.
[0090] A node query request is a request initiated by a user or external system to query the execution results of a specific task node. It typically includes a task node identifier, which uniquely identifies the task node being queried. For example, in a text-based image task, a user can use a node query request to obtain the results and related parameters of a specific image generation node.
[0091] A task node identifier is a unique identifier for each visual task node in the target workflow. It is used to distinguish different task nodes and locate their corresponding execution results. For example, a "style transfer node" in a text-based image task may have a unique identifier of "Node_001".
[0092] The target visual result is the final output generated by the external visual service after executing a specific task node, such as the generated image, video, or other multimedia content. These results are the core target of the user query.
[0093] Node output parameters are the output parameters generated by the task node during execution, including task status information, execution time, resource consumption and other data, which can help users understand the detailed execution status of the task node.
[0094] In actual applications, after the system completes the execution of the task of the external visual service, the user end can initiate a node query request containing the task node identifier. After receiving the query request, the system first determines the task node to be queried based on the task node identifier. For example, in the text image task, the user may want to query the result of a certain "image generation node". The system will locate the node based on the identifier in the request. Then, the system queries the target visual result and node output parameters corresponding to the node. The target visual result may be a high-definition image, and the node output parameters include detailed information such as image resolution, generation time, and the GPU model used. The system integrates the target visual result with the node output parameters to generate query result information. Regarding the method of generating the query result information in the step, one optional method is to encapsulate the target visual result and node output parameters into a JSON format data packet, and another optional method is to directly display the result and parameter information through a visual interface. Finally, the system sends the query result information to the user end for the user to view or further process.
[0095] In the embodiments of this specification, by supporting node-level query functions, the system can provide users with more fine-grained task execution feedback, enhance task transparency and user experience, and facilitate users to analyze and optimize task results.
[0096] For example, on a virtual task platform supporting text-to-image tasks, a user submitted a target workflow consisting of multiple nodes, including a "text parsing node," a "style transfer node," and an "image generation node." After completing the task, the user wished to query the execution results of the "style transfer node," thus initiating a node query request containing the task node identifier "Node_002." Upon receiving the request, the system first located the "style transfer node" based on the identifier and queried its target visual result and node output parameters. The target visual result was a style-transferred image, while the node output parameters included image resolution "1920x1080," execution time "5 seconds," and GPU model "NVIDIA A100." The system aggregated this information into query results, encapsulated them in JSON format, and sent them to the user. After receiving the query results, the user could view detailed execution information by parsing the JSON data packet or intuitively browse the generated images and parameters through a visual interface. This approach not only allows users to quickly obtain the required information but also provides important reference for subsequent task optimization.
[0097] Step 106: Render the visual results corresponding to each visual task node.
[0098] Visual results are the final outputs returned by external visual services after processing, such as generated images, videos, and other multimedia content. These visual results, as the output of each visual task node in the target workflow, need to be presented to the user in an intuitive form through the rendering step or used for subsequent processing.
[0099] In actual applications, after the system receives the visual results returned by the external visual service, it will use local computing resources to render these results. First, the system loads the visual results into the local memory and adjusts the rendering settings according to the display requirements in the task parameters (such as resolution, frame rate, etc.). Then, the system uses hardware resources such as the local graphics processing unit (GPU) and core processor (CPU) to perform rendering operations. For visual results containing complex animations or high-resolution images, the system can use multi-threading technology to fully utilize local computing resources and accelerate the rendering process. For the rendering method in the step, one optional method is to render and display the visual results directly on the local device in real time. Another optional method is to first render the visual results into a file in a specific format (such as an MP4 video file or a JPEG image file), and then display it to the user through a player or other display tools. Throughout the process, the system relies on the powerful local computing power to ensure that the visual results can be presented efficiently and with high quality.
[0100] In the embodiments of this specification, by utilizing local computing resources for rendering, the system can quickly respond to user needs and provide an instant visual experience, while ensuring rendering quality and enhancing the realism and immersion of the user experience.
[0101] For example, on a task platform that supports multimodal visual task processing, a user submits a workflow call request for image-to-video, hoping to convert a series of static images into a dynamic video. After receiving the video data returned by the external visual service, the system begins to use local computing resources for rendering. The system first loads the video data into memory and configures the rendering environment according to parameters such as resolution and frame rate specified by the user. Then, the system starts the local GPU and CPU to work together to render the video data frame by frame. In this process, in order to improve efficiency, the system uses multi-threaded processing technology, which greatly improves the rendering speed. After rendering is completed, the system saves the video file in MP4 format and displays this vivid video converted from static images to the user through the built-in media player. The whole process demonstrates how to use local computing resources to effectively achieve high-quality visual rendering results.
[0102] Furthermore, rendering the visual results corresponding to each visual task node includes: when obtaining the target visual result corresponding to the target task node, scheduling the internal rendering service, and rendering the target visual result in the target area of the display interface, wherein the target task node is any node among the visual task nodes, and the display interface is used to display the target workflow.
[0103] The internal rendering service is responsible for converting visual results into a user-friendly format. It utilizes local computing resources (such as the CPU and GPU) to ensure efficient and high-quality presentation of visual results. The internal rendering service supports a variety of display formats and optimization options to accommodate diverse display requirements.
[0104] The display interface is used to display the results of each visual task node during the execution of the target workflow. This interface is usually interactive and customizable, allowing users to select the results of a specific node or adjust the display method.
[0105] In actual applications, once the system obtains the target visual result corresponding to the target task node, it dispatches an internal rendering service to process the result. First, the system identifies the target task node and its corresponding target visual result. For example, in a workflow containing multiple nodes such as a text-based graph and a graph-based graph, suppose the current focus is on the image generated by the "text-based graph" node. Next, the system invokes the internal rendering service to prepare to render the target visual result in the target area of the display interface. The display interface can be customized based on user needs, such as setting a specific area to display the results of a specific node or using a split-screen mode to display the results of multiple nodes simultaneously. The internal rendering service uses the local graphics processing unit (GPU) and core processor (CPU) to optimize the processing based on the characteristics of the target visual result (such as resolution and color depth) and adapt it to the designated area of the display interface. Regarding the method of dispatching the internal rendering service in the step, one option is to directly utilize local hardware resources for real-time rendering and immediate display of the result. Another option is to first render the visual result into a file in a specific format (such as JPEG, MP4, etc.) and then load it into the display interface through a player or other display tool.
[0106] In the embodiments of this specification, by scheduling internal rendering services and rendering visual results in the target area of the display interface, the system not only achieves efficient visual result display, but also enhances the realism and immersion of the user experience, allowing users to intuitively evaluate the output effects of each visual task node.
[0107] For example, on a multimodal visual task processing platform, a user completes a workflow involving nodes such as "Text Image" and "Image Image." The task for the "Text Image" node is to generate a landscape painting based on a provided text description. After the system receives the image generated by this node from an external visual service, it dispatches the internal rendering service. The system first identifies the "Text Image" node as the target task node and obtains its corresponding target visual result—a high-definition landscape painting. The system then calls the internal rendering service to prepare to display this image in the left area of the display interface. The display interface is designed with a two-column layout, with the left side dedicated to displaying the results of the "Text Image" node and the right side reserved for other nodes. The internal rendering service utilizes the local GPU to optimize the image, ensuring high-quality presentation. After processing, the image is smoothly displayed in the designated area of the display interface. The user can directly view the generated landscape painting and, based on the display quality, decide whether to adjust parameters or resubmit the task. This approach allows users to intuitively see the task results without additional interaction, significantly improving the user experience.
[0108] Furthermore, before obtaining the target workflow in response to the call request of the target workflow, it also includes: in response to the workflow creation request, obtaining a preset visual task node template, and calling the internal rendering service to render the visual task node template in the workflow editing interface; in response to the node editing operation, adding at least one visual task node to the visual task node template, and generating a task node identifier corresponding to each visual task node, wherein the task node identifier is used for the front-end to query at least one visual task node; in response to the drag operation, determining the connection relationship between each visual task node and generating a workflow; in response to the call request of the target workflow, obtaining the target workflow, including: in response to the call request of the target workflow, obtaining the target workflow from the workflow.
[0109] A workflow creation request is a request initiated by a user or external system to generate a new target workflow. It typically contains requirements and configuration information for visual task nodes. For example, a user may wish to create a workflow containing nodes such as a document-to-graph or a graph-to-graph to complete a specific task.
[0110] Visual task node templates are predefined, general-purpose task node templates used to quickly build specific task nodes within a workflow. These templates typically include default task parameters and logical structures, which users can customize based on their needs. For example, a text image node template might include components such as a text input box and a resolution selector.
[0111] The workflow editing interface allows users to create and edit target workflows. It supports adding, editing, and connecting visual task nodes in a visual manner. This interface typically features a drag-and-drop function, making it easy for users to define the connections between nodes.
[0112] In actual application, the system first responds to the user's request to create a workflow, obtains the preset visual task node templates, and calls the internal rendering service to render these templates in the workflow editing interface. For example, if the user wants to create a Wenshengtu workflow, the system will load the Wenshengtu node template and display it in the editing interface. Subsequently, the user can add at least one visual task node to the template through node editing operations, such as adding a "style transfer" node or an "image enhancement" node. Each newly added node will be assigned a unique target node identifier for subsequent query and management. Then, the user can determine the connection relationship between each visual task node by dragging and dropping. For example, connect the output of the Wenshengtu node to the input of the style transfer node to form a complete task flow. After completing the node connection, the system will generate the target workflow according to the user's configuration. Regarding the method of generating workflows in the steps, one optional method is to directly complete the workflow design through interface interaction, and another optional method is to achieve automatic generation by importing a predefined workflow configuration file.
[0113] After generating the target workflow, users can further set the input and output parameters of the workflow. For example, in a document image task, users can add parameters such as "text description" and "image resolution" in the input parameter settings, and specify the format of the returned content (such as JPEG image) in the output parameter settings. In addition, users can also specify an alias for the workflow to facilitate subsequent calls. When publishing a workflow, the system will generate interface version management information to ensure that each update can be seamlessly replaced through an alias. For example, the user specifies the alias "t2i_0821" for the document image workflow and uses this alias to simplify the interface call process when calling.
[0114] In the embodiments of this specification, by providing a visual way to create workflows, the system significantly reduces the difficulty for users to build complex task processes, while improving the flexibility and scalability of workflows through templated design and parameterized configuration.
[0115] For example, on a virtual task platform, a user wishes to create a workflow containing text-based graphs and style transfer nodes. The user first initiates a workflow creation request. The system responds by loading a preset text-based graph node template and rendering it in the workflow editing interface. The user then adds a style transfer node through node editing and generates unique identifiers, "Node_001" and "Node_002," for each node. Subsequently, the user connects the output of the text-based graph node to the input of the style transfer node by dragging and dropping, completing the logical connection between the nodes. Next, the user adds "text description" and "image resolution" parameters to the workflow's inbound parameters and specifies the return value as "stylized image" in the outbound parameters. After completing the configuration, the user assigns the workflow the alias "text_to_style_image" and clicks the Publish button to generate the interface version. After publishing, the user obtains AKSK authentication information by creating an application and completes the interface call configuration using the Python SDK. Finally, the user automatically generates stylized images from text descriptions by calling the workflow with the alias "text_to_style_image." This approach not only simplifies the workflow creation process, but also improves the efficiency and flexibility of task execution.
[0116] Furthermore, the node editing operation includes a node adding operation and a parameter editing operation; in response to the node editing operation, at least one new visual task node is added to the visual task node template, including: in response to the node adding operation, at least one new initial node is added to the visual task node template; in response to the editing operation on the target initial node, the input parameters and output parameters of the target initial node are set to obtain a first visual task node, wherein the target initial node is any node among the multiple initial nodes, the input parameters are used to define the data required for the execution of the first visual task node, and the output parameters are used to define the data required for the execution of the second visual task node, and the second visual task node is a downstream node of the first visual task node.
[0117] Node editing is the process of adding, deleting, or modifying parameters of a vision task node in the workflow editing interface. This allows users to customize the functionality and parameter configuration of each vision task node based on their specific needs.
[0118] The node add operation is a type of node editing operation, specifically used to add new visual task nodes to a workflow. Through this operation, users can quickly import the required task nodes and further configure them in detail.
[0119] Input parameters are external data or configuration information required for the execution of the target vision task node, such as text description, image resolution, etc. These parameters define the specific content and requirements of the node execution.
[0120] Output parameters are the data or results generated after executing the target vision task node, and are typically used as input parameters for downstream nodes. For example, in a workflow that includes a Style Transfer node and a Style Image node, the output (generated image) of the Style Image node can be used as input for the Style Transfer node.
[0121] In practice, the system first responds to user node editing operations, particularly node addition operations, by adding at least one initial node to the visual task node template. For example, if a user wishes to create a text-to-image task, the system provides a series of preset node templates, such as "Text Parsing" and "Image Generation." The user can add these node templates to the workflow by clicking the corresponding button. Next, for each added initial node, the user can perform detailed parameter editing. Taking the "Text Parsing" node as an example, the user needs to set its input parameters, such as "Text Description," which defines the specific text content to be processed when the node executes. After completing the input parameter setting, the user also needs to define the node's output parameters—the format or content of the generated intermediate data. This is typically used as an input parameter for downstream nodes, such as the "Image Generation" node. For example, the "Text Parsing" node might generate a structured text description object as one of the input parameters for the "Image Generation" node, so that the latter can generate an image based on the parsed text content. Regarding the method of setting input parameters and output parameters in the steps, one optional method is to directly fill in or select parameter values in the node editing interface; another optional method is to import a pre-defined parameter configuration file to simplify the complex parameter setting process.
[0122] In the embodiments of this specification, by supporting node addition and parameter editing operations, the system enables users to flexibly build complex workflows, while ensuring that the input and output of each node are clear and unambiguous, facilitating the connection and data transfer between subsequent nodes, and improving the maintainability and scalability of the entire workflow.
[0123] For example, on a multimodal vision task processing platform, a user wants to create a workflow that generates an image with a specific style from a text description. The user first initiates a node addition operation, adding three initial nodes: "Text Parsing," "Style Transfer," and "Image Generation" in the workflow editing interface. Next, the user edits the "Text Parsing" node, setting its input parameter to "Text Description: A Tranquil Lake," which defines the specific text content to be processed when the node executes. The user also sets the node's output parameter to "Parsed Text Description Object," which will serve as one of the input parameters for the downstream "Image Generation" node. Subsequently, the user edits the "Image Generation" node, setting its input parameters to include the aforementioned "Parsed Text Description Object" and "Image Resolution: 1920x1080." After completing the parameter settings for all nodes, the user connects them by dragging and dropping, forming a complete text-to-image workflow. Finally, the user publishes and runs the workflow, successfully generating an image that meets the expectations. This approach not only allows users to intuitively construct complex task flows but also significantly improves the success rate and efficiency of task execution.
[0124] One embodiment of the present specification implements a method of obtaining a target workflow in response to a call request of the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate the corresponding visual task, and the call request carries the task parameters of the visual task; running the target workflow based on the task parameters, and when running to the visual task node in the target workflow, calling the external visual service to execute the visual task corresponding to the visual task node, and obtaining the visual result corresponding to the visual task node; and rendering the visual result corresponding to each visual task node. Flexible arrangement and execution of visual tasks are achieved through workflow scenario services. When the call request triggers the target workflow to run, the visual task node relies on the external visual service to complete the specific task, avoiding direct dependence on local computing power, thereby achieving decoupling of rendering and computing power, effectively improving resource utilization efficiency and system scalability, while ensuring that each visual task result can be rendered and displayed in real time to meet dynamic needs in multiple scenarios.
[0125] The following combined Figure 2 , taking the application of the method provided in this specification in the visual task workflow management platform as an example, the visual task processing method is further explained. Figure 2 This is an architectural diagram of a visual task workflow management platform provided in one embodiment of this specification. Figure 2The paper demonstrates the architecture of a visual task workflow management platform, which breaks down the rendering service, image generation service, and task query service into sub-categories. The rendering module is responsible for exposing native image generation parameters to users, ensuring that the page display and user experience are completely consistent with the locally deployed version in existing technologies, and supports the installation of the latest plug-ins. The image generation service includes a task distribution module and a task execution module. The task distribution module parses input parameters and routes requests to servers of appropriate specifications for execution. The task execution module receives requests and forwards them to the native image generation service, while maintaining the lifecycle of the task. The computing power management module dynamically adjusts computing power supply according to real-time load, cooperates with the task execution module to complete task processing, and finally returns the task execution results through the task query service, realizing low-threshold, maintenance-free, and highly available AI image generation capabilities.
[0126] See also Figure 3 , Figure 3 This is a flowchart of the processing process of a visual task processing method provided by one embodiment of this specification. It includes two stages: workflow deployment and code call. In the workflow deployment stage, users need to activate the platform, debug the workflow, publish the workflow as an excuse, specify parameters / alias, create an application, and obtain the AKSK step. In the code call stage, in the code interface call process, it is necessary to connect to the SDK, call the excuse, poll the results until the final state step. Finally, in the platform GPU computing power service, there are user exclusive resource pools and shared resource pools. The exclusive resource pool is called when the user performs production tasks, and the shared resource pool is called when the user performs testing tasks.
[0127] Figure 3 The paper demonstrates the process of a visual task processing method. First, a custom plug-in is used for mainstream image processing platforms, and the basic dynamic replacement function is implemented by dynamically importing native modules. Then, this capability is used to replace single-user functions with functions that support multiple users, and user context is stored in centralized storage to achieve distributed multi-user management. Finally, the native internal call to the image processing service is replaced with an HTTP interface call to the remote service. After receiving the image processing input, the task distribution module searches for the most matching image processing instance based on user configuration, input parameter details, current limiting configuration, and the real-time load of each image processing cluster on the cloud. After completing the task delivery, it returns the image processing task ID to the rendering service. The task execution module receives and calls the image processing service, maintains the task lifecycle until the execution succeeds or fails, and stores the results in the cloud. Finally, the task rendering module obtains the latest task image processing progress / results through the task query service, achieving efficient task management and resource scheduling, and providing low-threshold, maintenance-free, and highly available AI image processing capabilities. This process optimizes the balance between cost and performance, supporting the customized configuration and deep integration needs of enterprise customers.
[0128] See also Figure 4, Figure 4 This is a schematic diagram of the front-end interface of the first visual task workflow management platform provided by an embodiment of this specification. In the initial display interface of the visual task workflow management platform, a "New Workflow" control is included. After the user triggers it, a workflow setting page is displayed on the right, including the workflow name and the corresponding input box for the user to enter, the workflow description and the corresponding input box for the user to enter, and a selection control for selecting a workflow template. Specifically, the selection control includes a blank workflow, a text-to-image workflow, a picture-to-image workflow, a text-to-video workflow, and a picture-to-video workflow. Figure 3 The Chinese raw image workflow has been selected by the user. Click the confirmation control below to initiate further operations. Click the cancel control to delete all settings in the workflow settings page.
[0129] See also Figure 5 , Figure 5 This is a schematic diagram of the front-end interface of the second visual task workflow management platform provided by an embodiment of this specification. Figure 4 Triggered after the user clicks the confirmation control Figure 5 The interface contains multiple visual task node components (such as components 1 to 7). The "Forward Prompt Word" and "Reverse Prompt Word" nodes are used to input key parameters for generating images from text. The top toolbar provides function buttons for "Import," "Export," "Generate History," "Publish," and "Run Workflow" to support quick user operations. The left side of the interface displays the workflow name "Simple Text-Image Workflow 001," and the right side presents the workflow topology structure in the form of node connections. Users can adjust the logical relationship by dragging and dropping components, and the input and output parameters between components are dynamically bound through connections.
[0130] See also Figure 6 , Figure 6 This is a schematic diagram of the front-end interface of the third visual task workflow management platform provided by an embodiment of this specification. Figure 5 Triggered after the user clicks the Run Workflow control Figure 6 Interface, a new execution progress bar is added at the top of the interface to display the overall operation status of the workflow in real time. Image A is displayed in component 7 at the end of the workflow, and the historical generation result area displays the output content such as "Historical Image B" and "Historical Image C" in thumbnail form, supporting click preview and comparison. In the original component layout, the "Forward Prompt Word" and "Reverse Prompt Word" nodes retain the parameter configuration, the run button status is changed to "Executing", and the "Generate History" button is activated, and past task records can be viewed. The middle canvas area dynamically renders the execution status of each node, the completed nodes are marked with a green border, and the executing nodes display a pulse animation.
[0131] See also Figure 7 , Figure 7This is a schematic diagram of the front-end interface of the fourth visual task workflow management platform provided by an embodiment of this specification. Figure 6 Publish controls (such as Node 3) are displayed, and input parameter configurations are managed in the "Selected" area. Add new parameters using the "Add" button or remove redundant parameters using the "×" button. The interface includes a "Version Description" input box for entering a release description. The "Workflow Parameter Settings" area at the bottom supports dynamic expansion. Users can define interface-exposed parameters using the "+Add" button, complete the release process using "Submit Release," or return to the editing state by clicking "Cancel." Bindings for mapping input nodes to workflow nodes are selected from the drop-down menu.
[0132] See also Figure 8 , Figure 8 This is a schematic diagram of the front-end interface of the fifth visual task workflow management platform provided in an embodiment of this specification. The interface displays the configured input and output parameters in a tabular form, including two columns: "Parameter Name" and "Components and Fields in the Corresponding Workflow". Users can modify the parameter alias through "Rename", adjust the parameter order through "Move Up" and "Move Down", or achieve custom sorting by dragging and dropping. Parameter configuration supports batch operations to ensure that the interface fields are aligned with the data structure of the enterprise's existing system, reducing integration complexity.
[0133] See also Figure 9 , Figure 9 This is a schematic diagram of the front-end interface of the sixth visual task workflow management platform provided by an embodiment of this specification. The node topology structure (such as node 3) is displayed on the left side of the interface, and the "Workflow Output Parameter Settings" area on the right allows the user to select the output field (such as "Output Parameter 1") through the "+ Add" button, and supports previewing the metadata of the generated results. The input parameter configuration area displays a list of bound parameters (such as "Input Parameter 1" and "Input Parameter 2"), and the user can further adjust the parameter mapping relationship. After completing the configuration, click "Submit and Release" to generate a workflow interface service with a version number, or click "Cancel" to return to the editing state. The version description input box supports multiple lines of text, which is convenient for recording task semantics and change instructions.
[0134] See also Figure 10 , Figure 10 This is a schematic diagram of the front-end interface of the seventh visual task workflow management platform provided by an embodiment of this specification. Figure 9 After clicking Add, a new workflow output parameter setting called Output 1 is added to the display area of the published workflow, allowing users to click and edit it.
[0135] See also Figure 11 , Figure 11This is a schematic diagram of the front-end interface of the eighth visual task workflow management platform provided by an embodiment of this specification. The bottom function area contains the "Service Management", "Version Management" and "Alias Management" tabs. The current interface focuses on alias management, and lists the configured alias information in a table, including the alias name, description, corresponding version and "Edit" and "Delete" operation buttons. Users can complete the configuration update through "Submit Publish" or click "Cancel" to abandon the modification. Alias configuration supports dynamic binding of workflow instances of different versions, which makes it convenient for enterprise-level customers to call the interface through fixed aliases and avoid frequent code adjustments due to version iterations.
[0136] See also Figure 12 , Figure 12 This is a schematic diagram of the front-end interface of a visual task application management platform provided in an embodiment of this specification. The title of the interface is "Visual Task Application Management Platform", and the user enters the configuration page through the "Create Application" button. The page contains core configuration items such as the "Application Name" input box, the "Application Type" radio button group (optional "SDK Access" or "API Application"), and the "Application Description" input box. The application type defaults to "API Application", and the user can switch to "SDK Access" mode to adapt to different integration scenarios. After completing the configuration, click the "Confirm" button to generate an application instance and assign a unique AK / SK key pair, or click the "Delete" button to clear the current form. This interface provides enterprises with a unified access portal to ensure that application permissions, billing, and log management are seamlessly integrated with the company's existing systems.
[0137] through Figure 4-Figure 2 After editing the workflow, you can run the workflow to view the results, including visual task results and parameter results. The parameter results are as follows:
[0138] {
[0139] "status": 10,
[0140] "err_code": null,
[0141] "err_message": null,
[0142] "sub_err_code": null,
[0143] "sub_err_message": null,
[0144] "api_invoke_id": "i_66c5a89c9b80590025596496",
[0145] "data": {
[0146] "task_id": "01j5t1n78pcvmjsdbsp5jd52hq",
[0147] "images": [
[0148] "http: / / xxx"
[0149] ],
[0150] "info": {},
[0151] "parameters": null,
[0152] "status": "succeeded",
[0153] "imgs_bytes": null
[0154] }
[0155] }
[0156] Among them, "status": 10: represents the gateway status code. A value of 10 means that the request was successfully processed.
[0157] "err_code": null: Gateway error code. This field is empty when the request is successful, indicating that no gateway-level error occurs.
[0158] "err_message": null: Gateway error details. This field is ignored if successful and is empty.
[0159] "sub_err_code": null: Service error code. This field is ignored when the request is successful, indicating that no service-level error occurs.
[0160] "sub_err_message": null: Service error details, ignored if successful, this is empty.
[0161] "api_invoke_id": "i_66c5a89c9b80590025596496": A unique identifier generated by the system, used to identify this API call and to facilitate subsequent queries or troubleshooting.
[0162] "data": { ...}: Contains the specific result data object returned by the service.
[0163] "task_id": "01j5t1n78pcvmjsdbsp5jd52hq": The unique identifier of the image generation task, used to track the execution status of a specific task.
[0164] "images": ["http: / / xxx"]: The default output image list contains the accessible addresses of the generated images.
[0165] "info": {}: Additional information object, which may be empty or contain relevant metadata depending on the actual situation.
[0166] "parameters": null: Task parameters, or null if not specified.
[0167] "status": "succeeded": Task status. A value of "succeeded" indicates that the task was successfully executed.
[0168] "imgs_bytes": null: Image byte stream. If the image file is not returned directly, this field is empty.
[0169] Corresponding to the above method embodiment, this specification also provides a visual task processing device embodiment, Figure 13 This is a schematic diagram of the structure of a visual task processing device provided by an embodiment of this specification. Figure 13 As shown, the device includes:
[0170] An acquisition module 1302 is configured to acquire a target workflow in response to a call request for the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate a corresponding visual task, and the call request carries task parameters of the visual task;
[0171] The calling module 1304 is configured to run the target workflow based on the task parameters, and when running to the visual task node in the target workflow, call the external visual service to execute the visual task corresponding to the visual task node, and obtain the visual result corresponding to the visual task node;
[0172] The rendering module 1306 is configured to render the visual result corresponding to each visual task node.
[0173] Optionally, the acquisition module 1302 is further configured to determine the task type and task parameters of the visual task in response to a call request of the target workflow; acquire the initial workflow based on the task type; and update the initial workflow based on the task parameters to obtain the target workflow.
[0174] Optionally, the calling module 1304 is further configured to parse the computing power requirement information corresponding to the visual task node through the task distribution service; based on the computing power requirement information, determine the target service node among the candidate service nodes, wherein the candidate service node is used to provide external visual services; and call the target service node to execute the visual task corresponding to the visual task node.
[0175] Optionally, the calling module 1304 is further configured to determine the task stage of the target workflow based on the task parameters; when the task stage is the production stage, the calling request is parsed to obtain the user identity identifier, and the corresponding proprietary service node is determined based on the user identity identifier, and the proprietary service node is determined as the service node to be selected; when the task stage is the testing stage, the public service node is determined as the service node to be selected.
[0176] Optionally, the visual task processing device also includes a query module, which is configured to respond to a node query request containing a task node identifier, determine the task node to be queried based on the task node identifier; query the query visual results and node output parameters corresponding to the task node to be queried, and generate query result information based on the query visual results and node output parameters; and send the query result information to the user end.
[0177] Optionally, the rendering module 1306 is further configured to schedule an internal rendering service to render the target visual result in a target area of the display interface when the target visual result corresponding to the target task node is obtained, wherein the target task node is any node among the visual task nodes, and the display interface is used to display the target workflow.
[0178] Optionally, the visual task processing device also includes a creation module, which is configured to respond to a workflow creation request, obtain a preset visual task node template, and call an internal rendering service to render the visual task node template in the workflow editing interface; respond to a node editing operation, add at least one visual task node to the visual task node template, and generate a task node identifier corresponding to each visual task node; respond to a drag operation, determine the connection relationship between each visual task node, and generate a workflow.
[0179] Optionally, the node editing operation includes a node adding operation and a parameter editing operation; accordingly, the creation module is further configured to add at least one initial node to the visual task node template in response to the node adding operation;
[0180] In response to an editing operation on a target initial node, the input parameters and output parameters of the target initial node are set to obtain a first visual task node, wherein the target initial node is any one of multiple initial nodes, the input parameters are used to define the data required for the execution of the first visual task node, and the output parameters are used to define the data required for the execution of the second visual task node, and the second visual task node is a downstream node of the first visual task node.
[0181] Applied to the visual task processing device, first, the acquisition module 1302 responds to the call request of the target workflow to obtain the target workflow containing at least one visual task node. The call request carries the task parameters of the visual task to provide necessary input for subsequent steps. Then, the calling module 1304 runs the target workflow based on these task parameters. When encountering a visual task node, it completes the specific visual task by calling an external visual service, avoiding direct dependence on local computing power, and realizing the decoupling of computing power and rendering power, thereby effectively improving resource utilization efficiency and system scalability. Finally, the rendering module 1306 is responsible for real-time rendering of the visual results corresponding to each visual task node, ensuring that these results can be quickly displayed to the user to meet dynamic needs in multiple scenarios. In this way, the device not only realizes the flexible configuration and efficient execution of the visual task processing process, but also enhances the adaptability and scalability of the system.
[0182] The above is a schematic diagram of a visual task processing device according to this embodiment. It should be noted that the technical solution of this visual task processing device and the technical solution of the visual task processing method described above are based on the same concept. For details not described in detail in the technical solution of the visual task processing device, please refer to the description of the technical solution of the visual task processing method described above.
[0183] Figure 14 14 shows a block diagram of a computing device 1400 according to one embodiment of the present disclosure. Components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.
[0184] Computing device 1400 also includes an access device 1440 that enables computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1440 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0185] In one embodiment of the present specification, the above components of the computing device 1400 and Figure 14 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 14 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0186] Computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1400 can also be a mobile or stationary server.
[0187] The processor 1420 is configured to execute the following computer program / instructions, which implement the steps of the above-mentioned visual task processing method when executed by the processor.
[0188] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computing device embodiment is generally similar to the visual task processing method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the visual task processing method embodiment.
[0189] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above-mentioned visual task processing method when executed by a processor.
[0190] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences between the other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the visual task processing method embodiment, so its description is relatively simple. For relevant parts, refer to the description of the visual task processing method embodiment.
[0191] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned visual task processing method when executed by a processor.
[0192] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the aforementioned visual task processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned visual task processing method.
[0193] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0194] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0195] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0196] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0197] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for processing a visual task, comprising: In response to a call request for a target workflow, obtaining the target workflow, wherein the target workflow includes at least one visual task node, the visual task node is used to indicate a corresponding visual task, the call request carries task parameters of the visual task, and the target workflow includes a combined flow of text-to-graph, graph-to-graph, text-to-video, and graph-to-video visual tasks; Running the target workflow based on the task parameters, and when running to the visual task node in the target workflow, calling the external visual service to execute the visual task corresponding to the visual task node, and obtaining the visual result corresponding to the visual task node; The internal rendering service is scheduled to render the visual result corresponding to the visual task node, wherein the internal rendering service is used to render the visual result in a target area of the display interface.
2. The method according to claim 1, wherein the step of obtaining the target workflow in response to the call request of the target workflow comprises: In response to a call request of a target workflow, determining a task type and task parameters of a visual task; Obtaining an initial workflow based on the task type; The initial workflow is updated based on the task parameters to obtain a target workflow.
3. The method according to claim 1, wherein calling an external visual service to execute the visual task corresponding to the visual task node comprises: Analyze the computing power demand information corresponding to the visual task node through the task distribution service; Determining a target service node from among candidate service nodes based on the computing power requirement information, wherein the candidate service node is used to provide external visual services; The target service node is called to execute the visual task corresponding to the visual task node.
4. The method according to claim 3, before determining a target service node from among candidate service nodes based on the computing power requirement information, further comprising: Determining a task phase of the target workflow based on task parameters; In the case where the task stage is the production stage, parsing the call request to obtain a user identity identifier, determining a corresponding proprietary service node based on the user identity identifier, and determining the proprietary service node as a candidate service node; When the task phase is a testing phase, a public service node is determined as a candidate service node.
5. The method according to claim 1, after calling the external visual service to execute the visual task corresponding to the visual task node, further comprising: In response to a node query request including a task node identifier, determining a task node to be queried based on the task node identifier; Querying the query visual result and node output parameters corresponding to the task node to be queried, and generating query result information based on the query visual result and the node output parameters; The query result information is sent to the user terminal.
6. The method according to claim 1, wherein rendering the visual result corresponding to the visual task node comprises: When the target visual result corresponding to the target task node is obtained, the internal rendering service is scheduled to render the target visual result in the target area of the display interface, wherein the target task node is any node among the visual task nodes, and the display interface is used to display the target workflow.
7. The method according to any one of claims 1 to 6, before obtaining the target workflow in response to the call request of the target workflow, further comprising: In response to a workflow creation request, obtaining a preset visual task node template, and calling an internal rendering service to render the visual task node template in a workflow editing interface; In response to the node editing operation, at least one visual task node is added to the visual task node template, and a task node identifier corresponding to each visual task node is generated, wherein the task node identifier is used by the front end to query the at least one visual task node; In response to the drag operation, determining the connection relationship between the visual task nodes and generating a workflow; The step of obtaining the target workflow in response to the call request of the target workflow includes: In response to a call request of a target workflow, the target workflow is acquired from the workflow.
8. The method according to claim 7, wherein the node editing operation includes a node adding operation and a parameter editing operation; In response to the node editing operation, adding at least one visual task node to the visual task node template includes: In response to the node adding operation, adding at least one initial node in the visual task node template; In response to a parameter editing operation on a target initial node, the input parameters and output parameters of the target initial node are set to obtain a first visual task node, wherein the target initial node is any one of the at least one initial nodes, the input parameters are used to define the data required for the execution of the first visual task node, and the output parameters are used to define the data required for the execution of the second visual task node, and the second visual task node is a downstream node of the first visual task node.
9. A visual task processing device comprising: an acquisition module configured to, in response to a call request for a target workflow, acquire the target workflow, wherein the target workflow comprises at least one visual task node, the visual task node being used to indicate a corresponding visual task, the call request carrying task parameters of the visual task, and the target workflow comprising a combined flow of text-to-graph, graph-to-graph, text-to-video, and graph-to-video visual tasks; a calling module configured to run the target workflow based on the task parameters, and when running to the visual task node in the target workflow, call the external visual service to execute the visual task corresponding to the visual task node, and obtain the visual result corresponding to the visual task node; The rendering module is configured to schedule an internal rendering service to render the visual result corresponding to the visual task node, wherein the internal rendering service is used to render the visual result in a target area of the display interface.
10. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the visual task processing method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the visual task processing method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the visual task processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Calculation power scheduling method and device of calculation power network
CN116095179A
Hybrid computing power processing method and device based on trusted execution environment
CN116910745A