Large model process arrangement system, method, equipment and medium

By flexibly combining and automating the scheduling of modular nodes, the complexity and resource allocation issues in large-scale process orchestration are resolved, enabling rapid and visualized process management and a closed-loop end-to-end system, thereby improving task success rate and cross-team collaboration efficiency.

CN120929202APending Publication Date: 2025-11-11SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802077.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for large-scale model process orchestration and tool management suffer from problems such as long process setup cycles, inflexible resource allocation, reliance on professional technical personnel, difficulty in fault location, and difficulty in quickly integrating new tools or expanding business scenarios, which limit the development of large-scale models in the industry.

Method used

It adopts flexible combination and automated scheduling of modular nodes, and realizes process-oriented and standardized management through toolchain development modules, management modules, pre-launch testing modules, task execution modules and status monitoring modules. It supports dynamic parameter configuration and resource pre-allocation, and provides a visual drag-and-drop interface and full-link log tracking function.

Benefits of technology

Significantly shortens process setup time, improves task success rate and efficiency, reduces operation and maintenance costs, supports non-technical personnel to quickly build complex processes, achieves full-link closed-loop management and cross-team collaboration, and meets data security compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929202A_ABST
    Figure CN120929202A_ABST
Patent Text Reader

Abstract

The invention provides a large model process arrangement system, method and device and a medium, and belongs to the technical field of large models. The system comprises a tool chain development module, a tool chain management module, a node management module, a state monitoring module, a pre-online test module, a tool chain task operation module and a task management module. The method comprises the steps of tool chain template generation, node parameter configuration, resource pre-allocation, sandbox environment testing, task dynamic scheduling, log tracing, result linkage and the like. According to the method, the operation threshold is reduced through the visual dragging interface, dynamic parameter binding and GPU resource intelligent allocation are supported, and the process execution efficiency and stability are optimized. Meanwhile, functions of full-link log monitoring, error tracing and automatic result pushing are provided, and cross-module collaboration and service closed loop are realized. According to the system, the construction and execution efficiency of a large model task process can be remarkably improved, the flexibility and expandability are enhanced, the operation and maintenance cost is reduced, the data security and auditing performance are guaranteed, and technical support is provided for efficient application of a large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of large model technology, and more specifically relates to a large model process orchestration system, method, device and medium. Background Technology

[0002] In today's era of rapid technological advancement, artificial intelligence is revolutionizing various industries at an unprecedented pace. Industry-wide big data models play a crucial role throughout the entire lifecycle, from data preprocessing and model training to inference deployment, providing a powerful driving force for the intelligent upgrading of numerous fields. However, despite the immense application value of industry-wide big data models, existing technologies face significant bottlenecks in process orchestration and tool management, severely hindering their further development and application.

[0003] Currently, traditional toolchain development heavily relies on manually written scripts or configured parameters, lacking a standardized workflow template. This approach suffers from poor data format compatibility between different tools, requiring repeated debugging, resulting in long workflow setup cycles and significant issues of repetitive development.

[0004] In terms of computing resource allocation, existing technologies mainly rely on static configuration, which cannot be dynamically adjusted according to the actual task requirements. This leads to a situation where, during actual operation, the allocation of computing resources such as GPUs depends on static configuration and cannot be dynamically adjusted according to task needs, easily resulting in resource idleness or over-provisioning. Furthermore, the lack of fault tolerance mechanisms (such as retries and dependency rollbacks) during task execution results in high task failure rates and poor stability.

[0005] Existing toolchain management and process design heavily rely on technical professionals, making it difficult for non-technical personnel to participate in complex process design. Furthermore, the dispersed storage of toolchain logs makes error tracing extremely difficult and time-consuming when errors occur. In addition, redundant storage of intermediate data further increases the operational burden, resulting in persistently high maintenance costs.

[0006] Existing systems struggle to quickly integrate new tools or expand business scenarios; parameter configurations are rigid and unable to adapt to dynamic data scales. Cross-module collaboration capabilities are weak, and process results must be manually pushed to downstream systems, making it difficult to achieve end-to-end closed-loop management.

[0007] In summary, the numerous problems existing in current technologies regarding process orchestration and tool management have become key factors restricting the further development and application of large-scale industry models. Therefore, developing a technical solution that can effectively address these problems is of significant practical importance. Summary of the Invention

[0008] To address the above problems, the present invention aims to provide a large-scale model process orchestration system, method, device, and medium that achieves process-oriented and standardized management of large-scale model full lifecycle tasks through flexible combination and automated scheduling of modular nodes.

[0009] To achieve the above objectives, the present invention employs the following technical solution: In a first aspect, embodiments of this application provide a large-scale model workflow orchestration system, including: The toolchain development module is used to define modular nodes, including input nodes, tool nodes, and output nodes, configure modular nodes, and generate toolchain templates that include node execution order, resource quotas, and error handling strategies. The toolchain management module is used to define toolchain processes by dragging and dropping modular nodes, automatically verify the compatibility of data formats and the legality of topology between nodes, and configure toolchain parameters and pre-allocate resources. The node management module is used to classify modular nodes by function and control access permissions, and to perform node online and offline operations. The pre-launch testing module is used to automatically create a sandbox environment before the toolchain is released, simulate real resource allocation and data flow to perform full-process testing, and generate test reports; The toolchain task execution module is used to generate task instances and schedule their execution based on the toolchain template. The status monitoring and log management module is used to monitor and display the operation information of resources and modular nodes in real time during task execution, aggregate and store node logs and provide full-text search and link tracing functions; The execution result processing module is used to persist the execution results generated by the output node to a specified storage and generate an access link, and trigger downstream processes based on the execution results.

[0010] In one optional implementation, the toolchain development module includes: The node type definition unit is used to define input nodes, tool nodes, and output nodes, and to configure node parameters, input / output interfaces, and connection relationships through a visual interface. The input nodes are used for data access, data preprocessing, and format conversion; the tool nodes are used for model training, inference, evaluation, and deployment; and the output nodes are used for storing execution results, pushing data, and visualization. The toolchain template generation unit is used to generate toolchain templates based on the combination relationship of modular nodes, and encapsulates multiple tool nodes into composite nodes through nested sub-processes.

[0011] In an optional implementation, the toolchain management module includes: The toolchain creation unit is used to define the toolchain process by dragging and dropping modular nodes onto the canvas and configuring connection logic, and automatically verifying the data format compatibility and topology legality between modular nodes; The parameter configuration unit is used to configure the static parameters of modular nodes through a visual form, and to configure the dynamic parameters of modular nodes based on the running results of upstream nodes. The resource pre-allocation unit is used to generate a resource allocation scheme based on the resource requirements of each modular node in the toolchain.

[0012] In an optional implementation, the node management module includes: The node classification and permission control unit is used to divide modular nodes into public nodes and private nodes according to their functions, set permissions for private nodes, and perform version management and operation control for modular nodes. The node online / offline unit is used to bring modular nodes that have passed the pre-launch test online. It is also used to check whether there is a task using the corresponding toolchain template when a tool node needs to be edited. If not, it will take the node offline.

[0013] In an optional implementation, the toolchain task execution module includes: The task scheduling and execution unit is used to generate task instances based on the toolchain template, assign unique task IDs, and push them to the task queue; it also starts modular nodes in sequence according to the task queue and dynamically monitors resource usage. The real-time status synchronization unit is used to push task status through a callback interface and update the status of tasks and nodes periodically based on the task status.

[0014] In an optional implementation, the status monitoring and log management module includes: The multi-dimensional monitoring unit is used to monitor the global resource utilization, task throughput, and node success rate in real time during task execution. It can filter monitoring data by task, node, and time dimensions and generate visual reports. The log aggregation and tracing unit is used to read node logs and aggregate and store node logs by task ID; it also embeds full-text search, keyword highlighting, and log link tracing functions into the aggregated node logs.

[0015] In an optional implementation, the execution result processing module includes: The execution result archiving unit is used to read the execution results generated by the output node, persist the execution results to a specified storage, and generate a unique access connection; The execution result push and linkage unit is used to trigger downstream processes through Webhook based on the execution result.

[0016] Secondly, embodiments of this application also provide a large model flow orchestration method, including: Define modular nodes including input nodes, tool nodes, and output nodes, and configure node parameters and connection relationships to generate toolchain templates; The toolchain building process involves dragging and dropping modular nodes, verifying data format compatibility, and configuring dynamic parameters and resource allocation schemes. Manage the permissions and versions of modular nodes, and verify the integrity of the process by performing pre-launch tests; Schedule toolchain task instances to monitor resource usage and execute node sequences; It aggregates log data and provides cross-node error tracing, persists runtime results, and triggers downstream linkage processes.

[0017] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the large model flow orchestration method described in any of the above.

[0018] Fourthly, embodiments of this application also provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the large model flow orchestration method as described in any of the above.

[0019] As can be seen from the above technical solutions, the present invention has the following advantages: The large model workflow orchestration system provided in this application, through the flexible combination of modular nodes and an automated orchestration mechanism, can quickly construct standardized workflows for the entire lifecycle of large models (such as data preprocessing, model training, evaluation, and deployment), avoiding redundant development and manual configuration, and significantly shortening workflow setup time. The nested sub-workflow design of the toolchain templates and the parameter template reuse function further reduce operational complexity, enabling new task workflows to be deployed within minutes, effectively improving overall efficiency.

[0020] This application supports custom extensions of input nodes, tool nodes, and output nodes. Users can quickly register third-party tool nodes or encapsulate composite nodes through a visual interface to meet diverse business needs. Dynamic parameter configuration (such as JSONPath variable binding) and resource pre-allocation mechanisms enable the process to adapt to different data scales and hardware environments, flexibly responding to changes in business scenarios and effectively improving the efficiency of toolchain expansion.

[0021] This application employs a dynamic resource allocation and task queue scheduling strategy to intelligently balance computing load and avoid resource idleness or over-provisioning. The sandbox environment verification and conflict detection functions in pre-launch testing can expose process design flaws (such as circular dependencies and data format conflicts) in advance, reducing the failure rate of online tasks. During task execution, fault tolerance mechanisms such as error retries and dependency rollback are supported, effectively improving the task success rate.

[0022] This application features a visual drag-and-drop interface and automated parameter verification, enabling non-technical personnel to quickly build complex processes and reducing reliance on professionals. It also includes log tracing and cross-node error attribution capabilities, reducing fault location time from hours to minutes.

[0023] This application restricts unauthorized users' access to sensitive tools or data through tiered node permissions (public / private nodes) and version control of toolchain templates. By providing end-to-end encrypted storage of runtime logs and results, combined with a real-time task status monitoring dashboard, it meets data security compliance requirements. All operation records (such as node configuration changes and task execution history) are audited, supporting post-event traceability and accountability.

[0024] This application achieves automated linkage between output nodes and downstream systems, supports pushing runtime results to external services via Webhook, and realizes a closed-loop process encompassing "data processing - model training - evaluation - deployment." By monitoring metrics such as resource utilization and task throughput, it provides data support for cross-team collaboration, effectively improving collaborative efficiency. Attached Figure Description

[0025] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A schematic diagram of the structure of the large model process orchestration system provided in this application.

[0027] Figure 2 A flowchart illustrating the large model process orchestration method provided in this application.

[0028] Figure 3 A flowchart illustrating the usage of the large model workflow orchestration system provided in this application.

[0029] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0030] The various embodiments of this disclosure will be described more fully in the detailed system architecture and functions of the large model flow orchestration system described below. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.

[0031] In the following, the terms “comprising” or “may include”, which may be used in various embodiments of this disclosure, indicate the presence of the disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. Furthermore, as used in various embodiments of this disclosure, the terms “comprising,” “having,” and their cognates are intended only to indicate a particular feature, number, step, operation, element, component, or combination of the foregoing, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more combinations of the foregoing.

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figure 1 The diagram shown is a structural schematic of a large-scale model process orchestration system in a specific embodiment. The system includes: a toolchain development module, a toolchain management module, a node management module, a pre-launch testing module, a toolchain task execution module, a status monitoring and log management module, and an execution result processing module.

[0034] The toolchain development module is used to define modular nodes, including input nodes, tool nodes, and output nodes, configure these modular nodes, and generate toolchain templates that include node execution order, resource quotas, and error handling strategies. Specifically, it defines three basic types of nodes: input nodes, tool nodes, and output nodes. Node parameters, input / output interfaces, and connection relationships are configured through a visual interface, generating toolchain templates that include node execution order, resource quotas, and error handling strategies, and supports nested sub-processes.

[0035] In a specific implementation, the toolchain development module includes a node type definition unit and a toolchain template generation unit.

[0036] The node type definition unit is used to define input nodes, tool nodes, and output nodes, and to configure node parameters, input / output interfaces, and connection relationships through a visual interface. The input nodes are used for data access, data preprocessing, and format conversion. The tool nodes are used for model training, inference, evaluation, and deployment. The output nodes are used for storing running results, pushing data, and visualization.

[0037] This example defines three basic types of nodes—input nodes, tool nodes, and output nodes—as modular nodes. Users can configure node parameters, input / output interfaces, and connection relationships through a visual interface.

[0038] The input nodes support data access (such as databases, OSS buckets, and API interfaces), data preprocessing, and format conversion; the tool nodes cover tools for model training, inference, evaluation, and deployment, and support dynamic allocation of GPU resources; the output nodes are used for result storage (such as databases and file systems), data push, and visualization.

[0039] The toolchain template generation unit is used to generate toolchain templates based on the combination relationship of modular nodes, and encapsulates multiple tool nodes into composite nodes through nested sub-processes.

[0040] For example, a toolchain template can be generated based on the combination relationship of modular nodes, supporting nested sub-processes (such as encapsulating multiple tool nodes into a composite node). The template includes metadata such as node execution order, resource quotas, and error handling strategies (such as retry mechanisms and dependency rollback).

[0041] The toolchain management module is used to define toolchain processes by dragging and dropping modular nodes. It automatically verifies data format compatibility and topology validity between nodes, and performs toolchain parameter configuration and resource pre-allocation. Specifically, it allows users to define toolchain processes by dragging and dropping modular nodes, automatically verifies data format compatibility and topology validity between nodes, and supports dynamic parameter configuration, resource pre-allocation, and conflict detection.

[0042] In a specific implementation, the toolchain management module includes: a toolchain creation unit, a parameter configuration unit, and a resource pre-allocation unit.

[0043] The toolchain creation unit is used to define the toolchain process by dragging and dropping modular nodes onto the canvas and configuring connection logic, and automatically verifying the data format compatibility and topology legality between modular nodes.

[0044] For example, users define the toolchain flow by dragging and dropping nodes onto the canvas and configuring connection logic. The system automatically verifies the compatibility of data formats between nodes (such as the matching of input node outputs with tool node inputs) and the validity of the topology; if conflicts exist, it prompts the user to make adjustments.

[0045] The parameter configuration unit is used to configure the static parameters of modular nodes through a visual form, and to configure the dynamic parameters of modular nodes based on the running results of upstream nodes.

[0046] Each module node supports dynamic parameter configuration. Dynamic parameter configuration includes visual form input of static parameters (such as model hyperparameters and file paths) and dynamic parameters that are bound to the results of upstream nodes through JSONPath expressions; In addition, this system provides a parameter template library, which supports one-click reuse of historical configurations.

[0047] The resource pre-allocation unit is used to generate a resource allocation scheme based on the resource requirements of each modular node in the toolchain.

[0048] For example, based on the resource requirements of each node in the toolchain (such as the number of GPUs and memory quotas), a resource allocation scheme is automatically generated, and manual adjustments are supported. Resource conflicts (such as GPU over-provisioning) are verified on the backend, and a pre-execution report is generated for user review.

[0049] The node management module is used to classify modular nodes by function and control access permissions, and to perform node online and offline operations.

[0050] In a specific implementation, the node management module includes a node classification and permission control unit and a node online / offline unit.

[0051] The node classification and permission control unit is used to divide modular nodes into public and private nodes according to their functions, set permissions for private nodes, and perform version management and operation control for modular nodes.

[0052] For example, nodes are divided into public nodes (built into the platform) and private nodes (user-defined) based on their functions. Private nodes support hierarchical permission settings. The super administrator can manage node versions (such as upgrades and rollbacks) and perform global disable / enable operations. The hierarchical permission settings for private nodes include visibility control for the creator or team, such as visibility only to the creator or team.

[0053] The node online / offline unit is used to bring modular nodes that have passed the pre-launch test online. It is also used to check whether there is a task using the corresponding toolchain template when a tool node needs to be edited. If not, it will take the node offline.

[0054] For example, before a toolchain template is deployed online, its functionality must be verified through pre-deployment testing. If tool nodes need to be edited, they must be taken offline first, and before taking them offline, it must be verified whether any tasks are currently using the template.

[0055] The pre-launch testing module is used to automatically create a sandbox environment before the toolchain is released, simulate real resource allocation and data flow to perform full-process testing, and generate test reports.

[0056] For example, the pre-launch testing module is specifically used for: Before the toolchain is released, the system automatically creates a sandbox environment to simulate real resource allocation and data flow, and performs full-process testing.

[0057] The sandbox environment simulation verifies the correctness of data transmission, node execution status, and resource consumption. The test results include node execution status, resource consumption, data transmission correctness, and error logs, and generate a test report for user review.

[0058] The toolchain task execution module is used to generate task instances and schedule their execution based on the toolchain template. Specifically, it supports parallel execution of nodes without dependencies, sequential execution of nodes with strong dependencies, and task interruption / reset operations.

[0059] In a specific implementation, the toolchain task execution module includes: a task scheduling and execution unit and a real-time status synchronization unit.

[0060] The task scheduling and execution unit is used to generate task instances based on the toolchain template, assign unique task IDs, and push them to the task queue; it also starts modular nodes in sequence according to the task queue and dynamically monitors resource usage.

[0061] For example, on the backend, task instances are generated based on the toolchain template, assigned unique task IDs, and pushed to the task queue. On the algorithm side, nodes are started sequentially, and resource usage is dynamically monitored. For nodes without dependencies, execution is performed in parallel to improve efficiency; for nodes with strong dependencies, execution is performed sequentially, and if an upstream node fails, the downstream task is terminated and an alarm is triggered.

[0062] The real-time status synchronization unit is used to push task status through a callback interface and update the status of tasks and nodes periodically based on the task status.

[0063] For example, the algorithm pushes task status (running, queued, successful, failed) to the backend via a callback interface, and updates the task and node status on the frontend interface periodically. During runtime, users can interrupt running tasks or reset failed tasks to a specified node for re-execution.

[0064] The status monitoring and log management module is used to monitor and display the operation information of resources and modular nodes in real time during task execution, aggregate and store node logs, and provide full-text search and link tracing functions.

[0065] In a specific implementation, the status monitoring and log management module includes: a multi-dimensional monitoring unit and a log aggregation and tracing unit.

[0066] The multi-dimensional monitoring unit is used to monitor global resource utilization (such as GPU card utilization), task throughput, and node success rate in real time during task execution. It filters monitoring data by task, node, and time dimensions and generates visual reports (such as line charts and heatmaps). The multi-dimensional monitoring unit can also serve as a multi-dimensional monitoring dashboard to display global resource utilization, task throughput, and node success rate.

[0067] The log aggregation and tracing unit is used to read node logs and aggregate and store node logs by task ID; it also embeds full-text search, keyword highlighting, and log link tracing functions into the aggregated node logs.

[0068] Through the log tracing feature, users can locate the root cause of cross-node errors (such as abnormal data format transmission paths).

[0069] The execution result processing module is used to persist the execution results generated by the output node to a specified storage and generate an access link, and trigger downstream processes based on the execution results.

[0070] In a specific implementation, the execution result processing module includes: an execution result archiving unit and an execution result push and linkage unit.

[0071] The execution result archiving unit is used to read the execution results generated by the output node, persist the execution results to the specified storage, and generate a unique access connection.

[0072] For example, the output node persists the final result to a specified storage (such as an OSS path or database table) and generates a unique access link.

[0073] The results archiving unit is also used to automatically clean up intermediate results to free up storage space, and data that needs to be retained for a long time can be marked as "critical artifacts".

[0074] The execution result push and linkage unit is used to trigger downstream processes through Webhook based on the execution results, such as automatically triggering model deployment tasks based on model evaluation results.

[0075] In this embodiment, a standardized workflow for the entire lifecycle of large-scale model tasks is constructed through the flexible combination and automated orchestration of modular nodes (input nodes, tool nodes, and output nodes). This supports nested sub-processes and parameter template reuse, enabling new task workflows to be deployed within minutes and reducing manual intervention. The system introduces a dynamic GPU resource allocation mechanism and task queue scheduling strategy to intelligently balance computing load. Pre-deployment testing verifies the integrity of the workflow, and combined with error retries and dependency rollback mechanisms, significantly improving task execution success rates. The system provides a visual drag-and-drop interface and automated parameter verification, allowing non-technical personnel to quickly build complex workflows. Log tracing and automatic intermediate result cleanup shorten fault location time and reduce storage redundancy. The system supports custom node expansion and dynamic parameter configuration (such as JSONPath variable binding) to adapt to diverse business needs. Automated linkage between output nodes and downstream systems (such as Webhook push) achieves a closed-loop workflow from "data processing to training to evaluation to deployment," improving cross-team collaboration efficiency.

[0076] like Figure 2 As shown, the following are embodiments of the large model flow orchestration method provided in this disclosure. This system belongs to the same inventive concept as the large model flow orchestration system in the above embodiments. For details not described in detail in the embodiments of the large model flow orchestration method, please refer to the embodiments of the large model flow orchestration system described above.

[0077] A method for orchestrating large-scale model processes includes the following steps: S1: Define modular nodes including input nodes, tool nodes, and output nodes, configure node parameters and connection relationships to generate toolchain templates; S2: Drag and drop modular node building toolchain process, verify data format compatibility and configure dynamic parameters and resource allocation schemes; S3: Manages the permissions and versions of modular nodes, and verifies the integrity of the process by performing pre-launch tests; S4: Schedule toolchain task instance, monitors resource usage and executes node sequence; S5: Aggregates log data and provides cross-node error tracing, persists runtime results, and triggers downstream linkage processes.

[0078] The large-scale model workflow orchestration method provided in this embodiment achieves flexible orchestration of data processing workflows by generating toolchain templates through modular node definition and configuration; it automatically verifies the data format compatibility between nodes during toolchain management to ensure smooth connection of data processing workflows; it ensures the stability and reliability of data processing workflows before release through pre-launch testing in a sandbox environment; it monitors resource and node operation information in real time during task execution to provide dynamic protection for data processing; it aggregates and stores node logs and provides full-text search and link tracing functions to facilitate the investigation and tracing of data processing problems; finally, it persists the running results and generates access links, and can also trigger downstream processes based on the results to achieve efficient utilization and flow of data processing results, comprehensively improving the quality, efficiency and maintainability of data processing.

[0079] Accordingly, to better illustrate the application process of the large-scale model workflow orchestration system and method disclosed in this invention, see [link to documentation]. Figure 3 As shown, this invention also discloses a method for using a large-scale model workflow orchestration system. This method is applied to the user side, and through user operation, the orchestration of large-scale model workflows is realized. The method specifically includes the following steps: Step 1: Select a toolchain template.

[0080] First, access the large-scale workflow orchestration system's interface and browse available templates in the toolchain template list. The template list is categorized by creation time, name, and purpose. Then, use the search function to quickly locate the desired template by entering keywords (such as template name, related task type, etc.). After selecting a suitable toolchain template, the system will display detailed information about it, including metadata such as node composition, execution order, and resource quotas. After user confirmation, click the "Select" button.

[0081] Step 2: Upload the file via the input node.

[0082] After selecting a toolchain template, find the input node based on the node structure of the template displayed on the interface.

[0083] The input node offers multiple data access methods. Users can select the appropriate method based on the file storage location. For example, if they choose to upload from the local file system, they can click the "Upload" button and select the target file in the pop-up file selection box. If the file is stored in a database or OSS storage bucket, they can enter the corresponding connection information and file path.

[0084] The input node supports data preprocessing and format conversion. Users can configure relevant parameters according to their needs, such as setting data format conversion rules and data cleaning strategies. After completion, clicking "Confirm" will upload the file to the system for processing.

[0085] Step 3: Edit the tool node parameters.

[0086] First, locate the tool node in the toolchain, and click on the tool node to enter the parameter editing interface.

[0087] For static parameters, the system provides a visual form where users can fill in the corresponding information according to task requirements, such as hyperparameters of model training nodes (learning rate, number of iterations, etc.) and file paths.

[0088] For dynamic parameters, users can bind the results of upstream nodes through JSONPath expressions, or select a suitable template from the parameter template library to reuse historical configurations with one click.

[0089] If the tool node involves GPU resource allocation, the resource configuration parameters such as the number of GPUs can be adjusted according to the task's computational requirements.

[0090] Step 4: Save and execute.

[0091] After the user completes editing all node parameters, clicking the "Save" button will store the user's configured toolchain process, node parameters, and other information.

[0092] After successful saving, the system automatically triggers a verification mechanism, which verifies the data format compatibility and topology legality between nodes by creating a new unit through the toolchain in the toolchain management module.

[0093] If the verification passes, the user clicks the "Execute" button, and the system generates a task instance based on the toolchain template. The task scheduling and execution unit of the toolchain task running module assigns a unique task ID and pushes it to the task queue to wait for execution.

[0094] Step 5: Execute node tasks according to logic.

[0095] The system's task scheduling and execution unit starts modular nodes sequentially according to the task queue order. For nodes with no dependencies, the system uses parallel execution to improve task execution efficiency; for nodes with strong dependencies, they are executed sequentially.

[0096] During node execution, the system dynamically monitors resource usage, such as GPU card utilization and memory usage. If an upstream node fails, the system automatically terminates the downstream task and triggers an alarm mechanism to notify the user that an error has occurred.

[0097] Step 6: Check the status of templates and node tasks.

[0098] Find the status viewing function entry in the system interface and enter the task status monitoring page.

[0099] The task status monitoring page displays relevant information obtained by the multi-dimensional monitoring unit of the status monitoring and log management module during the task execution process in real time, including global resource utilization, task throughput, node success rate, etc., presented in the form of visual reports (such as line charts and heatmaps).

[0100] Users can filter monitoring data by task, node, and time dimension to view the running status of specific tasks or nodes, such as running, queued, successful, or failed. Additionally, users can click on a specific node to view its detailed execution logs and related parameter information.

[0101] Step 7: Rename the status of each node task.

[0102] On the node task status viewing interface, locate the node task that needs to be renamed and click the edit button for that node task status. After entering edit mode, the user enters the new status name, and after confirmation, the system saves the changes and updates the display of the new node task status name on the interface.

[0103] Step 8: Go live.

[0104] If no errors occur during task execution and the results meet expectations, the user can click the "Go Online" button. If the task has not undergone pre-deployment testing, the system will first trigger the pre-deployment testing module, automatically create a sandbox environment, simulate real resource allocation and data flow, and execute full-process testing.

[0105] After the test is completed, a test report is generated. After the user reviews the report and confirms that there are no problems, the system uses the node online / offline unit of the node management module to go online the relevant modular nodes and officially deploy them to the production environment for subsequent use.

[0106] Figure 4 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.

[0107] The large-scale model flow orchestration method provided in this application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiments of this invention does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. In the embodiments of this invention, the electronic device includes, but is not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.

[0108] Electronic devices may include processors, external memory interfaces, internal memory, universal serial bus (USB) interfaces, charging management modules, power management modules, batteries, wireless communication modules, audio modules, speakers, microphones, sensor modules, buttons, cameras, displays, and SIM card interfaces, etc.

[0109] A processor may include one or more processing units, such as: a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0110] The processor can serve as the nerve center and command center of an electronic device. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0111] The processor may also include memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or that are used repeatedly. If the processor needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.

[0112] An external storage interface (ESI) can be used to connect external memory cards, such as microSD cards, to expand the storage capacity of electronic devices. The external memory card communicates with the processor through the ESI to perform data storage functions, such as saving music and video files on the external memory card.

[0113] Internal memory can be used to store computer executable program code, which includes instructions. The processor executes various functional applications and data processing of electronic devices by running the instructions stored in internal memory. Internal memory can include a program storage area and a data storage area. Internal memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0114] Wireless communication functionality in electronic devices can be achieved through antennas, wireless communication modules, modem processors, and baseband processors.

[0115] Wireless communication modules can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0116] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.

[0117] Electronic devices can achieve shooting functions through ISPs, cameras, video codecs, GPUs, displays, and application processors.

[0118] Electronic devices can achieve display functions through GPUs, displays, and application processors.

[0119] A GPU is a microprocessor for image processing, connected to the display screen and application processor. GPUs are used to perform mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information.

[0120] A display screen is used to display images, videos, etc. A display screen includes a display panel.

[0121] The aforementioned electronic equipment realizes the large model process orchestration method of this application through the flexible combination and automated scheduling of modular nodes, and achieves the process-oriented and standardized management of tasks throughout the entire life cycle of the large model. This achieves the beneficial effects of ensuring flexible orchestration, stable operation, efficient monitoring, traceability of problems, and effective utilization of results in the data processing process.

[0122] The storage medium provided in this application stores a program product capable of implementing a large-scale model process orchestration method.

[0123] Large model process orchestration methods include: Define modular nodes including input nodes, tool nodes, and output nodes, and configure node parameters and connection relationships to generate toolchain templates; The toolchain building process involves dragging and dropping modular nodes, verifying data format compatibility, and configuring dynamic parameters and resource allocation schemes. Manage the permissions and versions of modular nodes, and verify the integrity of the process by performing pre-launch tests; Schedule toolchain task instances to monitor resource usage and execute node sequences; It aggregates log data and provides cross-node error tracing, persists runtime results, and triggers downstream linkage processes.

[0124] In some possible implementations, the large-scale model process orchestration method of this disclosure can be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of this disclosure.

[0125] The storage medium disclosed herein can take the form of any combination of one or more readable media. A readable medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0126] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A large-scale model process orchestration system, characterized in that, include: The toolchain development module is used to define modular nodes, including input nodes, tool nodes, and output nodes, configure modular nodes, and generate toolchain templates that include node execution order, resource quotas, and error handling strategies. The toolchain management module is used to define toolchain processes by dragging and dropping modular nodes, automatically verify the compatibility of data formats and the legality of topology between nodes, and configure toolchain parameters and pre-allocate resources. The node management module is used to classify modular nodes by function and control access permissions, and to perform node online and offline operations. The pre-launch testing module is used to automatically create a sandbox environment before the toolchain is released, simulate real resource allocation and data flow to perform full-process testing, and generate test reports; The toolchain task execution module is used to generate task instances and schedule their execution based on the toolchain template. The status monitoring and log management module is used to monitor and display the operation information of resources and modular nodes in real time during task execution, aggregate and store node logs and provide full-text search and link tracing functions; The execution result processing module is used to persist the execution results generated by the output node to a specified storage and generate an access link, and trigger downstream processes based on the execution results.

2. The large model workflow orchestration system according to claim 1, characterized in that, The toolchain development module includes: The node type definition unit is used to define input nodes, tool nodes, and output nodes, and to configure node parameters, input / output interfaces, and connection relationships through a visual interface. The input nodes are used for data access, data preprocessing, and format conversion; the tool nodes are used for model training, inference, evaluation, and deployment; and the output nodes are used for storing execution results, pushing data, and visualization. The toolchain template generation unit is used to generate toolchain templates based on the combination relationship of modular nodes, and encapsulates multiple tool nodes into composite nodes through nested sub-processes.

3. The large model process orchestration system according to claim 2, characterized in that, The toolchain management module includes: The toolchain creation unit is used to define the toolchain process by dragging and dropping modular nodes onto the canvas and configuring connection logic, and automatically verifying the data format compatibility and topology legality between modular nodes; The parameter configuration unit is used to configure the static parameters of modular nodes through a visual form, and to configure the dynamic parameters of modular nodes based on the running results of upstream nodes. The resource pre-allocation unit is used to generate a resource allocation scheme based on the resource requirements of each modular node in the toolchain.

4. The large model flow orchestration system according to claim 3, characterized in that, The node management module includes: The node classification and permission control unit is used to divide modular nodes into public nodes and private nodes according to their functions, set permissions for private nodes, and perform version management and operation control for modular nodes. The node online / offline unit is used to bring modular nodes that have passed the pre-launch test online. It is also used to check whether there is a task using the corresponding toolchain template when a tool node needs to be edited. If not, it will take the node offline.

5. The large model flow orchestration system according to claim 4, characterized in that, The toolchain task execution module includes: The task scheduling and execution unit is used to generate task instances based on the toolchain template, assign unique task IDs, and push them to the task queue; it also starts modular nodes in sequence according to the task queue and dynamically monitors resource usage. The real-time status synchronization unit is used to push task status through a callback interface and update the status of tasks and nodes periodically based on the task status.

6. The large model process orchestration system according to claim 5, characterized in that, The status monitoring and log management module includes: The multi-dimensional monitoring unit is used to monitor the global resource utilization, task throughput, and node success rate in real time during task execution. It can filter monitoring data by task, node, and time dimensions and generate visual reports. The log aggregation and tracing unit is used to read node logs and aggregate and store node logs by task ID; it also embeds full-text search, keyword highlighting, and log link tracing functions into the aggregated node logs.

7. The large model flow orchestration system according to claim 6, characterized in that, The execution result processing module includes: The execution result archiving unit is used to read the execution results generated by the output node, persist the execution results to a specified storage, and generate a unique access connection; The execution result push and linkage unit is used to trigger downstream processes through Webhook based on the execution result.

8. A method for orchestrating large-scale model processes, characterized in that, The method employs the large model flow orchestration system as described in any one of claims 1 to 7; The method includes: Define modular nodes that include input nodes, tool nodes, and output nodes, and configure node parameters and connection relationships to generate toolchain templates; The toolchain construction process involves dragging and dropping modular nodes, verifying data format compatibility, and configuring dynamic parameters and resource allocation schemes. Manage the permissions and versions of modular nodes, and verify the integrity of the process by performing pre-launch tests; Schedule toolchain task instances to monitor resource usage and execute node sequences; It aggregates log data and provides cross-node error tracing, persists runtime results, and triggers downstream linkage processes.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the large model flow orchestration method as described in claim 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the large model flow orchestration method as described in claim 7.