Workflow scheduling intelligent maintenance method and system based on large model

By adopting a workflow scheduling intelligent maintenance method based on a large model, a closed loop of autonomous analysis and execution is realized, which solves the problems of low operation and maintenance efficiency and insufficient system stability in existing technologies, and improves the operation and maintenance efficiency and system stability of big data workflow scheduling.

CN121501435APending Publication Date: 2026-02-10GUANGZHOU FAISCO INFORMATON TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511496296.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing big data workflow scheduling systems struggle to achieve comprehensive automated maintenance when faced with massive, complex workflows and dynamically changing operating environments, requiring deep manual intervention, resulting in low operational efficiency and insufficient system stability.

Method used

A workflow scheduling intelligent maintenance method based on a large model is adopted. By combining the large model decision layer with the data interface layer and the execution interface layer, a closed loop of autonomous analysis and execution is achieved. Maintenance decisions are generated by semantic understanding, logical reasoning and causal reasoning, and the results are fed back through the interaction layer to reduce human intervention.

Benefits of technology

It improved the operational efficiency of big data workflow scheduling, enhanced system stability, reduced workflow scheduling delays, and increased resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501435A_ABST
    Figure CN121501435A_ABST
Patent Text Reader

Abstract

The invention discloses a workflow scheduling intelligent maintenance method and system based on a large model, and relates to the technical field of big data analysis, and the method comprises the steps: receiving and triggering an instruction; acquiring and processing data; carrying out large model deep analysis and decision generation; on-demand multi-round iteration depth analysis is carried out; executing an operation instruction; and feeding back a result and updating a state. According to the method, an'autonomous analysis-execution 'closed loop of workflow maintenance is realized to the greatest extent by fully utilizing the inference analysis capability and the tool execution capability of a large model, so that the operation and maintenance efficiency and the system stability of big data workflow scheduling are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data analytics technology, and in particular to a method and system for intelligent maintenance of workflow scheduling based on a large model. Background Technology

[0002] In today's big data era, big data processing and analysis have become essential core capabilities for internet companies. Workflow scheduling systems, as a critical underlying infrastructure for big data processing and analysis, handle the scheduling and execution of hundreds of thousands of data processing workflows daily. The rationality of their scheduling dependency configurations, the stability of system operation, and the timeliness of fault recovery are crucial to ensuring business continuity and data quality. Therefore, this places high demands on the daily work of data developers and operations personnel, requiring them to periodically check workflow dependency configurations and respond and intervene promptly when faults occur.

[0003] Although improving platform functionality and tool support has significantly reduced the workload of development and operations personnel, the static maintenance rules on which existing platforms and tools rely are often insufficient to fully adapt to the dynamic operating environment and complex scenarios when faced with massive, complex workflows and extremely intricate dependencies. Deep manual intervention is still required for their use. Summary of the Invention

[0004] The main objective of this application is to propose an intelligent maintenance method and system for workflow scheduling based on a large model, so as to improve the operation and maintenance efficiency and system stability of big data workflow scheduling.

[0005] To achieve the above objectives, one aspect of this application proposes an intelligent maintenance method for workflow scheduling based on a large model, the method comprising the following steps: Receive input commands via dialog box, or send corresponding input commands when the system status triggers the set conditions; Utilizing the large model decision layer based on Function Calling capabilities, it connects to the big data workflow scheduling platform through the data interface layer to obtain real-time operational data of the workflow and connects to the knowledge base to obtain operation guidance manuals. The real-time running data, the operation guide manual, and the input instructions are input into the decision layer of the large model. The semantic understanding, reasoning, causal reasoning, and predictive reasoning of the large model are used to analyze the workflow status or scheduling system status, and then generate analysis results corresponding to the input instructions. If the information required for the analysis is insufficient or a deeper tracing is needed, additional rounds of data acquisition and in-depth analysis are triggered until the large model deems the information sufficient, and then the analysis results are obtained. By utilizing the large model decision layer based on the Function Calling capability and the decision results corresponding to the analysis results, the target operation and maintenance interface encapsulated by the execution interface layer is selected and called to execute several corresponding operation instructions and obtain the execution results, so as to realize the automated closed loop of maintenance operations; The execution results and scheduling system workflow status are fed back to the user through the interaction layer.

[0006] In some embodiments, the large model includes: Interaction Layer: Configured to provide an interaction interface for data developers and platform operators, used to receive workflow-related query requests and operation instructions input by users, and output the decision results, workflow status information and system response generated by the large model decision layer; Data interface layer: configured to connect to the big data workflow scheduling platform, acquire and process the operation data of the scheduling platform, and provide the reference data required for decision-making of the big model decision layer; Execution Interface Layer: Configured to encapsulate the management and maintenance interface of the big data workflow scheduling platform, which can be called by the big model decision layer to execute operation instructions on the scheduling platform; Large Model Decision Layer: Configured to be based on a large model, dynamically combining the scheduling platform operation data provided by the data interface layer, autonomously analyzing the task execution status of the big data workflow scheduling system, generating maintenance decisions; responding to user query requests received by the interaction layer, generating workflow-related Q&A results; and outputting operation instructions through the execution interface layer to change the operation status of the scheduling system.

[0007] In some embodiments, the interaction method of the interaction layer includes: using a natural language-based multimodal dialogue bot as the interaction interface to provide intelligent interaction for data developers and platform operators, specifically including: Dialogue bot architecture: It adopts a combination of multi-turn dialogue and intent recognition. It uses the natural language understanding capabilities of a large model to analyze user intent, converts user input into input for the decision layer, and summarizes and refines the output of the decision layer of the large model into content that is easy for users to understand. Natural language query function: Users describe their query needs in natural language, and the big model automatically parses the query intent and uses the intent as input to the decision layer to generate corresponding results; Proactive early warning push: When the workflow status of the scheduling system meets the configuration rules and is triggered, the system pushes abnormal status information based on the current workflow status and historical operation data.

[0008] In some embodiments, the data in the data interface layer includes: workflow metadata, historical execution data, scheduling platform management and user manual, and manual adjustment memos; The workflow metadata includes: workflow node information, workflow task information, workflow dependency configuration information, and workflow output SLA configuration. The historical operation data includes: historical workflow operation status, historical operation scheduling time, historical operation time, and cluster resource load. The dispatching platform management and user manual includes: an introduction to the dispatching process, business domain rules, best practices for the dispatching platform, and an emergency handling manual; The manual adjustment memo includes: manual processing notes and workflow target configuration.

[0009] In some embodiments, the operation interface functions of the execution interface layer include: Control operations for workflow task instances, including starting, pausing, and resuming; Dynamic modification of workflow task and task node configuration parameters, including modification timing and dependency order.

[0010] In some embodiments, the large model decision layer is constructed based on a multimodal thinking and reasoning large model, used to acquire platform data and reference operation manuals on demand, realize intelligent analysis and decision support for the big data workflow scheduling system, and output the analysis results or execute system operations. The large model decision layer includes: On-demand acquisition of workflow metadata and runtime data: By understanding user input intent, the Function Calling capability based on the large model selects the appropriate interface from the data structure layer, constructs the parameter input, and makes the call; On-demand access to reference manuals and documents: Based on the understanding of user input intent, the system analyzes potential operational guidelines through large-scale model-based prediction strategies and selects relevant content from the platform's operation specifications, best practices, and emergency response manuals knowledge base at the data interface layer. In-depth analysis and reasoning decision-making: Based on the causal reasoning capabilities of a multimodal thinking and reasoning model, it integrates workflow operation status, historical operation data and related reference documents obtained on demand through data interfaces to analyze user input questions; Decision generation and operation instruction output: The analysis results are summarized into a structured target form and returned to the user through the interaction layer, providing clear explanations and evidence; if the system decision also includes the execution of operation instructions, the selection and invocation of instructions are based on the Function Calling capability of the large model. On-demand multi-round iterative analysis: If the data or reference documents obtained in the current round of analysis are insufficient to support the analysis and decision-making process or require continuous source tracing analysis, the decision-making system will conduct several additional rounds of data and document acquisition until the decision-making system determines that the acquired information is sufficient to support the analysis and problem-solving.

[0011] In some embodiments, the on-demand multi-round iterative analysis includes the following steps: First round of analysis: Based on the instruction content, the current running status of the scheduling system or workflow, and the operation and maintenance manual, preliminary analysis results and action plans are generated; Iterative deepening: Call the data interface layer to obtain supplementary data, verify or revise the analysis results and action plan; Convergence Decision: When the confidence level reaches the threshold, output the final maintenance strategy, operation instructions, and total user feedback; Execution feedback: After the operation instructions are executed at the interface execution layer, the system status after execution will be used to determine whether the expected goal is met. If the expected goal is not met, the system will be iterated and deepened again.

[0012] In some embodiments, the analysis results include: Root cause analysis: Based on the upstream dependencies of the current workflow, a dependency DAG graph is constructed. By combining the execution status, completion time and execution logs of each workflow, the root cause workflow of the failure or delay and its scope of impact are located through large model reasoning. Predictive maintenance recommendations: Based on workflow dependencies, historical runtime and time consumption trends, output workflow delay warnings and expected workflow end times; Fault recovery strategy: Based on the upstream dependencies of the current workflow, workflow task priorities and operation guidelines, the system outputs key recovery task paths and priority-based recovery strategies through large-scale model inference and analysis, thereby improving the task recovery efficiency of the scheduling system. Configuration optimization strategy: Based on workflow dependencies, historical runtime, and cluster resources, the system uses large-scale model reasoning to analyze critical paths and identify unreasonable timed schedules. The system then adjusts the timing sequence of workflow schedules to improve cluster resource utilization and optimize workflow data output time.

[0013] To achieve the above objectives, another aspect of this application proposes a workflow scheduling intelligent maintenance system based on a large model, the system comprising: The instruction input unit is used to receive input instructions through a dialog box, or to send corresponding input instructions when the system status triggers a set condition; The data acquisition unit is used to leverage the Function Calling capability of the large model decision layer to connect to the big data workflow scheduling platform through the data interface layer, thereby obtaining real-time running data of the workflow and connecting to the knowledge base to obtain operation guidance manuals. The data analysis unit is used to input the real-time running data, the operation guide manual, and the input instructions into the large model decision layer, and use the semantic understanding, reasoning, causal reasoning, and predictive reasoning of the large model to analyze the workflow status or scheduling system status, and then generate analysis results corresponding to the input instructions. The reanalysis unit is used to trigger additional rounds of data acquisition and in-depth analysis if the information required for analysis is insufficient or if a deeper tracing is needed, until the large model deems the information sufficient, and then the analysis results are obtained. The decision execution unit is used to select and call the target operation and maintenance interface encapsulated by the execution interface layer based on the function calling capability and the decision result corresponding to the analysis result of the large model decision layer, execute several corresponding operation instructions, and obtain the execution result, so as to realize the automated closed loop of maintenance operation. The feedback interaction unit is used to feed back the execution results and scheduling system workflow status to the user through the interaction layer.

[0014] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0015] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0016] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0017] The embodiments of this application include at least the following beneficial effects: This application provides a method and system for intelligent maintenance of workflow scheduling based on a large model. The solution of this application fully utilizes the reasoning and analysis capabilities and tool execution capabilities of the large model to maximize the "autonomous analysis-execution" closed loop of workflow maintenance, thereby effectively improving the operation and maintenance efficiency and system stability of big data workflow scheduling. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating an intelligent maintenance method for workflow scheduling based on a large model, provided as an embodiment of this application; Figure 2 A schematic diagram of the structure of a workflow scheduling intelligent maintenance system based on a large model is provided in this application embodiment; Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.

[0022] This application provides a method and system for intelligent maintenance of workflow scheduling based on a large model, relating to the field of big data analytics. The method and system provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing a method for intelligent maintenance of workflow scheduling based on a large model, but is not limited to the above forms.

[0023] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0024] Reference Figure 1 This application provides an intelligent maintenance method for workflow scheduling based on a large model. This method may include, but is not limited to, steps S100 to S600, as detailed below: S100: Command Reception and Triggering: The system receives commands from developers or operations personnel through the bot dialog box, or sends corresponding commands when the system status triggers the set conditions, such as "Please handle a certain workflow delay".

[0025] S200: Real-time data acquisition and processing: The large model decision layer, based on Function Calling capabilities, connects to the big data workflow scheduling platform through the data interface layer to obtain workflow metadata and runtime data, and connects to the knowledge base to obtain operation guidance manuals.

[0026] S300: Large Model Deep Analysis and Decision Generation: Input real-time running data, operation manuals and received instructions into the large model decision layer, and use the semantic understanding, reasoning, causal reasoning and predictive reasoning of the large model to perform deep analysis on the workflow status or scheduling system status, and generate intelligent analysis results corresponding to the input instructions, including: root cause analysis, predictive maintenance suggestions, fault recovery strategies or configuration tuning strategies.

[0027] S400: On-demand multi-round iterative in-depth analysis: In step S3, if the information required for the analysis is insufficient or further in-depth tracing is needed, additional rounds of data acquisition and in-depth analysis are actively triggered until the large model considers the information sufficient to support high-quality analysis.

[0028] S500: Operation command execution: Based on the Function Calling capability, the large model decision layer can optionally automatically select and call specific operation and maintenance interfaces encapsulated in the execution interface layer based on the generated decision results, and execute several corresponding operation commands (such as starting / stopping workflow, modifying configuration, etc.) to realize the automated closed loop of maintenance operations.

[0029] S600: Result Feedback and Status Update: The execution results and scheduling system workflow status are fed back to the user through the interaction layer.

[0030] The S100 receives and triggers commands. The dialogue bot's input is natural language and supports multi-turn dialogue. It recognizes user intent based on the natural language processing capabilities of the large model and converts the user's intent into input for the large model's decision layer. Simultaneously, the system also supports inputting predefined commands into the large model's decision layer via structure when certain conditions are met.

[0031] S200 data acquisition and processing in real time includes workflow metadata such as workflow node information, workflow task information, workflow dependency configuration information, and workflow output SLA configuration; workflow operation data includes: historical workflow operation status, historical operation scheduling time, historical operation time, and cluster resource load; and management and user manuals include: scheduling process introduction, business domain rules, best practices for the scheduling platform, and emergency handling manual.

[0032] S300 large-scale model in-depth analysis and decision-making, the core analysis and decision-making steps include: on-demand acquisition of workflow metadata and runtime data; on-demand acquisition of management and user manual documents; in-depth analysis and reasoning decision-making; decision generation and operation instruction output; on-demand multi-round iterative analysis; The S300 large-scale model performs in-depth analysis and decision-making, and outputs the following content based on different instructions and scenarios: root cause analysis, which is to build a DAG dependency graph to locate the fault root cause workflow and its impact scope; predictive maintenance suggestions, which is to output delay warnings and expected end times; fault recovery strategies, which is to generate priority recovery paths; and configuration tuning strategies, which is to optimize timing and resource allocation. The S400 performs in-depth analysis on demand through multiple iterations. If insufficient information is required, it can automatically trigger multiple rounds of data acquisition and iterative reasoning until the confidence level is met, ensuring the depth and accuracy of the analysis.

[0033] The S400's on-demand, multi-round iterative in-depth analysis process includes the following steps: first-round analysis of fusion instructions, platform data, and user manuals to generate preliminary conclusions; calling data interfaces to obtain more data for further analysis to verify or revise the analysis conclusions; outputting the final conclusions or operation instructions after the confidence level meets expectations; and, based on the system or workflow status after executing the instructions, optionally continuing multiple rounds of iteration.

[0034] The S500 operation execution command execution interface provides the following operation interface capabilities: control operations of workflow task instances, such as start, pause and resume; dynamic modification of workflow task and node configuration parameters, such as modifying timing and dependency configurations; The core innovations of this application include: The large-model-driven intelligent analytics engine leverages the function calling capabilities of large models to achieve on-demand data acquisition and API calls, breaking the limitations of traditional fixed processes. Simultaneously, through causal reasoning and predictive analysis, it deciphers the root causes of problems from dependent metadata and historical operational data, enabling root cause localization and trend prediction. Finally, on-demand, multi-round iterations ensure the depth and accuracy of the analysis.

[0035] The system architecture is layered and decoupled. The solution divides the system into: an interaction layer, which adopts a multi-turn dialogue bot to support natural language query and proactive message push; a data interface layer, which uniformly encapsulates data elements such as metadata, runtime data, and management and usage documents; an execution interface layer, which provides atomic operation and maintenance capabilities such as workflow control and configuration modification; and a large model decision layer, which serves as the core hub and dynamically coordinates the various layers to achieve a closed loop of "analysis-decision-execution".

[0036] The following sections will provide a detailed description and explanation of some optional embodiments of this application, using specific application examples.

[0037] This embodiment discloses an intelligent maintenance method for workflow scheduling based on a large model. The method includes: an interaction layer receiving manual instructions or system-triggered instructions from a dialogue bot; a large model decision layer autonomously acquiring the required data and operation guidance documents through a data interface layer; the large model decision layer performing in-depth analysis and decision generation based on the data and documents, outputting root cause analysis, predictive maintenance suggestions, fault recovery strategies, or configuration optimization strategies; when data is insufficient to support the analysis, the large model decision layer autonomously conducting multiple rounds of iterative in-depth analysis until the analysis results meet the confidence requirements; based on the generated decision results, the large model decision layer autonomously selecting and executing the structure of the execution interface layer; and the decision results and analysis results being fed back to the user or system through the interaction layer. This method reduces manual intervention through automated closed-loop analysis and maintenance driven by a large model, significantly improving the maintenance efficiency of workflow scheduling. Simultaneously, the method's autonomous analysis and optimization also reduces workflow scheduling output delays and improves workflow scheduling stability.

[0038] Still refer to Figure 1 This embodiment of a workflow scheduling intelligent maintenance method based on a large model includes the following steps: S100: Command reception and triggering.

[0039] The system receives natural language commands (e.g., "Please handle a workflow delay") from developers or operations personnel via a chatbot interface, or automatically generates corresponding commands when the system status meets preset trigger conditions. The chatbot is developed based on the WeChat Work chatbot interface and supports integration with collaboration platforms such as DingTalk and Slack. For natural language input, the large model will automatically trigger a multi-turn dialogue mechanism until it accurately understands the user's intent when intent recognition is unclear. System status-triggered commands are generated based on preset rules; for example, when a workflow fails or its end time exceeds the SLA (Service Level Agreement) time limit, the system will automatically generate a command containing the current trigger information.

[0040] S200: Data Acquisition and Processing.

[0041] The large model decision layer utilizes its Function Calling capability to connect to the big data scheduling platform and knowledge base through the data interface layer to obtain the required information. Function Calling capability refers to the large model's ability to intelligently select one or more tools from a predefined and detailed list of Tools based on actual needs, construct input parameters according to their parameter requirements, and return them to the caller. Therefore, this system selects a large model with inference capabilities and support for Function Calling (such as GLM4.5, or similar models like Doubao and Tongyi Qianwen) as the decision core. This layer receives instructions (user natural language input or system instructions) obtained in step S100, understands the user's intent, and uses this intent and potential data requirements as the basis for selecting Tools. The data interface layer encapsulates three types of key data: 1. Workflow metadata: including workflow node information, task information, dependency configuration information, and output SLA configuration.

[0042] 2. Workflow historical operation data: including workflow operation status, historical scheduling time, historical operation time, and cluster resource load.

[0043] 3. Dispatch Platform Management and User Manual: Includes an introduction to the dispatch process, business domain rules, platform best practices, and emergency handling manual.

[0044] Optionally, the management and user manual documents are manually organized and segmented, and RAG (Retrieval Augmentation) technology is used as the underlying retrieval mechanism to ensure that the large model can efficiently obtain the most relevant document fragments while avoiding context length overflow.

[0045] S300: Large-scale model in-depth analysis and decision generation.

[0046] The user intent parsed in step S100, the workflow metadata obtained in step S200, historical execution data, and relevant management manual fragments, combined with preset system prompts, constitute the input context of the large model. Simultaneously, the three types of data interfaces defined in step S200 are passed to the large model as a Tools list. The system calls the large model interface to initiate the initial inference process: 1. If the large model determines that the current context information is sufficient to make the final analysis and decision, the decision result will be output directly, and the calling instructions of the interface layer Tool can be selectively attached.

[0047] 2. If the large model determines that the current context information is insufficient to support the final analysis and decision, it outputs the preliminary decision results and the planning steps for subsequent analysis and decision, and may optionally include the calling instructions of the data interface layer Tool to obtain supplementary information.

[0048] S400: On-demand multi-round iterative in-depth analysis.

[0049] When the confidence level of the context information in the large model of S300 is insufficient, the system enters the on-demand multi-round iterative deep analysis phase. Based on the Tool call information output by the large model in the previous round, the system initiates one or more calls to the data interface layer. The interface call results are added to the system's "memory." Here, "memory" refers to the complete historical record of the system's interaction with the large model, including all output data from the initial system prompt to the latest Tool execution. After each Tool call, the output of the large model and the interface return results of this round are appended to the context history as new dialogue data. Each new call to the large model requires a complete context containing all historical rounds as input. In this step, the system cyclically executes large model calls and data interface calls, and the large model dynamically selects the required Tool based on the current analysis and decision state. The iteration process terminates when any of the following conditions are met: 1. The large model initiates a call to the execution interface layer Tool; 2. The large model has already explicitly output the final analysis and decision results and does not require further tool calls; 3. The number of iterations reaches the maximum threshold preset by the system.

[0050] S500: Operation command execution.

[0051] In S300 or S400 systems, if the output of the large model contains call instructions for the execution interface layer Tool, the system will execute these interface calls sequentially. The interface call parameters are automatically constructed by the large model based on the Tool description information. After execution, the interface return results must be added to the system's "memory" as data for the next round of dialogue. The execution interface layer possesses the following core capabilities: 1. Workflow task instance control: such as starting, pausing, and resuming task instances.

[0052] 2. Dynamic adjustment of workflow task and node configurations: such as adjusting scheduling timing rules and scheduling dependencies.

[0053] S600: Result feedback and status update.

[0054] After S500 is completed, the system will integrate the iterative analysis outputs, interface call records, and return results from all previous steps (S100-S500) (i.e., the complete context stored in "memory"), and initiate another round of large-scale model interaction. This interaction is given preset summary prompts, aiming to guide the large-scale model to evaluate the entire process from decision analysis to execution results. The evaluation content includes: 1. Has a clear and confident analytical decision been made regarding the instruction? 2. After executing the Tool call at the interface layer, does the actual execution status of the system match expectations?

[0055] If the confidence level of the large model's decision analysis meets expectations, the final analysis execution result will be generated and returned to the user through the interaction layer. If the large model's decision still requires further iteration (e.g., insufficient decision confidence or execution status not meeting expectations), the system will continue to iterate through S400 (deep analysis) or S500 (execution operation) based on the output of the large model, until it re-enters S600 to complete the evaluation and output the result, or reaches the system's preset maximum number of iterations limit.

[0056] Reference Figure 2 This application also provides a workflow scheduling intelligent maintenance method system based on a large model, which can implement the above-mentioned workflow scheduling intelligent maintenance method based on a large model. The system includes: The instruction input unit is used to receive input instructions through a dialog box, or to send corresponding input instructions when the system status triggers a set condition; The data acquisition unit is used to leverage the Function Calling capability of the large model decision layer to connect to the big data workflow scheduling platform through the data interface layer, thereby obtaining real-time running data of the workflow and connecting to the knowledge base to obtain operation guidance manuals. The data analysis unit is used to input the real-time running data, the operation guide manual, and the input instructions into the large model decision layer, and use the semantic understanding, reasoning, causal reasoning, and predictive reasoning of the large model to analyze the workflow status or scheduling system status, and then generate analysis results corresponding to the input instructions. The reanalysis unit is used to trigger additional rounds of data acquisition and in-depth analysis if the information required for analysis is insufficient or if a deeper tracing is needed, until the large model deems the information sufficient, and then the analysis results are obtained. The decision execution unit is used to select and call the target operation and maintenance interface encapsulated by the execution interface layer based on the function calling capability and the decision result corresponding to the analysis result of the large model decision layer, execute several corresponding operation instructions, and obtain the execution result, so as to realize the automated closed loop of maintenance operation. The feedback interaction unit is used to feed back the execution results and scheduling system workflow status to the user through the interaction layer.

[0057] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0058] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method of this application. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0059] It is understood that the content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the methods of this application, and the beneficial effects achieved are the same as those achieved by the methods of this application.

[0060] Please see Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 302 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called and executed by the processor 301. Input / output interface 303 is used to implement information input and output; The communication interface 304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 305 transmits information between various components of the device (e.g., processor 301, memory 302, input / output interface 303, and communication interface 304); The processor 301, memory 302, input / output interface 303, and communication interface 304 are connected to each other within the device via bus 305.

[0061] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of this application.

[0062] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0063] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0064] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0065] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0066] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0067] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0068] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0069] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0070] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0071] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0072] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0074] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0075] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A workflow scheduling intelligent maintenance method based on a large model, characterized in that, The method includes the following steps: Receive input commands via dialog box, or send corresponding input commands when the system status triggers the set conditions; Utilizing the large model decision layer based on Function Calling capabilities, it connects to the big data workflow scheduling platform through the data interface layer to obtain real-time operational data of the workflow and connects to the knowledge base to obtain operation guidance manuals. The real-time running data, the operation guide manual, and the input instructions are input into the decision layer of the large model. The semantic understanding, reasoning, causal reasoning, and predictive reasoning of the large model are used to analyze the workflow status or scheduling system status, and then generate analysis results corresponding to the input instructions. If the information required for the analysis is insufficient or a deeper tracing is needed, additional rounds of data acquisition and in-depth analysis are triggered until the large model deems the information sufficient, and then the analysis results are obtained. By utilizing the large model decision layer based on the Function Calling capability and the decision results corresponding to the analysis results, the target operation and maintenance interface encapsulated by the execution interface layer is selected and called to execute several corresponding operation instructions and obtain the execution results, so as to realize the automated closed loop of maintenance operations; The execution results and scheduling system workflow status are fed back to the user through the interaction layer.

2. The intelligent maintenance method for workflow scheduling based on a large model according to claim 1, characterized in that, The large model includes: Interaction Layer: Configured to provide an interaction interface for data developers and platform operators, used to receive workflow-related query requests and operation instructions input by users, and output the decision results, workflow status information and system response generated by the large model decision layer; Data interface layer: configured to connect to the big data workflow scheduling platform, acquire and process the operation data of the scheduling platform, and provide the reference data required for decision-making of the big model decision layer; Execution Interface Layer: Configured to encapsulate the management and maintenance interface of the big data workflow scheduling platform, which can be called by the big model decision layer to execute operation instructions on the scheduling platform; Large Model Decision Layer: Configured to be based on a large model, dynamically combining the scheduling platform operation data provided by the data interface layer, autonomously analyzing the task execution status of the big data workflow scheduling system, generating maintenance decisions; responding to user query requests received by the interaction layer, generating workflow-related Q&A results; and outputting operation instructions through the execution interface layer to change the operation status of the scheduling system.

3. The intelligent maintenance method for workflow scheduling based on a large model according to claim 2, characterized in that, The interaction methods of the interaction layer include: using a multimodal dialogue bot based on natural language as the interaction interface to provide intelligent interaction for data developers and platform maintenance personnel, specifically including: Dialogue bot architecture: It adopts a combination of multi-turn dialogue and intent recognition. It uses the natural language understanding capabilities of a large model to analyze user intent, converts user input into input for the decision layer, and summarizes and refines the output of the decision layer of the large model into content that is easy for users to understand. Natural language query function: Users describe their query needs in natural language, and the big model automatically parses the query intent and uses the intent as input to the decision layer to generate corresponding results; Proactive early warning push: When the workflow status of the scheduling system meets the configuration rules and is triggered, the system pushes abnormal status information based on the current workflow status and historical operation data.

4. The intelligent maintenance method for workflow scheduling based on a large model according to claim 2, characterized in that, The data in the data interface layer includes: workflow metadata, historical operation data, scheduling platform management and user manual, and manual adjustment memos; The workflow metadata includes: workflow node information, workflow task information, workflow dependency configuration information, and workflow output SLA configuration. The historical operation data includes: historical operation status of the workflow, historical operation scheduling time, historical operation time, and cluster resource load. The dispatching platform management and user manual includes: an introduction to the dispatching process, business domain rules, best practices for the dispatching platform, and an emergency handling manual; The manual adjustment memo includes: manual processing notes and workflow target configuration.

5. The intelligent maintenance method for workflow scheduling based on a large model according to claim 2, characterized in that, The operation interface functions of the execution interface layer include: Control operations for workflow task instances, including starting, pausing, and resuming; Dynamic modification of workflow task and task node configuration parameters, including modification timing and dependency order.

6. The intelligent maintenance method for workflow scheduling based on a large model according to claim 2, characterized in that, The large-scale model decision layer is constructed based on a multimodal thinking and reasoning large-scale model. It is used to acquire platform data and reference operation manuals on demand, realize intelligent analysis and decision support for the big data workflow scheduling system, and output the analysis results or execute system operations. The large-scale model decision layer includes: On-demand acquisition of workflow metadata and runtime data: By understanding user input intent, the Function Calling capability based on the large model selects the appropriate interface from the data structure layer, constructs the parameter input, and makes the call; On-demand access to reference manuals and documents: Based on an understanding of user input intent, the system analyzes potential operational guidelines through large-scale model-based predictive strategies and selects relevant content from the platform's operational specifications, best practices, and emergency response manuals knowledge base at the data interface layer. In-depth analysis and reasoning decision-making: Based on the causal reasoning capabilities of a multimodal thinking and reasoning model, it integrates workflow operation status, historical operation data and related reference documents obtained on demand through data interfaces to analyze user input questions; Decision generation and operation instruction output: The analysis results are summarized into a structured target form and returned to the user through the interaction layer, providing clear explanations and evidence; if the system decision also includes the execution of operation instructions, the selection and invocation of instructions are based on the Function Calling capability of the large model. On-demand multi-round iterative analysis: If the data or reference documents obtained in the current round of analysis are insufficient to support the analysis and decision-making process or require continuous source tracing analysis, the decision-making system will conduct several additional rounds of data and document acquisition until the decision-making system determines that the acquired information is sufficient to support the analysis and problem-solving.

7. The intelligent maintenance method for workflow scheduling based on a large model according to claim 6, characterized in that, The on-demand multi-round iterative analysis includes the following steps: First round of analysis: Based on the instruction content, the current running status of the scheduling system or workflow, and the operation and maintenance manual, preliminary analysis results and action plans are generated; Iterative deepening: Call the data interface layer to obtain supplementary data, verify or revise the analysis results and action plan; Convergence Decision: When the confidence level reaches the threshold, output the final maintenance strategy, operation instructions, and total user feedback; Execution feedback: After the operation instructions are executed at the interface execution layer, the system status after execution will be used to determine whether the expected goal is met. If the expected goal is not met, the system will be iterated and deepened again.

8. A workflow scheduling intelligent maintenance method based on a large model according to any one of claims 1 to 7, characterized in that, The analysis results include: Root cause analysis: Based on the upstream dependencies of the current workflow, a dependency DAG graph is constructed. By combining the execution status, completion time and execution logs of each workflow, the root cause workflow of the failure or delay and its scope of impact are located through large model reasoning. Predictive maintenance recommendations: Based on workflow dependencies, historical runtime and time consumption trends, output workflow delay warnings and expected workflow end times; Fault recovery strategy: Based on the upstream dependencies of the current workflow, workflow task priorities and operation guidelines, the system outputs key recovery task paths and priority-based recovery strategies through large-scale model inference and analysis, thereby improving the task recovery efficiency of the scheduling system. Configuration optimization strategy: Based on workflow dependencies, historical runtime, and cluster resources, the system uses large-scale model reasoning to analyze critical paths and identify unreasonable timed schedules. The system then adjusts the timing sequence of workflow schedules to improve cluster resource utilization and optimize workflow data output time.

9. A workflow scheduling intelligent maintenance system based on a large model, characterized in that, The system includes: The instruction input unit is used to receive input instructions through a dialog box, or to send corresponding input instructions when the system status triggers a set condition; The data acquisition unit is used to leverage the Function Calling capability of the large model decision layer to connect to the big data workflow scheduling platform through the data interface layer, thereby obtaining real-time running data of the workflow and connecting to the knowledge base to obtain operation guidance manuals. The data analysis unit is used to input the real-time running data, the operation guide manual, and the input instructions into the large model decision layer, and use the semantic understanding, reasoning, causal reasoning, and predictive reasoning of the large model to analyze the workflow status or scheduling system status, and then generate analysis results corresponding to the input instructions. The reanalysis unit is used to trigger additional rounds of data acquisition and in-depth analysis if the information required for analysis is insufficient or if a deeper tracing is needed, until the large model deems the information sufficient, and then the analysis results are obtained. The decision execution unit is used to select and call the target operation and maintenance interface encapsulated by the execution interface layer based on the function calling capability and the decision result corresponding to the analysis result of the large model decision layer, execute several corresponding operation instructions, and obtain the execution result, so as to realize the automated closed loop of maintenance operation. The feedback interaction unit is used to feed back the execution results and scheduling system workflow status to the user through the interaction layer.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 8.