Non-technician-oriented script intelligent adaptation and visual arrangement method and system

Through the intelligent script adaptation and visual orchestration method for non-technical personnel, the problem of non-technical personnel's dependence on script environment and high technical threshold for orchestration tools is solved, and the compatibility and reuse efficiency of scripts in different environments is improved, so that they can independently complete the deployment and operation of automated processes.

CN120085847AActive Publication Date: 2025-06-03BEIJING YULORE INNOVATION TECH

Patent Information

Application Number
CN202510561683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

After non-technical personnel generate scripts through AI tools, they face the problem of environmental dependence and are unable to independently complete the entire process from script generation to deployment and operation. They lack intelligent path identification and correction capabilities, the technical threshold for process orchestration tools is high, and the script reuse capabilities are limited.

Method used

It provides a smart script adaptation and visual orchestration method for non-technical personnel, including dependency library analysis, static code conversion, input and output feature analysis and visual orchestration, forming a DAG definition of the workflow, and performing function-level packaging and interface standardization processing to realize reusable modules and their version management.

Benefits of technology

It realizes automatic parsing of script dependencies, builds an isolated execution environment, ensures the compatibility of scripts in different environments, reduces the technical threshold of orchestration tools, improves script reuse efficiency, and enables non-technical personnel to independently complete the entire process from script generation to production-level automation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085847A_ABST
    Figure CN120085847A_ABST
Patent Text Reader

Abstract

The invention provides a non-technician-oriented script intelligent adaptation and visual arrangement method and a non-technician-oriented script intelligent adaptation and visual arrangement system. The method comprises the following steps: acquiring a script file uploaded by a user, extracting a dependency library through an AST analysis technology, and constructing an isolation container environment; performing intelligent identification and conversion on a file path and database connection in the script to generate a standardized API call; carrying out input and output analysis on the basis of the converted script, and carrying out DAG workflow arrangement in a visual mode; performing function-level packaging and standardization processing on the script component in the workflow to realize component multiplexing and version management; and finally, configuring a scheduling strategy and carrying out automatic operation monitoring. According to the method, a script environment depends on automatic processing and intelligent path adaptation, and visual arrangement and modular multiplexing are combined, so that non-technical personnel can independently complete the whole process from AI script generation to a production-level automatic process, the problem of'last kilometer 'in script deployment is solved, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to technical fields such as artificial intelligence, automated script processing, visual workflow orchestration, container environment management, and script asset reuse. In particular, it relates to a method and system for intelligent adaptation and visual orchestration of scripts for non-technical personnel. Background Art

[0002] With the rapid development of artificial intelligence and large language model technologies, non-technical personnel can already create various script programs with the help of AI generation tools, which theoretically greatly improves work efficiency. This technical field mainly involves multiple aspects such as script automation, low-code platforms, process orchestration, and containerization technologies.

[0003] In the current technical environment, common automation tools such as Apache Airflow provide workflow definition and scheduling capabilities, CI / CD tools such as Jenkins support continuous integration and deployment of scripts, and interactive environments such as Jupyter Notebook allow users to write and execute code on a web interface. These tools have achieved good application effects in their respective fields.

[0004] However, when non-technical personnel generate scripts through AI tools, they often encounter environment dependency problems. For example, the script requires a specific version of the Python library to run, or contains hard-coded local file paths and database connection information. Existing technologies such as virtualenv and Docker can create isolated environments, but they require manual identification of dependencies and manual configuration, and cannot dynamically resolve cross-script dependency version conflicts. In addition, these tools usually require users to master configuration syntax such as YAML / JSON to create task dependencies and lack intuitive visual debugging capabilities.

[0005] The main defects of the current technical solutions are as follows: First, there is a "last mile" problem between the AI-generated script and the actual running environment, resulting in non-technical personnel being unable to independently complete the entire process from script generation to deployment and operation; second, there is a lack of intelligent path recognition and correction capabilities, and it is unable to automatically adapt to different storage environments; third, the technical threshold of the process orchestration tool is high, and it is difficult for non-technical personnel to achieve the visual construction of complex workflows; fourth, the script reuse ability is limited. The scripts developed by technical personnel usually exist in an overall form and cannot be split into function-level modules, resulting in the need for repeated development for similar requirements. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for intelligent adaptation and visual orchestration of scripts for non-technical personnel, to solve the problem in the prior art that non-technical personnel are unable to independently complete the entire process from AI-generated scripts to production-level automation processes.

[0007] To achieve the above object, the present invention provides a method for intelligent adaptation and visual orchestration of scripts for non-technical personnel, including: Obtain the script file uploaded by the user, perform dependency library analysis on the script file, and obtain an isolated container environment; Utilize the isolated container environment to perform static analysis and conversion on the file operation statements and database connection codes in the script file, and generate standardized API calls; According to the standardized API calls, perform input-output feature analysis and visual orchestration to form a DAG definition of the workflow; Perform function-level encapsulation and interface standardization processing on the script components in the DAG definition of the workflow to obtain reusable modules and their version management schemes; Based on the DAG definition of the workflow and in combination with the version management scheme of the reusable modules, configure the workflow scheduling strategy and perform automated operation and monitoring, and output the execution status and exception handling results of the workflow.

[0008] Preferably, the performing dependency library analysis on the script file to obtain an isolated container environment includes: Perform abstract syntax tree (AST) analysis on the script file, extract import statements and from-import statements, and obtain third-party library reference information; Construct a dependency relationship graph based on the third-party library reference information and perform version conflict detection. When a version conflict is detected, adopt a dependency isolation strategy to generate different isolated environment configurations; Utilize the isolated environment configuration to call the container API to build an image and store it in a private repository to form the isolated container environment.

[0009] Preferably, the performing static analysis and conversion on the file operation statements and database connection codes in the script file to obtain standardized API calls includes: Adopt abstract syntax tree (AST) analysis and regular expression matching to process the script file, and generate a file operation list and a database access list; Perform code pattern recognition and classification according to the file operation list and the database access list to form a conversion rule mapping; Utilize the conversion rule mapping to replace the file operation statements and database connection codes with platform standard interfaces to obtain the standardized API calls.

[0010] Preferably, the performing input-output feature analysis and visual orchestration according to the standardized API calls to form a DAG definition of the workflow includes: Analyze the variable definitions and function return values in the standardized API calls to generate the input and output feature representations of the script components; Configure the node parameters and dependency relationships in a visual drag-and-drop manner according to the input and output feature representations to form an orchestration result; Perform connection validity verification and format conversion on the orchestration result, and output the DAG definition of the workflow.

[0011] Preferably, configure the workflow scheduling policy, perform automated operation and monitoring, and output the execution status and exception handling results of the workflow, including: Based on the DAG definition of the workflow and the version management scheme of the reusable module, obtain the workflow configuration information, convert the execution time rule into a timing expression, and obtain the scheduling policy; Collect the operation status data of the workflow nodes based on the scheduling policy to obtain the execution progress information; Perform exception detection and recovery policy execution based on the execution progress information to obtain the processing result and notification message.

[0012] Preferably, the method further includes: Obtain the historical execution data of the workflow, extract the time dimension features and resource dimension features, and obtain the workflow feature vector; Group the workflows using the density clustering algorithm based on the workflow feature vector to obtain workflow groups with similar execution characteristics; Generate a pseudo-random detection point sequence and collect status snapshots based on the workflow groups to obtain anomaly scores.

[0013] Preferably, the method further includes: Perform a multiplication graph model conversion on the DAG definition of the workflow to obtain the node mapping relationship; Based on the node mapping relationship, use the spanning tree covering algorithm to find the optimal connection path to obtain real-time connection suggestions; Based on the real-time connection suggestions, use the virtual force model to perform layout optimization to obtain a display effect with reduced wire crossings.

[0014] Preferably, the method further includes: Obtain the DAG definition of the workflow, construct a graph representation labeled with function labels and attribute markers to obtain a multi-level workflow model; Perform pattern recognition using the connected maximum common subgraph algorithm based on the multi-level workflow model to obtain shared patterns; Calculate the importance scores and perform standardized conversion based on the shared patterns to obtain new reusable modules.

[0015] Preferably, function - level encapsulation and interface standardization are performed on the script components in the DAG definition of the workflow to obtain reusable modules and their version management solutions, including: Function splitting and interface definition are performed based on the script components to obtain standardized function modules; Based on the standardized function modules, metadata descriptions of function classification and version information are created to obtain a component market interface; Based on the component market interface, version updates are tracked and the impact of compatibility is evaluated to obtain upgrade suggestions.

[0016] The present invention also provides a script intelligent adaptation and visualization orchestration system for non - technical personnel, including: A dependency analysis module, which is used to obtain the script file uploaded by the user, perform a dependency library analysis on the script file, and obtain an isolated container environment; A path adaptation module, which is used to perform static analysis and conversion on the file operation statements and database connection codes in the script file based on the isolated container environment to obtain standardized API calls; A workflow orchestration module, which is used to perform input - output feature analysis and visualization orchestration based on the standardized API calls to obtain a DAG definition of the workflow; A module management module, which is used to perform function - level encapsulation and interface standardization on script components based on the DAG definition of the workflow to obtain reusable modules and their version management solutions; A scheduling and execution module, which is used to configure a workflow scheduling strategy, perform automated operation and monitoring based on the DAG definition of the workflow and in combination with the version management solution of the reusable modules, and output the execution status and exception handling results of the workflow.

[0017] The present invention has the following beneficial effects: 1. An automated dependency management system based on AST static analysis and container technology can automatically parse script dependencies without manual intervention and build an isolated execution environment, solving the environmental configuration problems faced by non - technical personnel.

[0018] 2. A file path intelligent recognition and conversion mechanism combining a rule engine and an AI large - model automatically replaces local file paths and database connections with platform - standard APIs to ensure the compatibility of scripts in different environments.

[0019] 3. A visual DAG workflow orchestration system can achieve zero - code concatenation of complex script processes by dragging and dropping, and automatically generate the configuration code required by the underlying execution engine.

[0020] 4. Method for defining function - level script modularization and standardized interfaces, supporting version control and compatibility management of script components, and realizing efficient reuse of technical assets.

[0021] 5. Full - link code - free solution for integrated environment adaptation, path conversion, process orchestration, and automatic scheduling, enabling non - technical personnel to independently complete the whole process from AI - generated scripts to production - level automated processes. Brief Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of the method for intelligent script adaptation and visual orchestration for non - technical personnel provided by the embodiments of the present invention; Figure 2 It is a flowchart of the method for analyzing dependency libraries of script files provided by the embodiments of the present invention; Figure 3 It is a flowchart of the static analysis and conversion of file operation statements and database connection codes provided by the embodiments of the present invention; Figure 4 It is a flowchart of the visual workflow orchestration provided by the embodiments of the present invention; Figure 5 It is a flowchart of the function - level encapsulation and interface standardization processing of script components provided by the embodiments of the present invention; Figure 6 It is a flowchart of the automatic scheduling and monitoring provided by the embodiments of the present invention; Figure 7 It is a flowchart of workflow grouping and anomaly detection based on historical data provided by the embodiments of the present invention; Figure 8 It is a flowchart of workflow layout optimization based on a doubling graph provided by the embodiments of the present invention; Figure 9 It is a flowchart of the modularization method based on common sub - graph mining provided by the embodiments of the present invention; Figure 10 It is a block diagram of the system for intelligent script adaptation and visual orchestration for non - technical personnel provided by the embodiments of the present invention. Detailed Embodiments

[0024] The following will further elaborate on the present invention in combination with the drawings and embodiments.

[0025] As Figure 1As shown below, an embodiment of the present invention provides a method for intelligent adaptation and visual orchestration of scripts for non-technical personnel, including the following steps: Step S1: Obtain the script file uploaded by the user, perform a dependency library analysis on the script file, and obtain an isolated container environment; In an actual application scenario, step S1 first receives the Python script file uploaded by the user through a Web interface. The user only needs to directly drag the script generated by the AI tool to the upload area provided by the system, and the system will immediately perform file format verification and security checks to ensure that only legitimate script files are received. After passing the verification, the script file will be stored in the distributed file system, and at the same time, the system will create a script metadata record in the database, including basic information such as script ID, name, upload time, user ID, etc., to establish an index for subsequent dependency analysis and processing.

[0026] Next, use the AST (Abstract Syntax Tree) library of Python to deeply analyze the uploaded script. AST analysis can parse Python code into a tree structure, enabling the system to accurately identify the import statements and from...import statements in the code. For example, when analyzing statements such as "import pandas as pd" or "from numpy import array", the system will identify "pandas" and "numpy" as the required third-party libraries. At the same time, it will also search for comments and strings in the script through regular expressions to find version constraint information such as "#requires tensorflow>=2.0.0". In this way, all third-party library dependencies required by the script and their version requirements can be comprehensively extracted to generate a complete dependency list.

[0027] According to the extracted dependency list, construct a dependency relationship graph to analyze the mutual dependencies and potential version conflicts between libraries. For example, when multiple scripts in a workflow require different versions of the same library (such as one script requires pandas 0.25 while another requires pandas 1.3), the system will detect this conflict and automatically decide whether to create an isolated environment. For cases where version conflicts are detected, the system adopts a dependency isolation strategy to create independent runtime environment configurations for incompatible script groups to ensure that each script can be correctly executed in a suitable environment.

[0028] Subsequently, a Dockerfile is automatically generated based on the dependency analysis results. This file defines all the steps required to build a container image. The Dockerfile usually starts from a base Python image and then adds instructions for installing necessary system dependencies and Python libraries. For example, for a script that requires data processing capabilities, the system will generate installation commands for libraries such as pandas and numpy; for a script that requires machine learning capabilities, instructions for installing scikit-learn or tensorflow will be added. The system automatically executes the build process through Docker API calls to generate a container image containing all necessary dependencies. The built image is pushed to a private repository, and the mapping relationship between the image ID and the script is recorded in the system database to form a reusable container environment.

[0029] To improve efficiency, a mapping table of dependency configurations and container images is also maintained. When a new script is uploaded, the system compares its dependency configuration with the existing environments. If it is found that the existing environment already contains all necessary dependency libraries, the environment will be directly reused instead of rebuilding the image. This caching and reuse mechanism significantly improves the efficiency of script deployment, reduces the consumption of system resources, and at the same time ensures the consistency and reliability of the environment. In actual use, as the number of scripts in the system increases, the environment reuse rate will gradually increase, and most newly uploaded scripts can find compatible existing environments, thus achieving near-instant deployment capabilities.

[0030] Step S2: Utilize the isolated container environment to perform static analysis and transformation on the file operation statements and database connection codes in the script file to generate standardized API calls; In step S2, first, in the isolated container environment created in step S1, a deep static analysis of the script file is performed. The analysis process combines the use of AST (Abstract Syntax Tree) technology and regular expression matching to comprehensively scan the file operation statements and database connection codes in the script. For file operations, the system can identify common Python file handling functions, such as basic operations like open(), read(), write(), etc., as well as data file handling functions such as pd.read_csv(), pd.read_excel(), pd.to_csv() in the pandas library. The file path strings and operation modes (read / write) in these function calls will be extracted to determine the access method and location of the file. For example, when the code "with open('C: / Users / data / input.txt', 'r') as f:" is detected, the system will recognize that this is a local file reading operation and record the file path "C: / Users / data / input.txt".

[0031] Similarly, the database connection code in the script will also be recognized. Different database access libraries have different connection patterns, such as the "create_engine" function of SQLAlchemy, the "connect" method of pymysql, the connection string of psycopg2, etc. The system has designed recognition rules for various common database connection patterns and can accurately extract key information such as the database type (e.g., MySQL, PostgreSQL, SQLite, etc.), connection address, username, password, database name, etc.

[0032] Based on the generated manifest, code pattern recognition and classification will be performed. A conversion rule library for file operations and database connections is maintained, and corresponding conversion strategies are defined for different types of operations. For example, for local file paths, they will be converted to the platform-standard file access API; for database connection strings, they will be converted to use the platform-unified data source configuration. These rules are organized into a conversion rule mapping table, and each operation mode has a corresponding conversion template. When a file read operation is recognized in the script, an appropriate conversion rule will be selected based on the file type and read method, and the hard-coded local path will be replaced with the platform path.

[0033] When encountering complex situations that cannot be covered by the rule library, the built-in AI model will be called for intelligent analysis. The AI model has been specially trained to understand the code context and intent and infer the most appropriate alternative. For example, when the script contains complex file path construction logic (such as path concatenation, environment variable reference, etc.), the AI model can analyze the entire logic chain and propose multiple possible alternatives. The system will sort these alternatives based on confidence, select the best alternative for replacement, or request user confirmation when necessary. This hybrid intelligent method enables the system to handle various complex file operation and database connection scenarios, not limited to simple pattern matching.

[0034] After applying the conversion rules, the recognized file operation statements and database connection code will be replaced with platform-standard API calls. For example, "open('C: / data / file.csv', 'r')" is converted to "get_platform_file('user_files / file.csv', 'r')", or the direct database connection string is replaced with "get_database_connection('database_alias')". These standard APIs encapsulate the platform's unified access logic for files and databases, ensuring that scripts can correctly access the required resources in the platform environment without relying on specific local paths or hard-coded connection information. Through this conversion, scripts that could only run in a specific environment become portable and can be stably executed in any container environment provided by the platform.

[0035] When performing code conversion, the mapping relationship between the original code and the modified code is saved and all modification contents are visually displayed in the user interface. Users can review each modification and choose to accept or reject specific changes. The system provides a code comparison view, using colors to mark the added, deleted, and modified parts, enabling users to clearly understand each change. For non-technical personnel, the system also provides simplified explanations, describing the purpose and impact of each change in plain language. The system also provides a complete modification history, supporting rollback to previous versions, ensuring the controllability and transparency of the code modification process. This interactive conversion process not only ensures the convenience of automation but also retains the user's control over key decisions, especially suitable for non-technical personnel to handle AI-generated scripts. Through this step, the system solves the environmental dependency problem in the script, laying a foundation for stable operation on the platform in the future.

[0036] The AI model in the system is a large language model based on the Transformer architecture, specifically fine-tuned for the code conversion task. The model adopts an encoder-decoder structure. The encoder consists of 12 layers of self-attention layers, each layer containing 8 attention heads, and the hidden layer dimension is 768; the decoder structure is similar, but a cross-attention mechanism is added to focus on the key parts of the input code. The model is trained using a dataset consisting of 2 million pairs of code conversion samples, which cover various conversion modes from local file paths to platform APIs, from direct database connections to data source abstraction APIs. The training process uses the teacher forcing strategy, with a cross-entropy loss function with label smoothing, the AdamW optimizer (learning rate of 3e-5, weight decay of 0.01), and is trained distributively on 8 GPUs for a total of 5 epochs.

[0037] The model is evaluated using three metrics: the conversion accuracy rate (exact match score) reaches 87.3%, the functional equivalence rate (execution result consistency) reaches 93.5%, and the average edit distance is 4.2. In actual deployment, the model is stored in the quantized int8 format, significantly reducing the model size while only losing less than 1% of the accuracy. During the code conversion process, the model receives the original code snippet and its context as input, generates multiple possible conversion results, and calculates a confidence score for each result. The system preferentially selects the result with the highest confidence, but when the confidence of all results is lower than the threshold (0.85), multiple candidate conversions are presented to the user for selection. This human-machine collaboration method ensures the accuracy and controllability of code conversion.

[0038] Step S3: Based on the standardized API calls, perform input-output feature analysis and visual orchestration to form a DAG definition of the workflow; In step S3, first, conduct an in-depth analysis of the standardized API calls generated in step S2 to understand the input-output features of the script. This analysis process focuses on the functional essence of the script rather than just its surface code. It checks the variable definitions, function parameters, return values, and the patterns of standardized API calls in the script to infer the possible input parameters required by the script and the output results it will produce. For example, when it detects that there are unassigned variables at the beginning of the script that are subsequently used, the system will identify them as possible input parameters; when it finds that certain variables are returned or written to a file at the end of the script, they will be marked as output results. The system also analyzes the data flow, traces the entire processing path of variables from input to output, and identifies key data conversion steps and branch points.

[0039] For some complex scripts, automatic inference may not be accurate enough, and a function for users to manually define interfaces is provided. Users can specify through an intuitive interface which variables should be used as input parameters for the script and which should be used as output results. They can also add information such as parameter types, default values, and documentation. A friendly form interface is provided to guide users to complete the interface definition step by step. Users do not need to understand complex programming concepts and only need to answer simple questions such as "What input data does this script require?" and "What results will be produced after running?" The system will automatically convert this information into a technical definition. This human-machine combination method ensures the accuracy of the script interface definition and lays a foundation for correctly connecting script components in the subsequent workflow.

[0040] Through analysis, standardized input and output feature representations are generated for each script, including information such as parameter names, data types, data formats, and whether they are required. These feature representations are stored in JSON format for easy internal processing within the system and interaction with the front-end interface. For example, the feature representation of a data preprocessing script may include detailed descriptions of the input parameter "data_file" (CSV file type) and the output result "processed_data" (DataFrame type). Based on these feature representations, a visual representation of the script is automatically generated, including the appearance of graphical nodes, the display of input and output interfaces, and interaction behaviors. This ensures that the visual elements seen by the user accurately reflect the actual functions and interface requirements of the script.

[0041] Based on these input and output feature representations, a powerful visual orchestration interface is provided. This interface is built using modern web technologies, based on the React framework and SVG / Canvas drawing capabilities, to achieve a smooth drag-and-drop workflow design experience. The interface layout is clear. On the left is the component library panel, which displays available script components, organized by function category and supporting search and filtering. In the center is the workspace where users can drag and drop components and connect data flows. On the right is the property panel, which displays detailed information and configurable options for the selected component. The entire interface uses an intuitive visual design with soft color coding and clear icons, making it easy for non-technical users to understand and operate.

[0042] Users can drag and drop script components from the component library into the workspace. Each component is represented by an intuitive graphical node, on which the component name and main function are displayed. The size, shape, and color of the nodes are carefully designed to facilitate user identification of different types of components. For example, data source components may use blue, transformation components use green, and output components use orange. Small icons are also displayed on the nodes to intuitively represent the functional types of the components, such as database icons, file icons, or algorithm icons. When a user selects a node, the system displays the detailed information and configurable parameters of the component, and the user can set these parameters through a form interface without writing any code. The system provides intelligent parameter setting assistance, such as data type checking, value range verification, and auto-completion, to help users avoid common errors.

[0043] The visualization interface of the system supports specifying the data flow direction through connection lines. Users only need to drag from the output port of one component to the input port of another component to create a data transfer relationship between components. During the connection process, the system provides visual feedback, such as highlighting compatible ports and previewing the connection path. When creating a connection, the compatibility of data types is automatically checked to ensure that only ports with matching types can be connected. For example, a component that outputs a CSV file can be connected to a component that accepts CSV input, but cannot be directly connected to a component that requires JSON input. When a user attempts to connect incompatible ports, the system displays a friendly error message and suggests possible solutions, such as adding a format conversion component. This real-time verification mechanism greatly reduces errors in workflow design and improves the work efficiency of non-technical personnel.

[0044] In addition to basic nodes and connections, advanced workflow control structures are also supported, such as conditional branching (if-else logic), loop execution (foreach loop), and parallel tasks. These control structures are also represented by intuitive graphical elements, and users can implement complex business logics through simple dragging and configuration. For example, users can add a conditional branching node, set the conditional expression "data.rows>1000" through a drop-down menu or a simplified expression editor, and then connect different processing paths to the true and false branches respectively. The system provides a conditional preview function, and users can input sample data to view the evaluation results of the conditions to ensure that the logic meets the expectations. These advanced control structures enable non-technical personnel to create complex workflows and achieve true business automation.

[0045] When the user completes the workflow design, the validity of the entire workflow will be verified to check for issues such as unconnected required inputs and circular dependencies. The verification results are presented as clear messages, indicating the specific problem locations and repair suggestions. For complex workflows, the "workflow checker" function is provided to automatically analyze the structure and data flow of the entire workflow and provide a comprehensive quality report, including potential performance bottlenecks, error risk points, and optimization suggestions. After passing the verification, the system converts the visual workflow created by the user on the interface into a standardized DAG (Directed Acyclic Graph) definition format, usually in JSON or YAML format. This DAG definition details each node (script component) in the workflow, the dependencies between nodes, execution conditions, parameter configurations, etc., and is a complete technical representation of the workflow. To ensure portability and compatibility, the system adopts open standard formats, facilitating integration with other tools and platforms.

[0046] Finally, according to the DAG definition, configuration code suitable for the underlying execution engine (such as Airflow, Prefect, or a custom executor) is automatically generated. This code includes task definitions, dependency settings, and parameter passing logic to ensure that the workflow can be correctly parsed and run by the execution engine. The generation process takes into account the characteristic differences of different execution engines to ensure that the generated code can achieve the best performance of each engine. The entire process is completely transparent to the user. Non-technical personnel do not need to understand the details of the underlying execution engine to create a professional-level workflow automation solution. Through this visual orchestration method, the system has successfully transformed complex technical work into intuitive graphical operations, enabling non-technical personnel to independently create and manage complex data processing and automation workflows.

[0047] Step S4: Perform function-level encapsulation and interface standardization on the script components in the DAG definition of the workflow to obtain reusable modules and their version management solutions; In step S4, efforts are made to transform the script components in the workflow into reusable standardized modules to achieve the accumulation and sharing of technical assets within the organization. This step first deeply analyzes the script components in the workflow DAG definition to identify code blocks with relatively independent functions and reuse value. Check the functional boundaries, input and output interfaces, and internal logic of the components to evaluate their independence and generality. For those components that implement a clear and single function and have a wide range of application scenarios, the system will mark them as potential reusable modules. This analysis combines static code features and usage scenario statistics to intelligently discover the most valuable reuse opportunities.

[0048] The system will guide technical personnel (or with AI assistance) to refactor these code blocks into independent function modules, following good software engineering practices, such as the single responsibility principle, clear input and output interfaces, etc. The refactoring process focuses on improving the readability, maintainability, and reusability of the code while maintaining the integrity of the original function. The system provides code refactoring suggestions, pointing out potential optimization points, such as parameter merging, logic separation, or naming improvement. For scripts created by non-technical personnel, the system can automatically perform basic refactoring with AI assistance to improve the code quality and make it more suitable for sharing and reuse in an enterprise environment.

[0049] During the function-level encapsulation process, it helps developers add standardized function signatures and detailed docstrings to each function. The function signature clearly defines the parameter list, type annotations, and return value type, making the usage of the function clear at a glance. The docstring contains the function's functional description, parameter explanations, return value interpretations, and usage examples, following industry-standard documentation formats. The system provides a documentation generation tool that can extract information from code analysis and usage scenarios to assist in creating comprehensive function documentation. Such detailed documentation enables non-technical personnel to understand the purpose and usage method of the function, greatly reducing the technical threshold. The encapsulated function module retains its original functionality but has better structure and readability, facilitating reuse in different workflows.

[0050] After completing the function-level encapsulation, create structured metadata descriptions for each script component. These metadata are stored in JSON format and contain rich attribute information, such as functional classification (e.g., data processing, machine learning, report generation, etc.), applicable scenario descriptions, detailed explanations of input parameters (including type, format requirements, value range, default values, etc.), output format and structure descriptions, performance characteristics (typical execution time, resource consumption, etc.), author and maintainer information, version history and change logs, other components or libraries it depends on, and usage examples and best practices. These metadata are not only used for component retrieval and display but also support the system in performing intelligent recommendations and compatibility checks.

[0051] The metadata creation process is semi-automated. It extracts initial information from code analysis, usage history, and existing documents, and then guides users to supplement and improve it. The user interface is designed to be intuitive and friendly, guiding users step by step in a wizard form to fill in key information and providing intelligent suggestions and examples to ease the burden of documentation writing. It also supports collaborative editing, allowing multiple team members to jointly improve the component documentation. These metadata support multi-dimensional search and filtering, enabling users to quickly find the required components. For example, users can locate components that meet specific requirements by keyword search, browsing by category, or filtering based on input / output types.

[0052] Based on these standardized function modules and metadata, an intuitive component market interface is built. This interface displays available script components in a way similar to an app store, with each component having a thumbnail, brief description, and rating information. The home page of the component market shows popular components and newly added components, and recommends relevant components based on the user's historical usage preferences. Users can browse the component library through category navigation or search functions and view the detailed description page of the component, which shows the complete documentation, usage examples, performance metrics, and user reviews of the component. The component market also provides a component comparison function, allowing users to compare the functions and features of multiple similar components side by side to make the best choice.

[0053] Users can directly add the selected components from the component market to their own workflows, and the system will automatically handle the import and initial configuration of the components. The component market also supports user evaluation and review functions. Users can rate the components and share their usage experiences, provide improvement suggestions or report problems. These feedbacks form a community knowledge base, which helps other users select suitable components and also provides improvement directions for component authors. The system will calculate the popularity of components based on ratings and usage frequencies, increase the exposure rate of high-quality components, and form a virtuous cycle.

[0054] In terms of version management, a perfect component version control mechanism has been implemented. The version changes of each component are recorded in detail, including modification content, improvement points, known problems and compatibility descriptions. The system uses semantic version numbers to clearly indicate the compatibility relationship between versions. A change in the major version number indicates an incompatible API modification, a change in the minor version number indicates a backward-compatible new feature addition, and a change in the patch version number indicates a backward-compatible problem fix. This version management convention enables users to quickly understand the scope of impact of version upgrades.

[0055] Provide version release tools for component developers to guide them to follow the standard version management process. Developers need to provide a detailed change log, clearly indicating the improvement content and potential compatibility problems of the new version. When releasing a new version, a series of verification tests will be automatically executed to ensure the quality of the component and generate a difference report with the previous version. When the component interface changes, the system will automatically evaluate the compatibility impact of these changes on existing workflows and generate a detailed impact analysis report, indicating the workflows that may be affected and the necessary adjustments.

[0056] For incompatible version upgrades, migration suggestions are provided to guide users on how to adjust their workflows to adapt to the new version of the component. For example, if the new version of the component modifies the parameter names or adds required parameters, the system will specifically indicate the places that need to be modified and even provide an automatic migration option. Whenever possible, the system will generate adaptation code to automatically handle parameter format conversion or default value filling, minimizing manual adjustments by users. This intelligent version management ensures that workflows depending on this component will not be accidentally interrupted due to upgrades, greatly reducing the burden of version maintenance.

[0057] Intuitive parameter configuration forms are also generated for each component based on the parameter definitions in the component metadata. These forms include parameter descriptions, default values, value ranges and format validations, guiding users to set parameters correctly and preview the parameter effects in real time, reducing the possibility of configuration errors. For complex parameters, the system provides visual configuration assistance, such as date pickers, color pickers or data structure visual editors, further simplifying the operations for non-technical personnel. The system also supports the parameter template function, allowing users to save common parameter combinations as templates and quickly apply them in different workflows to improve work efficiency.

[0058] Through this function-level encapsulation and interface standardization process, the system effectively transforms scattered script codes into a structured, manageable, and reusable technology asset library, greatly enhancing the knowledge precipitation and technology reuse efficiency within the organization. As the asset library continues to grow, users can quickly build complex workflows by combining existing components, significantly reducing repetitive development and improving the business response speed. This modular approach also promotes the dissemination and standardization of best practices, enhancing the technical quality and consistency of the entire organization.

[0059] Step S5: Based on the DAG definition of the workflow and combined with the version management scheme of the reusable modules, configure the workflow scheduling strategy and perform automated operation and monitoring, and output the execution status of the workflow and the exception handling results.

[0060] In step S5, the workflow defined in the previous steps is transformed into an actual running automated process, and comprehensive monitoring and management capabilities are provided. First, the system builds a complete workflow configuration based on the DAG definition of the workflow and the version management information of the reusable modules. This configuration includes detailed execution parameters, resource requirements, dependency relationships, and error handling strategies for each node, ensuring that all required resources and component versions can be correctly obtained during the workflow execution. The system checks each component version used in the workflow to ensure their compatibility and resolve possible version conflicts. For some key components, the system also automatically adds version locking settings to prevent accidental version changes from affecting the workflow stability.

[0061] An intuitive timing policy setting interface is provided, supporting various scheduling requirements. The interface design is simple and clear, mainly divided into basic scheduling and advanced scheduling parts. In the basic scheduling part, users can use the graphical calendar view to select the execution time, such as "9 am on Mondays, Wednesdays, and Fridays" or "5 pm on the last working day of each month". The interface provides quick options for common scheduling patterns, such as "every day", "weekday", or "beginning of month", and users can apply them with just one click. For more complex requirements, users can switch to the advanced mode and use a more flexible time selector to specify complex repeating patterns. In either case, the system automatically converts these natural language-based time rules into standard Cron expressions, such as "09 * * 1,3,5" (9 am on Mondays, Wednesdays, and Fridays).

[0062] To ensure that users understand the effects of the settings, the specific time points of the next few expected executions will be displayed, presented in the form of a calendar view and a time list. This enables users to visually verify whether the scheduling settings meet expectations, especially for complex scheduling rules such as "the second Wednesday of each month" or "the last working day of each quarter". The system also provides a time zone selection function to ensure that scheduling times are consistent among users in different regions, avoiding scheduling errors caused by time zone differences.

[0063] In addition to basic scheduled execution, it also supports the initiation of workflows triggered by events, such as events like file upload completion, database update, or API callback. Users can select trigger conditions through simple drop-down menus and tabs, specifying the event source and the specific event type. For example, users can set the workflow to be triggered "when the order table in the sales database is updated" or "when a new file is received in the specified FTP directory". A rich set of predefined event types are provided, covering common business scenarios, and custom events are also supported to meet special requirements. For more technically demanding event configurations, the system provides visual assistance tools such as database table selectors or file path browsers to simplify the configuration process.

[0064] It also allows setting conditional execution rules to increase the intelligence of workflow initiation. Users can define preconditions, such as "this workflow will only be executed when the previous workflow has been successfully completed and the number of generated data rows is greater than 100". The condition setting interface uses an intuitive form where users can select the condition type (such as workflow status, data characteristics, time window, etc.) and specify the specific condition values through drop-down menus and input boxes. The system supports the combination of multiple conditions, and users can create complex conditional expressions using "AND" and "OR" logical connectors. For advanced users, the system also provides a conditional expression editor to support more complex logical definitions. These advanced scheduling functions enable automated processes to respond more intelligently to business needs, reduce unnecessary executions, and improve the utilization efficiency of system resources.

[0065] After configuration is completed, a complete scheduling policy will be generated and submitted to the underlying scheduling engine. The scheduling engine is responsible for starting the workflow at the specified time or when the event occurs and allocating appropriate computing resources. To optimize resource utilization, the system analyzes the historical execution data of the workflow, predicts the execution time and resource requirements of each node, and then, based on these analysis results and the current system load, intelligently allocates computing resources to avoid resource contention and improve the overall execution efficiency. The concept of a resource pool is implemented, and different resource quotas can be allocated to different workflows to ensure that critical business processes obtain sufficient resources while preventing a single workflow from occupying too many system resources.

[0066] After the workflow is started and executed, a comprehensive monitoring process begins. The monitoring system adopts a multi-level architecture, including infrastructure monitoring, container monitoring, process monitoring, and application layer monitoring, to comprehensively grasp the running status of the workflow. It tracks the execution status, running time, resource usage, and log output of each node in the workflow in real time. The running metrics of each node are stored in a time series database, supporting historical trend analysis and anomaly detection. These monitoring information are presented to users through an intuitive visualization interface, which adopts a responsive design to adapt to display devices of different sizes, facilitating users to view the workflow status anytime and anywhere.

[0067] Nodes on the workflow diagram will display the current status in different colors: nodes waiting to be executed are gray, nodes being executed are blue, nodes successfully completed are green, nodes that failed are red, and nodes that are paused or skipped have special markings. The node icons will also display a progress indicator, intuitively reflecting the completion percentage of long-running tasks. Users can clearly understand the progress of the entire workflow at a glance and quickly identify problem nodes. The system supports operations such as zooming in, zooming out, and panning of the workflow diagram, facilitating the viewing of large and complex workflows. For workflows with a large number of nodes, the system also provides node grouping and folding functions, allowing users to focus on specific parts and reducing visual complexity.

[0068] For nodes that are being executed, real-time performance metric monitoring is provided, such as CPU usage, memory occupancy, data processing speed, etc., as well as a detailed log viewing function. The performance metrics are presented in the form of charts, clearly showing the trend and fluctuations of resource usage. Users can expand the nodes to view more details, or click the log button to view the complete execution log. The log viewer supports real-time scrolling, keyword search, and log level filtering, helping users quickly locate key information. For nodes that process a large amount of data, the system also provides data sampling preview, allowing users to view the data samples currently being processed and understand the data quality and processing effect. These detailed information help users understand the execution of the workflow, especially quickly locating the cause when problems occur.

[0069] An intelligent anomaly detection and automatic recovery mechanism is also implemented. During the monitoring process, the system continuously analyzes the execution data to find possible anomalies, such as: the execution time of a node is abnormally extended, the resource usage suddenly increases, the output data volume deviates significantly from the historical average, etc. Anomaly detection is based on multiple technologies, including statistical analysis, time series prediction, and machine learning models. A normal behavior baseline will be established for each workflow, and then abnormal patterns deviating from this baseline will be identified. For example, it can be detected that a data processing node that usually takes 5 minutes to complete has been running for 15 minutes, but the progress is only 30%, which may indicate that there are problems with data processing.

[0070] When these anomalies are detected, corresponding measures will be automatically taken according to the preset strategies. The response measures are divided into multiple levels, from light to heavy, including: recording warnings but continuing to execute, attempting mild interventions (such as reallocating resources), performing recovery operations (such as restarting specific nodes), rolling back to a safe point, or completely terminating the workflow. Which measure to take specifically depends on the nature, severity, and importance of the anomaly. For nodes that fail to execute, the system can automatically retry, and the user can configure the number of retries, the interval time, and the maximum waiting time. When performance issues caused by insufficient resources are detected, the system can dynamically allocate more resources, such as increasing the container memory limit or boosting the CPU priority. In case of data anomalies, it can automatically switch to a predefined alternative processing path, such as using cached data or simplifying the processing logic.

[0071] For unrecoverable errors, the workflow will be terminated in a timely manner and relevant personnel will be notified to prevent the spread of errors. The termination process is controlled. It will attempt to complete the critical nodes that have already started to ensure data consistency, while canceling the tasks that have not yet started and releasing resources. These automatic recovery mechanisms greatly improve the reliability of the workflow, reduce the need for manual intervention, and are especially suitable for non-technical personnel to manage complex automated processes. Even when problems occur, the system can autonomously attempt to repair them and only request manual intervention when the problems exceed its automatic processing capabilities.

[0072] In terms of notification management, multi-level notification strategy configurations are provided. Users can define different notification rules at the workflow level and node level to precisely control which events need to be notified and how to notify. The notification levels are divided into multiple grades, such as information, warning, error, and critical error. Users can set different notification methods for each level. The notification channels include multiple options, such as in-app messages, emails, text messages, mobile app push notifications, and integrations with enterprise instant messaging tools (such as Slack or WeCom). Users can flexibly combine these channels. For example, ordinary warnings only send in-app notifications, while critical errors are sent via email, text message, and instant messaging simultaneously.

[0073] It also supports different notification strategies for working hours and non-working hours to avoid disturbing users' rest time in non-emergency situations. Users can define their working hours and time zones, and the system will adjust the notification behavior according to these settings. For example, during working hours, general warnings may be sent via in-app messages, while during non-working hours, only critical errors will trigger notifications, and high-priority channels such as text messages will be preferred. The system also implements a notification escalation mechanism. If no one responds to an important alert within a specified time, the system will automatically notify the supervisor or backup personnel to ensure that critical issues are handled in a timely manner. The notification content is also intelligently generated, including problem descriptions, affected scopes, possible causes, and recommended actions, enabling the recipients to quickly understand the situation and take actions.

[0074] The complete execution results, performance data, and exception information are recorded by the system in the database to form the execution history of the workflow. The historical data storage adopts a hierarchical design. The hot data is retained in the high-performance storage for immediate query, while the historical data is gradually moved to the archival storage to ensure the cost-effectiveness of long-term data. Users can view the execution records for any past period through the historical query interface, which provides rich filtering conditions such as time range, execution status, keywords, and resource usage. Users can compare the execution performance at different times, analyze long-term trends, and identify performance degradation or improvement. Multiple visualization charts are provided, such as execution time trend charts, resource usage heatmaps, and status distribution pie charts, to help users understand the execution patterns and identify anomalies. These historical data can also be exported in CSV or Excel format for external reporting or in-depth analysis.

[0075] The historical data is not only used for user queries but also provides valuable resources for the system's own optimization. The system continuously learns from this data to improve resource allocation strategies, anomaly detection models, and prediction accuracy. For example, by analyzing the historical execution data, the system can more accurately predict the running time and resource requirements of the workflow and optimize the scheduling decisions. The system also identifies performance bottlenecks and common failure points in the workflow and generates optimization suggestions to help users improve the workflow design. Through this continuous learning and optimization, the system can become increasingly intelligent over time, providing more accurate monitoring and more efficient resource management.

[0076] Through this complete scheduling, execution, and monitoring system, non-technical personnel can manage complex automated workflows, ensure that the processes run reliably as planned, detect and solve problems in a timely manner, and ultimately achieve stable automation of business processes. The design concept of the system is to hide the complex technical details behind an intuitive user interface, allowing users to focus on business goals rather than technical implementation. This approach greatly reduces the technical threshold of automation, enabling more business personnel to participate in process automation and accelerating the digital transformation process of the organization.

[0077] In a preferred embodiment of the present invention, as Figure 2 shown, in step S1, the script file is analyzed for dependency libraries to obtain an isolated container environment, including: Step S11: Perform abstract syntax tree (AST) analysis on the script file to extract import statements and from-import statements to obtain third-party library reference information; Specifically, the deep AST analysis is implemented based on the Python's ast module, which can traverse the entire Python syntax tree structure. The following is the core implementation logic of the AST analysis: First, parse the Python script into an AST tree, and then implement a custom NodeVisitor class to traverse the entire tree structure, specifically identifying Import and ImportFrom nodes. For each Import node, extract the module name; for ImportFrom nodes, record both the module name and the specific functions or classes being imported. It can also identify alias imports (such as import numpy as np) and correctly associate them with the original library name. In addition to directly parsing import statements, the system also uses regular expressions to identify version requirement information in comments, such as # requires pandas>=1.0.0 or # dependency: tensorflow==2.4.0, etc. For complex conditional imports (such as imports in try-except blocks or runtime dynamic imports), a combination of static analysis and heuristic rules is adopted to achieve a dependency recognition accuracy of over 95%.

[0078] Step S12: Construct a dependency graph based on the third-party library reference information and perform version conflict detection. When a version conflict is detected, adopt a dependency isolation strategy to generate different isolated environment configurations; Specifically, the system implements a directed graph-based dependency relationship model, using an adjacency list to represent the dependency relationship between libraries. During the graph construction process, nodes represent libraries, edges represent dependency relationships, and each edge also contains version constraint information. The system uses a semantic version parser to process version strings, supporting complex version constraint expressions such as "> = 1.2.0, <2.0.0". The version conflict detection algorithm is based on depth-first search, starting from the entry script to traverse the dependency graph and maintaining a global version constraint set. When an incompatible version constraint for a certain library is detected, record the conflict source and constraint conditions to provide a decision basis for subsequent dependency isolation.

[0079] Adopt an optimized "connected component analysis" algorithm to divide incompatible library groups. This algorithm first represents all mutually incompatible version constraint relationships as an undirected graph, where nodes are "library + version constraint" and edges represent conflict relationships. Then apply a variant of the Tarjan algorithm to find all strongly connected components, and each component represents a set of libraries that need to be isolated. This method can isolate the minimum number of library groups in different environments while maximizing library sharing, optimizing the image size and build time.

[0080] Step S13: Use the isolated environment configuration to call the container API to build an image and store it in a private repository to form the isolated container environment.

[0081] Specifically, in step S11, the AST library of Python is used to perform abstract syntax tree analysis on the script uploaded by the user, identify import statements and from...import statements, and extract all third-party library references. At the same time, the comments and strings in the script are analyzed through regular expressions to find possible version requirement information (such as #requires numpy>=1.20), and a complete list of dependencies and version constraints is output.

[0082] In step S12, according to the third-party library reference information, a dependency graph is constructed to detect potential version conflicts. When a conflict is detected (for example, one script requires pandas<1.0 while another requires pandas>=2.0), the system adopts a dependency isolation strategy to create independent runtime environment configurations for each group of incompatible scripts to ensure that all scripts can run in a suitable environment.

[0083] In step S13, based on the dependency analysis results, a Dockerfile is dynamically generated. Starting from the base Python image, necessary system dependencies and Python library installation instructions are added. The image is built by calling the Docker API, and after the build is completed, the image is pushed to a private repository and the mapping relationship between the image ID and the script is recorded to form a reusable container environment. In addition, the system also maintains a mapping table between the dependency configuration and the container image. When a new script is uploaded, the compatibility between its dependency configuration and the existing environments is compared. If a compatible existing environment is found, that environment is directly reused to avoid repeated builds, significantly improving the script deployment efficiency and saving system resources.

[0084] In a preferred embodiment of the present invention, as Figure 3 shown, in step S2, static analysis and conversion are performed on the file operation statements and database connection codes in the script file to obtain standardized API calls, including: Step S21: Use abstract syntax tree AST analysis and regular expression matching to process the script file to generate a file operation list and a database access list; Step S22: Perform code pattern recognition and classification according to the file operation list and the database access list to form a conversion rule mapping; Step S23: Use the conversion rule mapping to replace the file operation statements and database connection codes with platform standard interfaces to obtain the standardized API calls.

[0085] Specifically, in step S21, the system uses a combination of AST analysis and regular expression matching to scan the file operation statements in the script, such as open(), pd.read_csv(), pd.read_excel(), etc., extract the file path strings and operation modes (read / write), and establish a list of script file operations. At the same time, the system identifies common database connection patterns in the script, including connection strings of libraries such as SQLAlchemy, pymysql, and psycopg2, extracts information such as database type, address, and username, and generates a database access list.

[0086] The system's static analyzer combines two techniques: AST analysis and regular expressions. For AST analysis, the system specifically identifies function calls related to file operations, including the built-in open function, functions in the os.path module, pathlib.Path operations, and file read / write functions of common data processing libraries (such as pandas.read_csv, numpy.load, etc.). For each identified function call, the system extracts relevant parameter information, especially the file path string and operation mode. The system can handle complex parameter passing patterns, including keyword arguments, positional arguments, default arguments, etc.

[0087] For database connection identification, a rule library containing common database connection patterns is maintained, covering: the create_engine function of SQLAlchemy, the connect method of pymysql / mysqlclient, the connection function of psycopg2, the connect function of sqlite3, and the MongoClient constructor of MongoDB, etc. Each connection pattern has corresponding parameter extraction rules, which can parse key information such as server address, port, username, password, and database name from the connection string or parameters.

[0088] In step S22, a set of transformation rule libraries for file operations and database connections is maintained. For the operations identified in step S21, the corresponding rules are applied for code transformation. For complex situations that cannot be covered by the rules, the system calls the built-in AI model to analyze the context, infer the code intention, generate alternative code adapted to the platform, and sort according to the confidence level, and select the optimal solution for replacement.

[0089] In step S23, convert the local file path such as "C: / data.csv" into a platform-unified storage API call such as "get_platform_file('user_files / data.csv')" to ensure that the script can correctly access the file in the platform environment. At the same time, after the application path conversion, the system saves the mapping relationship between the original code and the modified code, and visually displays the modified content in the user interface, allowing the user to confirm or reject specific modifications, providing a modification history, and supporting rollback to a previous version to ensure the controllability and transparency of code modifications.

[0090] In a preferred embodiment of the present invention, as Figure 4 shown, in step S3, according to the standardized API call, perform input / output feature analysis and visual arrangement to form a DAG definition of the workflow, including: Step S31: Analyze the variable definitions and function return values in the standardized API call to generate an input / output feature representation of the script component; Step S32: According to the input / output feature representation, configure the node parameters and dependencies in a visual drag-and-drop manner to form an arrangement result; Step S33: Perform connection validity verification and format conversion on the arrangement result, and output the DAG definition of the workflow.

[0091] Specifically, in step S31, analyze the variable definitions, function return values, and file operations of the script, infer the possible input parameters and output results of the script, and generate a script I / O feature description. For cases where automatic inference is not possible, provide a user manual definition interface to ensure that the script can be correctly connected in the workflow.

[0092] In step S32, the front end is based on React and SVG / Canvas technologies to implement a drag-and-drop workflow design interface, providing a visual representation of nodes (representing scripts) and connections (representing data flows). The user can intuitively drag and drop script components, set execution conditions and parameters, and specify the data flow direction through connection lines, and the system automatically checks the connection validity.

[0093] In step S33, the visual workflow created by the user on the interface is converted into a standardized DAG definition format (JSON format), including node information, dependency relationships, execution conditions, and parameter configurations. The system verifies the legality of the DAG, detects errors such as circular dependencies, and ensures the correct workflow structure. At the same time, based on the DAG definition, the system automatically generates configuration code suitable for the underlying execution engine (such as Airflow), including task definitions, dependency relationship settings, and parameter passing logic. In addition, the system also provides a workflow simulation execution function, allowing users to verify the correctness of data flow and parameter passing without actually running the complete workflow, helping users discover and fix potential problems.

[0094] In a preferred embodiment of the present invention, as Figure 5 shown, in step S4, the script components in the DAG definition of the workflow are subjected to function-level encapsulation and interface standardization processing to obtain reusable modules and their version management solutions, including: Step S41: Based on the script components, perform function splitting and interface definition to obtain standardized function modules; Step S42: Based on the standardized function modules, create metadata descriptions of function classification and version information to obtain a component market interface; Step S43: Based on the component market interface, track version updates and evaluate the impact of compatibility to obtain upgrade suggestions.

[0095] Specifically, in step S41, technicians are guided to split the script into independent functions according to functions and add standardized function signatures and docstrings. The processed function modules contain clear parameter definitions, return value descriptions, and usage examples, enabling non-technical personnel to understand the module functions and usage methods.

[0096] In step S42, a structured metadata description is created for each script component, including function classification, applicable scenarios, input parameter descriptions, output formats, version information, etc. The metadata is stored in JSON format, supporting search and filtering, and facilitating users to quickly find the required components. At the same time, the system constructs a component market interface to visually display available script components, including function descriptions, usage frequencies, ratings, and other information. Users can browse, search, and compare components and select appropriate components to add to their own workflows.

[0097] In step S43, a component version management mechanism is implemented to automatically track the component update history and record the content of each change. When the component interface changes, the system automatically evaluates the compatibility impact and provides migration suggestions when incompatibility occurs, ensuring that the workflows depending on the component will not be interrupted due to upgrades. When a new version of a component used in a workflow is detected, the system automatically evaluates the upgrade risk and provides a visual impact analysis report to help users make upgrade decisions.

[0098] In a preferred embodiment of the present invention, as Figure 6 shown, in step S5, configure the workflow scheduling policy and perform automated operation and monitoring, and output the execution status and exception handling results of the workflow, including: Step S51: Based on the DAG definition of the workflow and the version management scheme of the reusable module, obtain the workflow configuration information, convert the execution time rule into a timing expression, and obtain the scheduling policy; Step S52: Based on the scheduling policy, collect the operation status data of the workflow nodes to obtain the execution progress information; Step S53: Based on the execution progress information, perform exception detection and recovery policy execution to obtain the processing result and the notification message.

[0099] Specifically, in step S51, the system provides a visual timing policy setting interface, supporting precise scheduling based on Cron expressions and simplified configuration based on natural language (such as "9 o'clock on Monday morning"). The system automatically converts the user configuration into a standard Cron expression and displays the specific time points of the next few executions to ensure compliance with user expectations.

[0100] In step S52, the system real-time tracks the execution status of the workflow, collects the running time, resource usage, and output logs of each node. The front end visually displays the execution progress and status, including the completed nodes, the currently executing node, and the waiting nodes, enabling users to intuitively understand the workflow running situation.

[0101] In step S53, the system monitors abnormal situations during the workflow execution process, such as node execution failure, timeout, or insufficient resources. Automatically attempt to recover according to the preset policy, including task retry, resource reallocation, or alternative path execution, to ensure the completion of the workflow to the greatest extent and reduce the need for manual intervention. At the same time, the system configures multi-level notification policies and selects different notification methods according to the importance of the event to ensure that important information is conveyed in a timely manner without causing excessive disturbance.

[0102] In a preferred embodiment of the present invention, as Figure 7 shown, the method further includes: Step S6: Obtain the historical execution data of the workflow, extract the time - dimension features and resource - dimension features, and obtain the workflow feature vector; Step S7: Based on the workflow feature vector, use the density clustering algorithm to group the workflows and obtain workflow groups with similar execution characteristics; Step S8: Based on the workflow groups, generate a pseudo - random detection point sequence and collect status snapshots to obtain the anomaly scores.

[0103] Specifically, in Step S6, the system performs multi - dimensional feature extraction on the historical execution data of all workflows, including time - dimension features (execution duration, start time, end time, timing frequency) and resource - dimension features (CPU usage rate, memory occupancy, I / O operation frequency, network traffic). The system uses the sliding window technique to pre - process these raw data, eliminate noise and standardize the feature values. At the same time, it calculates the autocorrelation and cross - correlation of the time series of each feature to form the "behavior fingerprint" of the workflow.

[0104] During the processing of historical execution data, a multi - level data pre - processing strategy is adopted. First, data cleaning is performed on the original execution logs to identify and correct outliers, missing values, and duplicate records to ensure data quality. For time - series data, an overlapping sliding window with a length of 24 hours and a step size of 4 hours is applied. This configuration can effectively capture the intra - day fluctuation patterns. For each sliding window, the system calculates the exponentially weighted moving average with a decay factor of 0.85, which smooths short - term noise and retains trend changes. For resource data such as CPU and memory, multi - dimensional features such as mean, peak value, amplitude of fluctuation, and sudden increase frequency are captured; for I / O operations, professional metrics such as read - write ratio, average block size, and queue depth are additionally recorded. The system pays special attention to the burst mode of resource usage. By analyzing the first - order and second - order derivatives of resource metrics, the stages of rapid growth or decline are identified, and these features are crucial for predicting resource bottlenecks. In terms of time - pattern analysis, the system uses the fast Fourier transform (FFT) to detect periodicity and analyzes the time dependence through the autocorrelation function (ACF) to identify multiple periodicities at the hourly, daily, weekly, and monthly levels. For workflows with strong occasionality, the extreme value theory model is also applied to establish the extreme - value distribution characteristics of resource usage. All these extracted features are standardized (subtracting the mean and dividing by the standard deviation), and then dimensionality reduction is performed through principal component analysis (PCA), retaining the first N principal components that can explain 95% of the total variance. This complex feature engineering process ensures that the workflow feature vector can comprehensively and accurately characterize its running behavior characteristics.

[0105] In step S7, the system uses the DBSCAN (Density-Based Spatial Clustering) algorithm to cluster workflow features and identify workflow groups with similar execution characteristics. The system dynamically adjusts the neighborhood parameter ε to find the optimal clustering effect and evaluates the clustering quality through the silhouette coefficient. For each cluster, the system calculates the feature center point of the core samples and extracts the common features of the cluster, such as the typical execution duration range, resource usage pattern, failure rate, etc. These clustering results are used for resource planning and anomaly detection and are automatically updated every 24 hours to ensure that the model adapts to the dynamic changes of the system.

[0106] The density clustering algorithm of the system has achieved multiple innovative optimizations. First, in terms of parameter self-adaptation, the optimal neighborhood radius ε is automatically determined through k-distance graph analysis. Specifically, the system calculates the distance from each sample to its k-th nearest neighbor, sorts these distances and plots the k-distance graph, and then determines the value of ε by detecting the inflection point of the curve (using discrete curvature calculation). For the minimum number of samples MinPts, it is set to the number of dimensions + 1 (the number of principal components after PCA) and adjusted according to the total number of workflows to ensure that the clustering will not be over-segmented when the number of workflows increases. In the selection of distance metrics, the system implements Mahalanobis distance calculation, which considers the correlation between features and has better discrimination ability for high-dimensional data. To process large-scale datasets, a grid-accelerated DBSCAN variant is adopted, which divides the feature space into grid cells and only calculates the distances of the neighbors of non-empty grids, significantly reducing the computational amount. The system also implements a distributed DBSCAN version, which adopts a graph-based parallel processing framework to divide the data into multiple computing nodes, improving the processing speed while ensuring the correctness of the results. The clustering quality evaluation uses multiple indicators for comprehensive consideration: the Silhouette Coefficient evaluates the clustering compactness and separation, and the Davies-Bouldin index evaluates the similarity between clusters. Compared with using a single indicator alone, the combination of multiple indicators provides a more comprehensive quality evaluation. The system also implements an incremental clustering update mechanism. When new workflow data accumulates to a certain threshold or a significant decrease in clustering quality is detected (any evaluation indicator changes by more than 10%), the clustering recomputation is triggered. The recomputation process adopts an incremental update strategy, only re-evaluating the cluster membership of new data and samples in the boundary region, greatly improving the update efficiency. For each formed cluster, the system deeply analyzes its feature distribution, calculates the core interval (5% to 95% percentile) of each dimension feature, the feature correlation matrix, and the typical time series pattern. These cluster features are stored in a distributed cache system to support millisecond-level real-time queries and provide a benchmark reference for anomaly detection.

[0107] In step S8, the system generates a pseudo-random detection point sequence for each workflow based on a hash function and Poisson distribution, avoiding system load peaks caused by traditional fixed-interval detection and vulnerabilities that may be circumvented. The detection point density is dynamically adjusted according to the importance of the workflow, historical stability, and current system load. When each detection point is triggered, the system collects a snapshot of the current execution state and compares it with the predicted value of the clustering model to calculate the deviation score. The system also constructs a knowledge base containing various abnormal patterns, and through real-time matching and analysis, identifies potential abnormal situations, and conducts intelligent early warning grading and self-healing processing according to the scope of influence.

[0108] The pseudo-random detection point generation mechanism of the system uses a cryptographically secure random number generator, combined with the SHA-256 hash value of the workflow ID as the seed, ensuring that the detection point sequence is both unpredictable and deterministic (the same workflow will generate a consistent detection sequence under the same conditions). Based on the non-uniform Poisson process model, the system dynamically adjusts the detection intensity λ(t). Specifically, the calculation of λ(t) considers multiple factors: the importance weight of the workflow (based on the business impact rating, ranging from 1 to 10), the historical anomaly rate (the frequency of abnormal events in the past 30 days), the current system load level, and the time of day. The system sets the upper and lower limits of the detection interval to ensure a minimum monitoring guarantee even for the lowest-priority workflows, while avoiding over-monitoring of high-priority workflows that may cause system load. The system also implements an adaptive adjustment mechanism. When the anomaly rate of a certain type of workflow is detected to increase, the monitoring intensity of this type of workflow is automatically increased; when the overall system load approaches the warning threshold, the monitoring frequency of workflows with high stability is preferentially reduced. For state collection, the system implements two modes: lightweight mode and deep mode. The lightweight mode only collects basic metrics (CPU, memory, disk I / O, network traffic), with low overhead and high frequency; the deep mode additionally collects thread status, stack information, detailed log mode, etc., providing more comprehensive analysis data but with higher overhead, so the frequency is lower. The system also implements a state inference algorithm between detection points, using a conditional random field (CRF) model. Based on the temporal correlation of historical data, it infers the possible states at time points that are not directly observed, improving the monitoring coverage while avoiding the performance impact caused by continuous high-frequency sampling. The abnormal score calculation integrates multiple methods: Mahalanobis distance calculation based on the common neighborhood, deviation score based on the LSTM prediction model, and score of the expert rule engine. The final abnormal score integrates these three scores through a dynamic weight method, and the weight is automatically adjusted according to the historical detection accuracy. When the abnormal score exceeds the threshold (this threshold is dynamically set based on the importance of the workflow), the system triggers a multi-level response mechanism, from simple recording to active intervention, ensuring timely handling of potential problems while avoiding the stability risk caused by over-intervention.

[0109] In a preferred embodiment of the present invention, asFigure 8 As shown, the method further includes: Step S9: Perform a multiplication graph model transformation on the DAG definition of the workflow to obtain a node mapping relationship; Step S10: Based on the node mapping relationship, use a spanning tree covering algorithm to find an optimal connection path to obtain a real-time connection suggestion; Step S11: Based on the real-time connection suggestion, use a virtual force model to perform layout optimization to obtain a display effect with reduced wire crossings.

[0110] Specifically, in step S9, the system performs a deep dependency analysis on all script components, establishes a direct dependency graph between components, and converts this initial dependency graph into a multiplication graph structure. The system first identifies the input and output types, data formats, and call patterns of each component, and then uses a hash function to map components of similar types to adjacent nodes, ensuring that the distance between any two components that may need to be connected in the graph does not exceed the logarithmic level.

[0111] In the specific implementation of the multiplication graph model transformation, the system first constructs an initial dependency graph G=(V,E) for the workflow component set V, and the edge set E represents the direct dependency relationship between components. To convert this graph into a multiplication graph structure, the system implements an eigenvector representation method. Each component is mapped to a multi-dimensional feature space, and the eigenvector includes: a functional category encoding (using one-hot encoding to represent main functions such as data input, transformation, analysis, output, etc.), an input and output type signature (encoding data types and formats such as CSV, JSON, DataFrame, etc.), component performance characteristics (normalized values of metrics such as average execution time, resource consumption, etc.), and usage frequency characteristics (based on historical workflow usage statistics). The system uses the Locality-Sensitive Hashing (LSH) technique to map components with similar features to nearby hash buckets. Specifically, the SimHash algorithm is selected, using a 64-bit hash value. Compared with the standard LSH, SimHash is more suitable for maintaining the similarity of high-dimensional features. The system establishes d = log 2 (n)+2 virtual connections (n is the total number of components) for each component, connecting to other components with the closest hash value. To ensure the connection quality, the system applies a Hamming distance threshold filter, only retaining connections with a distance less than a preset threshold. This strategy controls the average degree of the graph within 25 in the experiment, while maintaining the shortest path between any two points not exceeding log 2(n) Jump. For version dependency management, the system implements a semantic version parser that supports complex version constraint expressions (such as "> = 1.2.0, <2.0.0, != 1.3.5"). Version constraints are represented as attributes on directed edges, and the system uses interval arithmetic algorithms to verify the compatibility between constraints. When multiple constraints act on a component simultaneously, the system calculates the effective version range through an intersection operation. If the intersection is empty, it indicates a version conflict. For detected conflicts, the system generates conflict resolution suggestions, including version upgrade paths, API compatibility analysis, and potential impact assessment, to help users make the best choice. The converted multiplication graph is stored using an adjacency list structure, and indexes are established for common query operations (such as shortest path finding, neighborhood query), significantly improving the execution efficiency of subsequent path discovery algorithms.

[0112] In step S10, the system implements a spanning tree covering algorithm based on the multiplication graph to discover the optimal connection paths between components. The system uses an improved Boruvka algorithm to construct a minimum spanning tree forest, which is particularly suitable for large-scale graph processing in a distributed environment. For the specified input and expected output by the user, the system can automatically recommend the shortest component connection path to minimize data conversion costs and execution complexity.

[0113] In the implementation of the spanning tree covering algorithm, components are first partitioned into functional domains through spectral clustering. Spectral clustering uses the eigenvectors of the Laplacian matrix of the component graph, applies K-means clustering in the feature space, and the value of K is automatically determined by the silhouette coefficient. For each functional domain, the system implements an improved Boruvka algorithm to construct a minimum spanning tree. The key optimizations of this algorithm include: parallel processing - using multiple threads to process independent subtrees simultaneously to achieve near-linear speedup on large component graphs; edge weight adaptivity - the edge weights comprehensively consider various factors and are dynamically adjusted by machine learning based on user feedback; sparse graph optimization - implemented using an adjacency list representation and a heap-optimized union-find set, optimizing the algorithm complexity to nearly O(E logV), where E is the number of edges and V is the number of nodes. To handle cross-domain connections, the system evaluates all possible bridging edges between each pair of functional domains and calculates the importance scores of the bridging edges. The importance scoring considers the usage frequency of the connected components (the number of times they co-occurred in the workflow in the past 90 days), the complexity of data type conversion (based on a predefined type conversion cost matrix), the co-occurrence pattern in the historical workflow (using association rule mining algorithms), and the centrality measure in the doubling graph (calculated using the PageRank algorithm). The top k bridging edges with the highest importance scores (k is the logarithm of the number of functional domains) are selected, and these bridging edges together with the intra-domain spanning trees form a forest structure that covers the entire component graph. In the real-time path recommendation section, the system utilizes the spanning tree forest characteristics to quickly find the optimal connection paths between components. Specifically, if two components are in the same tree, there is a unique path between them, which can be found in O(log V) time through the parent pointer array of the tree; if they are in different trees, the system needs to find the best cross-tree bridging path, and the algorithm checks all possible bridging combinations and uses dynamic programming to find the optimal path combination. The path scoring comprehensively considers the path length, data conversion consistency, execution efficiency, and reliability. An adaptive learning mechanism for path recommendation is also implemented. By recording the user's acceptance or rejection behavior of the recommended paths, the path scoring weights and recommendation strategies are continuously adjusted. Using a reinforcement learning framework, the user behavior is regarded as a reward signal to optimize the recommendation strategy to maximize the long-term acceptance rate. This continuous learning enables the system to adapt to the unique working patterns and user preferences of the organization and provide increasingly accurate path suggestions.

[0114] In step S11, based on the characteristics of the doubling graph, the system implements a congestion-aware automatic layout algorithm. First, the system divides the workflow graph into multiple layers with an equal number of nodes in each layer and the highest intra-layer node correlation. Then, an improved Sugiyama algorithm is applied for the initial layout, and a congestion weight function is introduced to evaluate the wire density of each region. For congested regions, the system uses a local optimization algorithm based on the spring model to readjust the node positions through the action of virtual forces until the congestion degree is reduced to an acceptable level.

[0115] The layout optimization algorithm of the system realizes the visual optimization of the workflow based on a physical mechanics model. First, using an improved Coffman-Graham hierarchical algorithm, the system converts the workflow DAG into a hierarchical structure. This algorithm is achieved through two-stage processing: in the first stage, priority numbers are assigned to the nodes, taking into account the depth and breadth of node dependencies; in the second stage, the nodes are assigned to levels according to the priorities, ensuring that dependent nodes are always located in the previous layer. The system extends the standard algorithm by adding a weight factor that considers the data flow size, making the paths with large data flows arranged vertically as much as possible and reducing long horizontal connections. After the layering is completed, the system applies a modified Sugiyama algorithm for the initial node arrangement, which includes four steps: cyclic removal, level assignment, in-level node sorting, and coordinate assignment. In the in-level node sorting step, a weighted median method is adopted, considering the connection weights and connection directions to keep related nodes as close as possible. The core of the layout optimization is a congestion-aware force-directed algorithm, which defines a congestion weight function w(x, y) representing the connection density in the area near the coordinate (x, y). Specifically, a two-dimensional Gaussian kernel function is used to calculate the contribution of each connection to each spatial point, and then a fast convolution algorithm is used to efficiently calculate the congestion map of the entire visible area. During the optimization process, the system iteratively calculates the forces acting on each node: the repulsive force between nodes (inversely proportional to the distance and proportional to the congestion degree), the attractive force generated by the connections (proportional to the distance and proportional to the importance of the connections), and the layer conservation force (keeping the nodes within their assigned layers). The system uses a modified Verlet integration method to update the node positions and applies a simulated annealing strategy to gradually reduce the maximum movement distance of the nodes, ensuring that the algorithm converges to a local optimal solution. For large workflows (more than 100 nodes), the system adopts a multi-scale optimization strategy: first, a rough layout optimization is performed on a simplified version of the workflow (merging similar nodes), and then it is gradually refined to optimize the local details while maintaining the overall structure. The system also performs special optimization on the connection visualization, implementing an intelligent line algorithm that adaptively selects straight lines, broken lines, orthogonal lines, or Bézier curves for different scenarios. In high-congestion areas, the system preferentially uses curves to reduce visual intersections; for critical paths, thick lines or special colors are used to highlight them. The connection grouping function is also supported. When there are multiple parallel connections between multiple nodes, they are visualized as a single thick connection, and users can expand and view the details as needed. This multi-level layout optimization strategy significantly improves the readability of complex workflows, enabling non-technical personnel to intuitively understand the data flow and easily perform editing and adjustment.

[0116] In a preferred embodiment of the present invention, as Figure 9 shown, the method further includes: Step S12: Obtain the DAG definition of the workflow, construct a graph representation annotated with function labels and attribute markers, and obtain a multi-level workflow model; Step S13: Based on the multi-level workflow model, use the connected maximum common subgraph algorithm for pattern recognition to obtain a shared pattern; Step S14: Based on the shared pattern, calculate the importance score and perform a normalization transformation to obtain a new reusable module.

[0117] Specifically, in step S12, the system constructs a multi-level labeled graph representation of the workflow, converting each workflow into a graph structure with rich attribute labels. The nodes not only contain the script component IDs but also additional multi-dimensional attribute labels such as function labels, input / output types, execution frequencies, etc. At the same time, the edges also contain marker information such as data flow types and transformation complexities. The system adopts a hierarchical storage architecture to support representing the same workflow at different abstraction levels to meet the requirements for discovering subgraphs of various complexities.

[0118] In the implementation of the multi-level workflow model, a refined labeled graph data structure was first constructed. Each workflow is represented as a graph G=(V,E,L,A), where the node set V represents script components, the edge set E represents the data flow direction, L is a function that maps nodes and edges to functional labels, and A maps them to a set of detailed attributes. Node labels contain multi-dimensional information: the functional category adopts a hierarchical classification system, such as "data processing / cleaning / missing value processing", supporting functional descriptions accurate to the third level; the input / output types adopt a unified data type system, defining basic types (such as numerical, text, time) and composite types (such as table, tree, graphic data), as well as the conversion relationships between types; the execution characteristics include the average running time (and its variance), the resource usage pattern (CPU-intensive, memory-intensive or I / O-intensive), and the degree of parallelism; the quality metrics record the historical failure rate, the average number of errors, and common error patterns. Edge labels are also rich: the data flow type indicates the data format and protocol for transmission; the data magnitude uses a logarithmic scale to represent the data volume size range; the conversion complexity represents the difficulty of data conversion during transmission, ranging from direct transfer (complexity 0) to structural changes requiring in-depth conversion (complexity 5). The system collects label information through three channels: automatically extracting structured information from component metadata; analyzing historical execution records and extracting performance characteristics through statistical methods; integrating user-added annotations and converting them into system labels through standardization processing. To support multi-level representation, a hierarchical abstraction algorithm was implemented. Layer L0 (detailed layer) retains the complete structure of the original graph; Layer L1 (intermediate layer) merges consecutive similar-functional nodes through aggregation rules, such as merging multiple consecutive data cleaning steps into a single "data cleaning" node; Layer L2 (overview layer) is further abstracted, only retaining the core functional blocks and the main data flow paths. This abstraction is achieved through functional similarity analysis. The system uses a semantic similarity algorithm to calculate the functional similarity between nodes, then applies a community detection algorithm to identify closely related node groups, and generates abstract nodes for each group, retaining the statistical summary of important attributes within the group. Different levels are associated through a bi-directional mapping function, supporting seamless drilling down from the high-level overview to the specific implementation details. The graph data is stored in a dedicated graph database, which is specially optimized for efficient graph retrieval: a multi-dimensional index based on functional labels, attribute characteristics, and topological structure is established; a caching mechanism for common query patterns is implemented; for frequently updated workflows, an incremental index update strategy is adopted to avoid the performance overhead caused by full-scale reconstruction. To ensure data consistency in a high-concurrency environment, a transaction processing mechanism based on MVCC (Multi-Version Concurrency Control) is implemented, supporting atomic updates and query consistency of graph data.

[0119] In step S13, the system implements an efficient connected maximum common subgraph algorithm to identify shared patterns in a large number of workflow labeled graphs. The system adopts an improved McGregor algorithm combined with a pruning optimization strategy. For large-scale graph sets, a distributed computing framework is introduced to partition and process the graphs in parallel. The algorithm pays particular attention to the label compatibility of nodes and edges. Two nodes are considered mappable only when key labels such as functional categories and input / output types match. At the same time, a fuzzy matching mechanism is implemented, allowing label differences within a certain threshold.

[0120] In the implementation of the connected maximum common subgraph algorithm, several optimizations were first made to the basic McGregor algorithm to handle large-scale workflow sets. The standard McGregor algorithm finds the common subgraph of two graphs through a state-space search method, and its core is a backtracking process that attempts to map the nodes of one graph to the nodes of another graph. Three key optimizations were introduced on the basis of the original algorithm: pruning strategy - quickly exclude search branches that cannot produce larger results through the upper bound of the subgraph size. The upper bound estimation is determined by comparing the number of unprocessed nodes and the current maximum match. When the upper bound is less than or equal to the current maximum match size, immediately terminate the search for that branch; node pre-sorting - prioritize nodes according to indicators such as degree, label rarity, and centrality, and preferentially explore node pairs that are more likely to form large matches. Experiments show that this heuristic strategy reduces the search space by more than 60% on average; similarity pre-filtering - calculate the similarity scores of node pairs before backtracking, and only consider node pairs that exceed a threshold (usually set to 0.7), greatly reducing the search space. In terms of node compatibility judgment, the system implements multi-dimensional similarity calculations: for functional category labels, use the semantic similarity of the hierarchical classification system, considering the distance in the category tree; for text description labels, use BERT-based semantic embeddings and cosine similarity; for numerical attributes, use the normalized Euclidean distance; for input and output types, use a predefined type compatibility matrix to represent the conversion difficulty between different data types. The system also introduces structural compatibility checks to ensure that the neighbor relationships of the mapped nodes are consistent in the two graphs. For large-scale graph sets, the system adopts a multi-stage distributed computing strategy: first, use the LSH (Locality-Sensitive Hashing) technique to perform preliminary clustering on the workflow graphs and group similar graphs; then, apply the optimized MCCS algorithm in parallel within each cluster to find the exact common subgraph; finally, merge the results and remove redundancies. The system uses Apache Spark as the distributed computing framework to achieve partitioned storage and parallel processing of workflow graphs. To ensure that the discovered subgraph meets the connectivity requirements, the system performs connectivity verification after each match is completed, using breadth-first search to confirm that the mapped nodes form a connected subgraph in the original graph. In addition, edge compatibility is also concerned, and the types and attributes of edges are checked for consistency in different workflows to ensure the coherence of data flow semantics. To further improve performance, the system implements a multi-level caching strategy: cache the results of subgraph isomorphism checks - avoid repeated calculation of the isomorphism of the same subgraph pairs; cache intermediate mapping results - save partial mapping results to accelerate the processing of similar graphs; cache feature vectors - store pre-computed graph feature vectors to accelerate similarity calculations. These optimizations enable the system to complete comprehensive pattern mining within an acceptable time in an enterprise environment containing thousands of workflows.

[0121] In step S14, the system develops a multi-dimensional scoring system to rank the discovered common subgraphs by importance and identify the most valuable workflow patterns. The scoring dimensions include subgraph size, usage frequency, business value, performance stability, and structural integrity, etc. The system automatically converts the identified high-value common subgraphs into standardized reusable modules, analyzes the boundary conditions of the subgraphs, determines clear input and output interfaces, extracts the parameter configuration patterns inside the subgraphs, and finally generates self-describing documents for the modules, including function descriptions, typical use cases, and configuration guides.

[0122] During the sharing mode scoring and conversion process, a multi-dimensional value evaluation framework is implemented. The scoring system considers five core dimensions: subgraph scale - quantified by the weighted sum of the number of nodes and edges, with the weight set to 3:1, reflecting the dominant position of nodes in the workflow; usage frequency - the number of times of complete occurrence and the number of partial matches in different workflows; business value - evaluated by the strength of association with key business indicators. The system analyzes the business labels and impact indicators of the workflows to which the subgraphs belong and uses a machine learning model to quantify this association; performance stability - calculates the coefficient of variation of time and resource consumption (standard deviation divided by the mean) based on historical execution data, and subgraphs with high stability receive higher scores; structural integrity - evaluates the self-contained degree of the subgraph as an independent functional unit, including the clarity of input and output interfaces, the encapsulation degree of internal dependencies, and the minimization of external dependencies.

[0123] As Figure 10 shown, an embodiment of the present invention also provides a script intelligent adaptation and visualization orchestration system for non-technical personnel, including: A dependency analysis module 100, configured to obtain a script file uploaded by a user, perform a dependency library analysis on the script file, and obtain an isolated container environment; A path adaptation module 200, configured to perform static analysis and conversion on file operation statements and database connection codes in the script file based on the isolated container environment to obtain standardized API calls; A workflow orchestration module 300, configured to perform input and output feature analysis and visualization orchestration based on the standardized API calls to obtain a DAG definition of the workflow; A module management module 400, configured to perform function-level encapsulation and interface standardization on script components based on the DAG definition of the workflow to obtain reusable modules and their version management schemes; A scheduling execution module 500, configured to configure a workflow scheduling policy and perform automated operation and monitoring based on the DAG definition of the workflow and in combination with the version management scheme of the reusable modules, and output the execution status of the workflow and the exception handling result.

[0124] The dependency analysis module 100 implements the function of step S1 in the above method, performs AST analysis on the script file, extracts third-party library reference information, constructs a dependency graph, detects version conflicts, generates an isolated environment configuration, and calls the container API to build an image.

[0125] The path adaptation module 200 implements the function of step S2 in the above method, performs static analysis on the script file, identifies file operation statements and database connection codes, generates a file operation list and a database access list, and performs code conversion according to a rule library or an AI model to achieve intelligent adaptation of file paths and data sources.

[0126] The workflow orchestration module 300 implements the function of step S3 in the above method, analyzes the input and output characteristics of the script, provides a visual drag-and-drop interface, supports node parameter configuration and dependency relationship setting, generates a standardized DAG definition, and provides a workflow simulation verification function.

[0127] The module management module 400 implements the function of step S4 in the above method, encapsulates the script components at the function level, creates a standardized interface and metadata description, constructs a component market interface, and implements version management and compatibility evaluation.

[0128] The scheduling and execution module 500 implements the function of step S5 in the above method, provides a timing policy configuration interface, converts the user configuration into a standard scheduling expression, monitors the execution status of the workflow, handles exception situations, and provides a multi-level notification mechanism.

[0129] An embodiment of the present invention may further include a workflow analysis module that implements the functions of steps S6 to S8 in the above method, extracts features from the workflow historical execution data, groups the workflows using a density clustering algorithm, generates a pseudo-random detection point sequence, and performs anomaly detection and warning.

[0130] An embodiment of the present invention may further include a layout optimization module that implements the functions of steps S9 to S11 in the above method, converts the workflow DAG into a multiplication graph model, uses a spanning tree covering algorithm to find the optimal connection path, and performs layout optimization through a virtual force model to improve the visualization effect.

[0131] An embodiment of the present invention may further include a pattern recognition module that implements the functions of steps S12 to S14 in the above method, constructs a multi-level workflow model, performs pattern recognition using a connected maximum common subgraph algorithm, discovers shared patterns, and converts them into new reusable modules.

[0132] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A script intelligent adaptation and visual arrangement method for non-technical personnel, characterized in that: include: Obtain the script file uploaded by the user, perform dependency library analysis on the script file, and obtain an isolated container environment; Using the isolated container environment, statically analyzing and converting the file operation statements and database connection codes in the script file to generate standardized API calls; According to the standardized API calls, input and output feature analysis and visual arrangement are performed to form a DAG definition of the workflow; Performing function-level encapsulation and interface standardization processing on the script components in the DAG definition of the workflow to obtain a reusable module and a version management solution thereof; Based on the DAG definition of the workflow and in combination with the version management solution of the reusable module, the workflow scheduling strategy is configured and automated operation and monitoring are performed, and the execution status and exception handling results of the workflow are output.

2. The method according to claim 1, characterized in that The step of analyzing the dependency library of the script file to obtain an isolated container environment includes: Performing an abstract syntax tree (AST) analysis on the script file, extracting import statements and from-import statements, and obtaining third-party library reference information; Build a dependency graph based on the third-party library reference information and perform version conflict detection. When a version conflict is detected, use a dependency isolation strategy to generate different isolation environment configurations; The isolated environment configuration is used to call the container API to build the image and store it in a private warehouse to form the isolated container environment.

3. The method according to claim 1, characterized in that The static analysis and conversion of the file operation statements and database connection codes in the script file to obtain standardized API calls includes: Process the script file using abstract syntax tree (AST) analysis and regular expression matching to generate a file operation list and a database access list; Perform code pattern recognition and classification according to the file operation list and the database access list to form a conversion rule mapping; The conversion rule mapping is utilized to replace the file operation statements and database connection codes with the platform standard interface to obtain the standardized API call.

4. The method according to claim 1, characterized in that: The input and output feature analysis and visual arrangement are performed according to the standardized API call to form a DAG definition of the workflow, including: Analyze the variable definitions and function return values ​​in the standardized API calls to generate input and output feature representations of the script components; According to the input and output feature representation, node parameters and dependencies are configured by visual dragging to form an orchestration result; Connection validity verification and format conversion are performed on the orchestration result, and a DAG definition of the workflow is output.

5. The method according to claim 1, characterized in that The configuration of workflow scheduling strategy and automatic operation and monitoring, output of workflow execution status and exception handling results, includes: Based on the DAG definition of the workflow and the version management scheme of the reusable module, the workflow configuration information is obtained, and the execution time rule is converted into a timing expression to obtain a scheduling strategy; Based on the scheduling strategy, collect the running status data of the workflow nodes to obtain the execution progress information; Based on the execution progress information, anomaly detection and recovery strategy execution are performed to obtain processing results and notification messages.

6. The method according to claim 1, characterized in that Also includes: Acquire historical execution data of the workflow, extract time dimension features and resource dimension features, and obtain a workflow feature vector; Based on the workflow feature vector, a density clustering algorithm is used to group the workflows to obtain workflow groups with similar execution characteristics; Based on the workflow group, a pseudo-random detection point sequence is generated and a state snapshot is collected to obtain an anomaly score.

7. The method according to claim 1, characterized in that Also includes: Performing a multiplication graph model conversion on the DAG definition of the workflow to obtain a node mapping relationship; Based on the node mapping relationship, a spanning tree cover algorithm is used to find the optimal connection path to obtain a real-time connection suggestion; Based on the real-time connection suggestion, a virtual force model is used to perform layout optimization, thereby obtaining a display effect of reducing line crossing.

8. The method according to claim 1, characterized in that Also includes: Obtaining a DAG definition of the workflow, constructing a graph representation annotated with function labels and attribute tags, and obtaining a multi-level workflow model; Based on the multi-level workflow model, a connected maximum common subgraph algorithm is used to perform pattern recognition to obtain a shared pattern; Based on the sharing mode, the importance score is calculated and standardized and transformed to obtain a new reusable module.

9. The method according to claim 1, characterized in that: The script components in the DAG definition of the workflow are encapsulated at the function level and the interface is standardized to obtain a reusable module and a version management solution thereof, including: Function splitting and interface definition are performed based on the script component to obtain a standardized function module; Based on the standardized function module, create a metadata description of function classification and version information to obtain a component market interface; Based on the component market interface, version updates are tracked and compatibility impacts are evaluated to obtain upgrade suggestions.

10. A script intelligent adaptation and visual arrangement device for non-technical personnel, characterized in that: include: A dependency analysis module is used to obtain script files uploaded by users, perform dependency library analysis on the script files, and obtain an isolated container environment; A path adaptation module, used to statically analyze and convert file operation statements and database connection codes in the script file based on the isolated container environment to obtain standardized API calls; A workflow orchestration module, used to perform input and output feature analysis and visual orchestration based on the standardized API call to obtain a DAG definition of the workflow; A module management module, used to perform function-level encapsulation and interface standardization on script components based on the DAG definition of the workflow, to obtain a reusable module and its version management solution; The scheduling execution module is used to configure the workflow scheduling strategy and perform automatic operation and monitoring based on the DAG definition of the workflow and in combination with the version management solution of the reusable module, and output the execution status and exception handling results of the workflow.

Citation Information

Patent Citations

  • Third-party API integrated management method based on groover script technology

    CN112925666A

  • Cross-platform script language deployment method

    CN113076109A

  • Visual DAG workflow task scheduling system and operation method thereof

    CN113254010A

  • Draggable front-end logic arrangement method and device for low-code platform

    CN115639980A

  • Well hole mode wave effective frequency dispersion pickup method and device based on density clustering

    CN116070085A

Cited By

  • Intelligent code analysis system and method based on graph database

    CN120951325A

  • Intelligent agent process arrangement method and device and storage medium

    CN121116556A

  • Script-based relay satellite control method and device and computer equipment

    CN121367534A

  • Agent parallel industrial data workflow method based on DAG

    CN121436507A

  • Data processing method for unstructured table document

    CN122088480A