Method, device, equipment, and medium for automatically orchestrating multi-version data extraction tasks in DevOps

By identifying the data flow between modules, generating multi-level numbers and dynamically adjusting the frequency in the DevOps platform, the problems of missing data dependency identification and imperfect multi-version management are solved, the stability and flexibility of data extraction are improved, and rapid version switching and on-demand delivery are supported.

CN120561189BActive Publication Date: 2025-09-26GUANGZHOU CANWAY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511073238.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-09-26
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

The DevOps platform has problems such as lack of data dependency identification, imperfect multi-version management, and static extraction frequency, which leads to high coupling of data extraction tasks, chaotic execution order, resource waste, and data lag.

Method used

By identifying the data flow between DevOps platform modules, generating multi-level numbers, building a table association diagram, dynamically adjusting the extraction frequency, and implementing multi-version job isolation through spatial isolation, we ensure that dependent tables are extracted first and version switching is fast.

Benefits of technology

It achieves explicit management of data dependencies across modules and within modules, improves the stability and flexibility of data extraction, supports on-demand delivery of jobs, simplifies version iteration management, and reduces operation and maintenance complexity and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561189B_ABST
    Figure CN120561189B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device, and medium for automatically arranging multi-version data extraction tasks in DevOps. The method comprises: identifying the data flow between multiple modules of a DevOps platform and generating first-level numbering for cross-module business data tables; generating second-level numbering for business data tables based on the dependency order of built-in tables in the modules; constructing a table association relationship diagram, automatically arranging the execution order of data extraction tasks according to the order in the diagram, and implementing task grouping; dynamically adjusting the extraction frequency based on module load, and utilizing data space to isolate multi-version tasks. The present invention solves the problems of high coupling degree, difficult version management, and static extraction frequency in traditional data extraction tasks, improves the efficiency and reliability of DevOps data processing, and is suitable for data governance throughout the entire life cycle of software development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of DevOps data processing technology, and in particular to a method, apparatus, device, and medium for automatically orchestrating multi-version data extraction tasks in DevOps. Background Art

[0002] In DevOps platforms, efficient orchestration of data extraction tasks and automated data collection are the core foundation for achieving data insights into software development processes. However, existing technologies have the following key issues:

[0003] Lack of Data Dependency Identification: DevOps platforms consist of multiple functional modules, with complex data flows and implicit dependencies between them. However, existing technologies lack a systematic mechanism for identifying dependencies between data tables across modules and within modules. This results in high coupling between data extraction tasks and a chaotic execution order. If an error occurs in a data extraction task in a particular module, it can easily cause the entire task chain to crash, seriously impacting the stability of data collection.

[0004] Imperfect multi-version data management mechanisms: During the iteration process of the DevOps platform, data extraction tasks must support the parallel execution of multiple versions, but existing technologies do not implement data isolation for multi-version jobs. This can easily contaminate online data during the debugging of new version tasks, and data rollbacks require manual operations, which is costly and inefficient.

[0005] Static extraction frequency: Existing data extraction tasks often use a fixed execution frequency, without dynamically adjusting to the module's real-time concurrency and data volume. When module load is low, this fixed frequency can lead to wasted resources. However, when data volume surges, untimely extraction can lead to data lags, impacting the accuracy of data analysis.

[0006] Currently, there is no solution for full-process DevOps data extraction that integrates dependency modeling, multi-level number management, dynamic frequency adjustment, and multi-version isolation technologies. Therefore, an innovative approach is urgently needed to address these issues and improve the stability, flexibility, and efficiency of data extraction. Summary of the Invention

[0007] The present invention provides a method, apparatus, device, and medium for automatically orchestrating multi-version data extraction tasks in DevOps to address at least one of the problems in related technologies. The technical solution is as follows:

[0008] In a first aspect, an embodiment of the present application provides a DevOps method for automatically orchestrating multi-version data extraction tasks, including:

[0009] Identify the data flow between multiple modules of the DevOps platform and generate the first-level number of cross-module business data tables;

[0010] Combined with the dependency order of the module's built-in tables, the second-level number of the business data table is generated to form a "first-level.second-level" combination identifier, which serves as a globally unique identifier;

[0011] Construct a table association diagram based on the combined identifiers of the business data tables, and automatically arrange the execution order of data extraction tasks according to the order in the table association diagram to ensure that dependent tables are extracted first;

[0012] Based on the scheduled task execution order, tasks are grouped according to whether they cross modules. Cross-module tasks are classified as "combination jobs" and single-module tasks are classified as "module jobs." This allows users who only purchase some modules to be provided with jobs for specified modules on demand.

[0013] Real-time collection of module concurrency and request data volume, dynamic adjustment of job extraction frequency, reducing the frequency for high-concurrency modules and increasing the frequency for data surge modules, and automatically executing jobs at the adjusted frequency to obtain DevOps data;

[0014] By isolating multi-version jobs in different spaces, the databases of different jobs are isolated. When a new version is launched, it is quickly switched to the new database through database routing to ensure the efficiency of version switching.

[0015] In one implementation method, generating a first-level number for a cross-module business data table includes:

[0016] Capture API call logs and data packets between modules through traffic collection agents deployed at the interface layer of each module;

[0017] Extract the source module, target module and data entity information of the data flow, combine the module dependency configuration, and build a directed acyclic graph of the data flow between modules;

[0018] Based on the graph topology sorting results, the first-level number is assigned to the business data tables that flow across modules. The numbering rule is "source module ID-target module ID".

[0019] In one implementation method, generating the second level number of the business data table includes:

[0020] Parse the table creation statements, foreign key constraints, and view definitions of all data tables in the module, and identify direct and indirect dependencies between tables;

[0021] The Kahn algorithm is used to perform topological sorting on the table dependencies within the module, and a second-level number is assigned to each table according to the sorting result, which is combined with the first-level number to form the globally unique identifier "source module identifier-target module identifier.sequence identifier".

[0022] In one implementation method, automatically arranging the execution sequence of data extraction tasks includes:

[0023] Based on the global unique identifier, a full-scale association relationship graph is constructed and stored in a graph database;

[0024] Calling a topological sorting interface to sort the table association relationship graph and generate a task execution sequence to implement the extraction task of the parent dependency table taking precedence over the child table;

[0025] Set up a dependency change monitoring mechanism to automatically trigger the reordering of task execution when table associations are updated, and retain historical sorting records for traceability.

[0026] In one implementation method, grouping tasks based on whether they span modules includes:

[0027] Scan the associated data table combination identifier of each task in the task execution sequence and extract the module identifier field therein;

[0028] Count the number of module identifiers involved in the task. If the number of module identifiers is ≥ 2, mark it as a "combination task". If the number of module identifiers is 1, mark it as a "module task".

[0029] For tasks with the same module ID and quantity, a task package containing multi-module dependency configuration is generated;

[0030] Generate independent task packages for module jobs, which have no dependencies with other job packages and can be deployed independently;

[0031] Read the module list purchased by the user in the configuration file, map it with the task package, and load the corresponding job package on demand according to the list.

[0032] In one implementation method, dynamically adjusting the job extraction frequency includes:

[0033] Collect the concurrency and data increment of each module and generate frequency adjustment instructions;

[0034] Reduce the extraction frequency when concurrency is high, and increase the extraction frequency when data surges. Automatically execute jobs at the adjusted frequency and push the results to the storage layer.

[0035] In one implementation, isolating multiple versions of a job by space includes:

[0036] Create an independent logical space for each job version and use a database schema isolation mechanism to automatically create databases with different prefixes and store data after executing tasks in different spaces, ultimately achieving physical isolation of job data storage.

[0037] The database prefix serves as a unique identifier for data in different spaces. It can locate the data in the corresponding space and quickly switch to a database with a different prefix to obtain data when a job execution exception occurs, while avoiding cross-version data pollution.

[0038] In a second aspect, an embodiment of the present application provides a DevOps device for automatically orchestrating multi-version data extraction tasks, including:

[0039] A data flow direction identification module configured to identify the data flow direction between modules and generate a first-level number;

[0040] A multi-level number generation module is configured to parse the table creation dependency order within the module and generate the second-level number;

[0041] A task scheduling module configured to construct an association relationship graph based on the combination identifier and schedule the execution order of the tasks;

[0042] A task grouping module configured to group tasks by module number;

[0043] A dynamic frequency adjustment module configured to adjust the decimation frequency in real time;

[0044] The multi-version isolation module is configured to create independent spaces and isolate multi-version data through schema.

[0045] In a third aspect, embodiments of the present application further provide an electronic device comprising: a memory and a processor. The memory stores instructions, which are loaded and executed by the processor to implement the method of any of the aforementioned embodiments. The memory and the processor communicate with each other via an internal connection path.

[0046] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, the method in any one of the above-mentioned embodiments is implemented.

[0047] The advantages or beneficial effects of the above technical solution include at least:

[0048] In terms of dependency management, through data flow identification and multi-level numbering mechanism, automatic identification and explicit management of data dependencies across modules and within modules are achieved, avoiding confusion in the order of task execution due to implicit dependencies, reducing the impact of single module exceptions on the overall task chain, improving data extraction stability, and eliminating the tedious manual sorting of dependencies, thereby reducing the complexity of operation and maintenance.

[0049] In terms of multi-version management, relying on the independent logical space and database schema isolation design, physical isolation of multi-version job data is achieved to prevent new version debugging from contaminating online data; combined with the database routing mechanism, it supports rapid version switching and one-click rollback, improving the security and efficiency of version iteration and simplifying the management process.

[0050] To address the problem of static extraction frequency, the dynamic frequency adjustment module implements adaptive frequency adjustment based on real-time monitoring of module load and data volume: reducing the frequency to reduce resource waste during high concurrency and increasing the frequency to ensure timely collection when data surges, thus balancing system resource utilization and data timeliness.

[0051] In addition, the task grouping mechanism supports on-demand delivery of tasks based on user-purchased modules, improving the platform's adaptability to different needs; the decoupled design of each module ensures the scalability of the system, and new modules can be quickly connected as the platform iterates, extending the applicability period of the technical solution.

[0052] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present application will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. In the accompanying drawings, unless otherwise specified, the same reference numerals throughout multiple drawings represent the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that the following drawings only illustrate certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can also be obtained based on these drawings without paying creative work.

[0054] Figure 1 A flowchart of a DevOps method for automatically orchestrating multi-version data extraction tasks provided in one embodiment of the present application;

[0055] Figure 2 A two-level numbering diagram of a data extraction task provided in one embodiment of the present application;

[0056] Figure 3 A schematic diagram of a multi-version space management page provided in one embodiment of the present application;

[0057] Figure 4 A schematic diagram of a prefix isolation page for a multi-version space management database provided in one embodiment of the present application;

[0058] Figure 5 A schematic diagram of a device for automatically orchestrating multi-version data extraction tasks in DevOps according to an embodiment of the present application;

[0059] Figure 6 A structural block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0060] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0061] This application discloses a method for automatically arranging multi-version data extraction tasks in DevOps. The flowchart of this method is as follows: Figure 1 As shown, the two-level numbering of data extraction tasks is as follows Figure 2 As shown, the following steps are included:

[0062] S101. Identify the data flow between multiple modules of the DevOps platform and generate the first-level number of the cross-module business data table.

[0063] In specific implementation, traffic collection agents deployed at each module's interface layer capture inter-module API call logs in real time. A log parsing engine extracts source module, target module, and data entity information from the data stream. Combined with the module dependency configuration in the metadata repository, a directed acyclic graph (DAG) of inter-module data flows is constructed. Based on the topological sorting results in the graph, first-level numbers are assigned to business data tables flowing across modules, using the "source module ID - target module ID" numbering scheme.

[0064] For example, Figure 2 As shown, after capturing the HTTP request logs and data packets from the CTest module sending test plan data to the CTeam module, the source module is parsed to determine CTest and the target module is CTeam. The corresponding data entity is the "Test Plan Data Table." This analysis reveals that the CTest module depends on the required task data of the CTeam module. A directed acyclic graph of data flows is constructed based on this dependency relationship. Based on the topological sorting of the graph, the execution order from CTeam to CTest is determined. The cross-module business data table, the "Test Plan Data Table," is then assigned a first-level number "2-1," where 2 represents the CTest module identifier and 1 represents the CTeam module identifier.

[0065] S102. Generate the second-level number of the business data table based on the dependency order of the module's built-in tables to form a "first-level.second-level" combination identifier.

[0066] During implementation, the table creation statements, foreign key constraints, and view definitions for all data tables within the module are parsed, and the direct and indirect dependencies between tables are identified through a syntax analyzer. The Kahn algorithm is used to topologically sort the dependencies between tables within the module. For tables with circular dependencies, the sorting priority is determined based on the table creation timestamp. A second-level number is assigned to each table based on the sorting result, and combined with the first-level number to form a globally unique identifier: "source module identifier - target module identifier. Sequence identifier." A source module may involve multiple modules, and the final globally unique identifier format is "source module identifier 1 - source module identifier 2 - source module identifier x - target module identifier. Globally unique identifier."

[0067] For example, Figure 2 As shown, when parsing the table creation statements and foreign key constraints of the data tables in the CTest module, it was found that the foreign key field of the "Test Result Table" was associated with the primary key of the "Test Plan Table", and the field information of the "Test Plan Table" was referenced in the view definition of the "Test Result Table", thereby identifying the direct relationship between the "Test Result Table" and the "Test Plan Table". After topologically sorting the table dependencies in the CTest module using the Kahn algorithm, the execution order was determined to be "Test Plan Table → Test Result Table". Based on the sorting result, the second-level number "1" is assigned to the "Test Plan Table", which, together with the first-level number "2-1" generated in step S101, forms the globally unique identifier "2-1.1" of the "Test Plan Table".

[0068] S103: Construct a table association relationship diagram based on the combined identifiers of the business data tables, and automatically arrange the execution order of the data extraction tasks according to the order in the table association relationship diagram to ensure that dependent tables are extracted first.

[0069] In specific implementation, a graph-building tool generates a full-table relationship graph based on the "first-level.second-level" globally unique identifiers of the data tables. The graph database stores the tables as nodes and the dependencies as edges. The graph database's built-in topological sorting interface is used to perform depth-first sorting on the relationship graph, generating a task execution sequence to ensure that extraction tasks for parent dependency tables are executed before those for child tables. A dependency change monitoring component is deployed to monitor changes to table foreign key constraints, view definitions, and other changes in real time. When dependency updates are detected, the task execution order is automatically reordered, and the historical sorting results are stored in a time series database for traceability.

[0070] S104. According to the arranged task execution order, tasks are grouped according to whether they cross modules. Cross-module tasks are classified as "combined jobs" and single-module tasks are classified as "module jobs" to support on-demand provision of jobs for specified modules for users who only purchase some modules.

[0071] In specific implementation, the task parser scans the data table combination identifiers associated with each task in the task execution sequence and extracts the module fields from the "source module identifier - target module identifier" pair (e.g., "2" and "1" are extracted from "2-1.1"). A counting tool is used to count the number of independent module identifiers involved in each task. Tasks with a count of ≥2 are marked as "combination jobs," and those with a count of 1 are marked as "module jobs." For "combination jobs" with the same module identifiers and counts, a packaging tool generates a task package containing multi-module dependency configurations. For "module jobs," independent task packages are generated without cross-module dependency configurations, allowing for independent deployment. Finally, a configuration parser reads the user's purchased module list (e.g., ["CTest", "CTeam"]), maps it to the module identifiers of the task package, and only loads matching task packages.

[0072] For example, there are two tasks in the task execution sequence: Task A is associated with "2-1.1" (CTest→CTeam) and "1-2.1" (CTeam→CTest), the extracted modules are identified as "2" and "1", the quantity is 2, and it is marked as "combined job", generating a task package "CTest-CTeam_group.job" containing the interaction configuration of CTest and CTeam; Task B is only associated with "2-1.2" (inside CTest), the module is identified as "2", the quantity is 1, and it is marked as "module job", generating an independent package "CTest_single.job". If the user's purchase list is ["CTest"], the configuration parser will only load "CTest_single.job" and will not load the cross-module "CTest-CTeam_group.job". For example Figure 3 As shown in the diagram of the multi-version space management page, all_module_job_v7 is a combined job, and other jobs are "module jobs."

[0073] S105. Collect the concurrency and request data volume of the modules in real time, dynamically adjust the job extraction frequency, reduce the frequency for high-concurrency modules, increase the frequency for data surge modules, and automatically execute jobs at the adjusted frequency to obtain DevOps data.

[0074] In specific implementation, a monitoring agent deployed on the module server collects the concurrency of each module (such as the number of API calls per second) and the amount of requested data (such as the number of new records per minute) in real time. The data is calculated by the stream processing engine to generate frequency adjustment instructions. When the module concurrency exceeds the preset threshold, the frequency reduction instruction is triggered; when the data increment exceeds the threshold, the frequency increase instruction is triggered. The adjusted frequency is effective through the task scheduler, which automatically executes the extraction job and pushes the results to the Doris database in the storage layer.

[0075] For example, when the CTest module monitoring agent collects concurrent data at a rate of 1,200 times per second, exceeding the threshold of 1,000, the stream processing engine generates a frequency reduction instruction, and the scheduler adjusts its extraction frequency from 5 minutes per time to 15 minutes per time, reducing the use of module resources; when the automatic testing tool generates a large number of test results, the data increment reaches 150,000 per hour, exceeding the threshold of 100,000, and a frequency increase instruction is generated at this time, and the frequency is adjusted to 1 minute per time, ensuring that the test result data is extracted to the Doris database in a timely manner to ensure the timeliness of subsequent analysis.

[0076] S106. Isolate multi-version jobs through space, and isolate different space job databases. When a new version is launched, quickly switch to the new database through database routing to ensure the efficiency of version switching.

[0077] During implementation, an independent logical space is created for each job version. The database schema isolation mechanism is adopted. Each space automatically generates a database with a version prefix (such as "v1" and "v2"). After the job is executed, the data is automatically stored in the database with the corresponding prefix to achieve physical isolation. The database prefix is ​​stored in the routing configuration table as the unique identifier of the space. When the new version is launched, the database routing rules are modified to switch data access from the old database "v1" to the new database "v2" to complete the version switch.

[0078] For example, Figure 4 As shown in the diagram for the database prefix isolation page for multi-version space management, the logical space allocation schema for a v1.0 job is "v1," the database prefix is ​​"v1_," and data is stored in the "v1_test" table. When v2.0 is launched, a "v2" schema is created, the database prefix is ​​"v2_," and data is stored in the "v2_test" table. If duplicate data extraction occurs during v2.0 execution, operations personnel can modify the routing configuration (changing the routing key from "v2_" to "v1_") to switch to the v1.0 database within 10 seconds, preventing impact to online business. Prefix isolation prevents data contamination between the two versions.

[0079] Figure 5 A device for automatically orchestrating multi-version data extraction tasks in DevOps according to an embodiment of the present application is shown, including:

[0080] A data flow direction identification module configured to identify the data flow direction between modules and generate a first-level number;

[0081] A multi-level number generation module is configured to parse the table creation dependency order within the module and generate the second-level number;

[0082] A task scheduling module configured to construct an association relationship graph based on the combination identifier and schedule the execution order of the tasks;

[0083] A task grouping module configured to group tasks by module number;

[0084] A dynamic frequency adjustment module configured to adjust the decimation frequency in real time;

[0085] The multi-version isolation module is configured to create independent spaces and isolate multi-version data through schema.

[0086] Figure 6 FIG. 1 shows a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device includes: a memory 310 and a processor 320. The memory 310 stores instructions, which are loaded and executed by the processor 320 to implement the DevOps method for automatically orchestrating multi-version data extraction tasks in the above embodiment. The number of memory 310 and processor 320 can be one or more.

[0087] The electronic device also includes:

[0088] The communication interface 330 is used to communicate with the diplomatic equipment and perform data exchange transmission.

[0089] If the memory 310, processor 320, and communication interface 330 are implemented independently, the memory 310, processor 320, and communication interface 330 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA bus), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0090] Optionally, in a specific implementation, if the memory 310, the processor 320 and the communication interface 330 are integrated on a chip, the memory 310, the processor 320 and the communication interface 330 can communicate with each other through an internal interface.

[0091] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the method provided in the embodiment of the present application is implemented.

[0092] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in a memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.

[0093] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0094] It should be understood that the processor described above may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0095] Furthermore, optionally, the above-mentioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may include random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (LDRAM) and direct rambus RAM (DR RAM).

[0096] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0097] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0098] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0099] Any process or method description in a flow chart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations in which the functions may be performed in a different order than shown or discussed, including in a substantially simultaneous manner or in a reverse order depending on the functions involved.

[0100] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as a sequenced list of executable instructions for implementing the logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).

[0101] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0102] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0103] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A DevOps method for automatically arranging multi-version data extraction tasks, characterized in that: include: Identify the data flow between multiple modules of the DevOps platform and generate the first-level number of cross-module business data tables; Combined with the dependency order of the module's built-in tables, the second-level number of the business data table is generated to form a "first-level.second-level" combination identifier, which serves as a globally unique identifier; Construct a table association diagram based on the combined identifiers of the business data tables, and automatically arrange the execution order of data extraction tasks according to the order in the table association diagram to ensure that dependent tables are extracted first; Based on the scheduled task execution order, tasks are grouped according to whether they cross modules. Cross-module tasks are classified as "combined tasks," while single-module tasks are classified as "module tasks." This allows users who have only purchased some modules to be provided with tasks for specific modules on demand. Real-time collection of module concurrency and request data volume, dynamic adjustment of job extraction frequency, reducing the frequency for high-concurrency modules and increasing the frequency for data surge modules, and automatically executing jobs at the adjusted frequency to obtain DevOps data; By isolating multi-version jobs in different spaces and isolating databases for different jobs, when a new version is released, database routing is used to quickly switch to the new database, ensuring efficient version switching. The execution sequence of the automatic arrangement of data extraction tasks includes: Based on the global unique identifier, a full-scale association relationship graph is constructed and stored in a graph database; Calling a topological sorting interface to sort the table association relationship graph and generate a task execution sequence to implement the extraction task of the parent dependency table taking precedence over the child table; Set up a dependency change monitoring mechanism to automatically trigger the reordering of task execution when table associations are updated, and retain historical sorting records for traceability; The grouping of tasks according to whether they are across modules includes: Scan the associated data table combination identifier of each task in the task execution sequence and extract the module identifier field therein; Count the number of module identifiers involved in the task. If the number of module identifiers is ≥ 2, mark it as a "combination task". If the number of module identifiers is 1, mark it as a "module task". For tasks with the same module ID and quantity, a task package containing multi-module dependency configuration is generated; Generate independent task packages for module jobs, which have no dependencies with other job packages and can be deployed independently; Read the module list purchased by the user in the configuration file, map it with the task package, and load the corresponding job package on demand according to the list.

2. The method according to claim 1, characterized in that The first level numbering of generating the cross-module business data table includes: Capture API call logs and data packets between modules through traffic collection agents deployed at the interface layer of each module; Extract the source module, target module and data entity information of the data flow, combine the module dependency configuration, and build a directed acyclic graph of the data flow between modules; Based on the graph topology sorting results, the first-level number is assigned to the business data tables that flow across modules. The numbering rule is "source module ID - target module ID".

3. The method according to claim 1, characterized in that The second level numbering of generating the business data table includes: Parse the table creation statements, foreign key constraints, and view definitions of all data tables in the module, and identify direct and indirect dependencies between tables; The Kahn algorithm is used to perform topological sorting on the table dependencies within the module, and a second-level number is assigned to each table according to the sorting result, which is combined with the first-level number to form the globally unique identifier "source module identifier-target module identifier.sequence identifier".

4. The method according to claim 1, wherein The dynamic adjustment of the job extraction frequency includes: Collect the concurrency and data increment of each module and generate frequency adjustment instructions; Reduce the extraction frequency when concurrency is high, and increase the extraction frequency when data surges. Automatically execute jobs at the adjusted frequency and push the results to the storage layer.

5. The method according to claim 1, wherein The multi-version operation by spatial isolation includes: Create an independent logical space for each job version and use a database schema isolation mechanism to automatically create databases with different prefixes and store data after executing tasks in different spaces, ultimately achieving physical isolation of job data storage. The database prefix serves as a unique identifier for data in different spaces. It can locate the data in the corresponding space and quickly switch to a database with a different prefix to obtain data when a job execution exception occurs, while avoiding cross-version data pollution.

6. A device for automatically arranging multi-version data extraction tasks in DevOps, characterized in that: A method for executing a DevOps automatic orchestration multi-version data extraction task according to any one of claims 1 to 5, comprising: A data flow direction identification module configured to identify the data flow direction between modules and generate a first-level number; A multi-level number generation module is configured to parse the table creation dependency order within the module and generate the second-level number; A task scheduling module configured to construct an association relationship graph based on the combination identifier and schedule the execution order of the tasks; A task grouping module configured to group tasks by module number; A dynamic frequency adjustment module configured to adjust the decimation frequency in real time; The multi-version isolation module is configured to create independent spaces and isolate multi-version data through schema.

7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores instructions, and the instructions are loaded and executed by the processor to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Application release arrangement method and device, equipment and storage medium

    CN117407049A

  • Task management methods and systems

    US20250209426A1