Pipeline data conversion method and related equipment
By defining a common loosely coupled abstract data structure, the automatic conversion of PAC structures of different cloud vendors is solved, and the problem of users needing manual operations in cross-cloud pipeline data relocation is improved, efficiency and accuracy are improved, and user experience is improved.
Patent Information
- Application Number
- CN202410231508.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-02-29
- Publication Date
- 2025-05-13
AI Technical Summary
There are large differences in the pipeline code (PAC) data structures of different cloud manufacturers, which leads to manual operations for format conversion and manual verification when relocating pipeline business data, reducing the experience and efficiency of cross-cloud usage pipelines.
Provide a pipeline data conversion method, which can realize the automatic conversion of PAC structures of different cloud manufacturers by defining a common loosely coupled abstract data structure, ensure the accuracy and completeness of the data and avoid user manual operations.
It realizes automated PAC structure conversion, improves the efficiency and accuracy of cross-cloud pipeline data conversion, improves user experience, and saves time and energy.
Smart Images

Figure CN119988466A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 13, 2023, with application number 202311523701.9 and invention name “A code processing method and related equipment”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of cloud computing technology, and in particular to a pipeline data conversion method, a pipeline data conversion system, a computing device cluster, a computer-readable storage medium, and a computer program product. Background Art
[0003] With the continuous development of cloud computing technology, major cloud vendors have launched Pipeline As Code (PAC) services. PAC is a method of defining software delivery processes as code. This method uses a standardized, structured description language to describe the complete phase division, dependencies, task details, and other information of a pipeline in a code-like style according to specific business keywords or rules.
[0004] PAC allows developers to automate the entire software delivery process and store it in a version control system so that collaborating developers can easily view, modify, and share. PAC can also be integrated with other tools and technologies, such as continuous integration / continuous delivery (CI / CD) tools, containerization technology, and cloud computing platforms to achieve more efficient software delivery. With PAC, developers can treat the software delivery process as code, making it easier to manage and maintain.
[0005] However, there are large differences between the PAC data structures of different cloud vendors, which leads to huge challenges for users when migrating pipeline business data. For cross-cloud pipeline business data migration needs, users usually need to manually convert the format and manually verify the converted pipeline data, which greatly reduces the experience and efficiency of using pipelines across clouds. Summary of the invention
[0006] The present application provides a pipeline data conversion method, which provides users with an automated PAC structure conversion solution, and in a feasible PAC data conversion scenario, ensures the accuracy and integrity of the data as much as possible, can assist users to smoothly and quickly convert pipeline data, avoid manual operations by users, and achieve rapid cloud or cross-cloud pipeline data conversion, and provide convenience for data migration and conversion for users to use pipeline services across clouds, thereby improving the user experience. The present application also provides a pipeline data conversion system, a computing device cluster, a computer-readable storage medium, and a computer program product corresponding to the above method.
[0007] In a first aspect, the present application provides a pipeline data conversion method. The method can be performed by a pipeline data conversion system. The pipeline data conversion system can be a software system, which can be an independent software or integrated with other software. The software system can be provided to the user in the form of a software package and deployed by the user. Or the software system can be provided to the user in the form of a cloud service. The above software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software to execute the pipeline data conversion method of the present application. In some possible implementations, the pipeline data conversion system can be a hardware system, such as a computing device cluster with pipeline data conversion capabilities, and when the computing device cluster is running, the pipeline data conversion method of the present application is executed.
[0008] Specifically, the pipeline data conversion system can receive the first pipeline code PAC structure input by the user and the desired output PAC type, and extract the task information of the input pipeline task and the orchestration relationship of the input pipeline task according to the first PAC structure. Then the pipeline data conversion system can map the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task. The intermediate general task supports mutual conversion with different types of tasks. Then the pipeline data conversion system can convert the intermediate general task according to the desired output PAC type to obtain the target task, assemble the target task according to the orchestration relationship, and obtain the second PAC structure. The type of the second PAC structure matches the desired output PAC type, and the first PAC structure and the second PAC structure are suitable for different cloud architectures.
[0009] The method defines a set of universal pipeline PAC standards, such as a loosely coupled abstract data structure, which is used to establish a mapping relationship with the PAC structure defined by other cloud vendors. In this way, when converting the PAC structures of different cloud vendors, the universal pipeline PAC standard can be used as an intermediate transition layer for conversion to convert the PAC structure input by the user into a transitional intermediate state, and then according to the needs, such as the desired output PAC type, the intermediate state is converted into the desired output PAC structure. The method provides users with an automated PAC structure conversion solution, and in feasible PAC data conversion scenarios, it ensures the accuracy and integrity of the data as much as possible, can assist users in smoothly and quickly converting pipeline data, avoids manual operation, manual conversion, and manual verification by users, saves users a lot of time and energy, and realizes rapid cloud or cross-cloud pipeline data conversion, and provides convenience for data migration and conversion for users who use pipeline services across clouds, thereby improving the user experience.
[0010] In some possible implementations, the pipeline data conversion system can also map the arrangement relationship of the input pipeline task to the intermediate PAC structure. The intermediate PAC structure includes the intermediate general task and the mapped arrangement relationship. The pipeline data conversion system can convert the intermediate PAC structure according to the desired output PAC type to obtain the target task.
[0011] The method maps the orchestration relationship of pipeline tasks to an intermediate PAC structure, and transforms the intermediate PAC structure into a desired PAC type, thereby improving the accuracy and availability of PAC structure transformation.
[0012] In some possible implementations, the pipeline data conversion system can also identify task keywords of intermediate general tasks, retrieve a keyword mapping set based on the task keywords of the intermediate general tasks and the expected output PAC types, obtain a conversion strategy corresponding to the task keywords of the intermediate general tasks, and then convert the task keywords of the intermediate general tasks according to the conversion strategy to obtain the target task.
[0013] Among them, the keyword mapping set can come from a custom rule set. The method supports the conversion of pipeline tasks of different architectures or formats through a custom rule set to meet different business needs and has high availability.
[0014] In some possible implementations, the pipeline data conversion system can also input the task information of the pipeline task into the PAC conversion model to obtain an intermediate PAC structure, where the intermediate PAC structure includes an intermediate general task. Accordingly, the pipeline data conversion system can input the intermediate general task and the desired output PAC type into the PAC conversion model to obtain the target task.
[0015] This method supports converting pipeline tasks into intermediate PAC structures through the PAC conversion model, using the intermediate PAC structure as a transition and converting it into the target task through the PAC conversion model, laying the foundation for the PAC structure conversion of different cloud vendors.
[0016] In some possible implementations, the PAC conversion model is constructed as follows:
[0017] Obtain training data, which includes data pairs of different types of pipeline tasks and intermediate general tasks;
[0018] Based on the training data, supervised fine-tuning is performed on the language model to obtain the PAC conversion model.
[0019] This method uses the data of different types of pipeline tasks and intermediate general tasks to supervise and fine-tune the language model to obtain the PAC conversion model, thereby realizing the mutual conversion of different types of pipeline tasks and intermediate general tasks, and further providing assistance for the conversion of PAC structures of different manufacturers or different architectures and different formats.
[0020] In some possible implementations, when monitoring the PAC change information of the cloud vendor, the pipeline data conversion system can also update the PAC knowledge base or the PAC conversion model according to the PAC change information. The knowledge in the PAC knowledge base is used to construct the training data of the PAC conversion model, or to perform retrieval enhancement on the PAC conversion model to generate RAG.
[0021] This ensures that new knowledge can be acquired and converted in a timely manner during the PAC structure conversion process, thereby improving the quality and effect of pipeline data conversion.
[0022] In some possible implementations, the PAC change information includes text change information and / or structure change information. Accordingly, the pipeline data conversion system can determine the type and impact of the PAC change information based on the text change information and / or structure change information. When the type and impact of the PAC change information indicate that the change is valid, the PAC change information is updated to the PAC knowledge base in units of tasks.
[0023] This method adds the effective information of the change to the PAC knowledge base. On the one hand, it can ensure that new knowledge can be acquired in time for transformation during the PAC structure transformation process. On the other hand, it only updates the effective information of the change instead of all the change information, which can reduce the frequency of updates and ensure user experience.
[0024] In some possible implementations, the loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data, and step layer data.
[0025] This method divides the PAC structure into four layers. Based on the custom four-layer pipeline stage PAC division standard (Pipeline / Stages / Jobs / Steps), the orchestration relationship (dependency) of pipeline tasks can be abstracted and modeled, thereby mapping the structure and orchestration relationship of the pipeline in the PAC structure to the intermediate transition state of the conversion layer, providing assistance for the conversion of PAC structures suitable for different architectures.
[0026] In a second aspect, the present application provides a pipeline data conversion system. The system comprises:
[0027] An identification module, configured to receive a first pipeline code PAC structure input by a user and a desired output PAC type, and extract task information of an input pipeline task and an arrangement relationship of the input pipeline task according to the first PAC structure;
[0028] A conversion module, used for mapping the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task, wherein the intermediate general task supports mutual conversion with different types of tasks;
[0029] The conversion module is further used to convert the intermediate general task to obtain the target task according to the desired output PAC type;
[0030] An assembling module is used to assemble the target task according to the orchestration relationship to obtain a second PAC structure, the type of the second PAC structure matches the type of the desired output PAC, and the first PAC structure and the second PAC structure are applicable to different cloud architectures.
[0031] In some possible implementations, the conversion module is further used to:
[0032] Mapping the arrangement relationship of the input pipeline task to an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task and the mapped arrangement relationship;
[0033] The conversion module is specifically used for:
[0034] The intermediate PAC structure is transformed according to the desired output PAC type to obtain the target task.
[0035] In some possible implementations, the conversion module is specifically used to:
[0036] Identify task keywords of the intermediate general task;
[0037] Retrieving a keyword mapping set according to the task keyword of the intermediate general task and the PAC type of the desired output, and obtaining a conversion strategy corresponding to the task keyword of the intermediate general task;
[0038] The task keywords of the intermediate general task are transformed according to the transformation strategy to obtain the target task.
[0039] In some possible implementations, the conversion module is specifically used to:
[0040] Inputting the task information of the pipeline task into a PAC conversion model to obtain an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task;
[0041] The conversion module is specifically used for:
[0042] The intermediate general task and the desired output PAC type are input into the PAC conversion model to obtain the target task.
[0043] In some possible implementations, the system further includes:
[0044] A training module is used to obtain training data, wherein the training data includes data pairs of different types of pipeline tasks and intermediate general tasks. Based on the training data, the language model is supervised and fine-tuned to obtain the PAC conversion model.
[0045] In some possible implementations, the system further includes:
[0046] An update module is used to update the PAC knowledge base or the PAC conversion model according to the PAC change information when monitoring the PAC change information of the cloud vendor, wherein the knowledge in the PAC knowledge base is used to construct the training data of the PAC conversion model, or to perform retrieval enhancement on the PAC conversion model to generate RAG.
[0047] In some possible implementations, the PAC change information includes text change information and / or structure change information, and the update module is specifically configured to:
[0048] Determine the type and impact of the PAC change information according to the text change information and / or the structure change information;
[0049] When the type and impact representation of the PAC change information are changed effectively, the PAC change information is updated to the PAC knowledge base in units of tasks.
[0050] In some possible implementations, the loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data, and step layer data.
[0051] In a third aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is used to execute instructions stored in the at least one memory, so that the computing device or the computing device cluster executes the pipeline data conversion method as described in the first aspect or any implementation of the first aspect.
[0052] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, wherein the instructions instruct a computing device or a computing device cluster to execute the pipeline data conversion method described in the above-mentioned first aspect or any one of the implementations of the first aspect.
[0053] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device or a computing device cluster, enables the computing device or the computing device cluster to execute the pipeline data conversion method described in the first aspect or any one of the implementations of the first aspect.
[0054] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required for use in the embodiments are briefly introduced below.
[0056] Figure 1 A schematic diagram of the architecture of a pipeline data conversion system provided in this application;
[0057] Figure 2 A schematic diagram of the implementation details of each component in a pipeline data conversion system provided by this application;
[0058] Figure 3 A flow chart of a pipeline data conversion method provided by this application;
[0059] Figure 4 A flowchart of a task conversion process based on a custom rule conversion engine provided in this application;
[0060] Figure 5 A schematic diagram of a process for task conversion based on a PAC conversion model provided in this application;
[0061] Figure 6 A schematic diagram of a process of updating a PAC knowledge base and adaptively updating a model provided in this application;
[0062] Figure 7 A flowchart of a domain knowledge update process provided for this application;
[0063] Figure 8 A schematic diagram of a model iterative update process provided in this application;
[0064] Fig. 9 A schematic diagram of an application scenario of a pipeline data conversion method provided in this application;
[0065] Fig.10 A schematic diagram of the structure of a pipeline data conversion system provided by this application;
[0066] Fig.11 A schematic diagram of the structure of a computing device provided for this application;
[0067] Fig.12 A schematic diagram of the structure of a computing device cluster provided for this application;
[0068] Fig.13 A schematic diagram of the structure of a computing device cluster provided for this application;
[0069] Fig.14 A schematic diagram of the structure of a computing device cluster provided in this application. DETAILED DESCRIPTION
[0070] The terms "first" and "second" in the embodiments of the present application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0071] First, some technical terms involved in the embodiments of the present application are introduced.
[0072] DevOps, a portmanteau of Development and Operations, is a culture, practice, or methodology that emphasizes communication and cooperation between "software developers (Dev)" and "IT operations and maintenance technicians (Ops)". By automating the processes of "software delivery" and "architecture changes", software building, testing, and release can be faster, more frequent, and more reliable.
[0073] The pipeline is a management tool for software development and release processes. The pipeline can divide the process from development to deployment into multiple stages and automatically transfer data and code between each stage. The pipeline connects the various links in the software development process, such as building, testing, deployment, and monitoring, in an automated and standardized manner, which can help the development and operation teams to automatically manage various processes and tasks in the software development process, such as code compilation, testing, packaging, and deployment. This can reduce errors and delays in manual operations, improve the speed and quality of software delivery, and achieve close collaboration between the development team and the operation team.
[0074] Pipeline As Code (PAC) is a method of defining the software delivery process as code. PAC uses a standardized, structured description language to describe a complete pipeline phase division, dependencies, task details and other information in a code-like style according to specific business keywords or rules. Cloud vendors have launched their own pipeline services, such as PAC services. The pipeline data (such as workflow data) generated by the PAC service is usually data in code form, which can be called a pipeline code structure, that is, a PAC structure. It should be noted that different PAC services can generate PAC structures according to their own defined PAC standard structure specifications. Based on this, the PAC structures generated by different PAC services can be different. For example, the PAC structure generated by the PAC service of vendor A and the PAC structure generated by the PAC service of vendor B can be quite different.
[0075] Many cloud vendors have launched cloud data migration services for cross-cloud service data migration scenarios. The scope of cloud data migration services mainly covers Infrastructure-as-a-Service (IaaS) and Platform-as-a-Service (PaaS), but does not cover Software-as-a-Service (SaaS) such as assembly lines.
[0076] Currently, the industry does not provide an automated data format conversion solution for the PAC structure of DevOps pipelines under different cloud service systems. If users need to migrate the PAC structure of DevOps pipelines across cloud services, they need to perform tedious manual operations and verifications, and the users need to manually complete the conversion of the PAC structure. This greatly reduces the experience and efficiency of using pipelines across clouds, increases the threshold and difficulty of using cloud services, and the probability of errors in manual operations is also high.
[0077] The inventor of this application has found through research that the pipeline services provided by mainstream cloud vendors can basically abstractly summarize and model the task scheduling relationship of the pipeline through the data structure of directed acyclic graph (DAG). In addition, for most common pipeline tasks, such as construction, testing, access control, auditing, etc., they are highly similar and universal.
[0078] To this end, the present application proposes a pipeline data conversion solution based on an intermediate state, which identifies, abstracts, and summarizes the PAC structures provided by different cloud vendors to design a set of universal PAC standards, which can be embodied as loosely coupled abstract data structures. The above-mentioned universal PAC standards or loosely coupled abstract data structures are used to establish a mapping relationship with the PAC structures of different cloud vendors. Based on this, the universal PAC standard can be used as an intermediate layer. When conversion between PAC structures of different formats or types is required, it can be first converted to an intermediate state, and then converted from the intermediate state to a target state. For example, when converting a PAC structure of type A (or format, type) to type B, the PAC structure of type A can be first converted into a broad and universal intermediate state structure, and then the intermediate state structure can be converted into an intermediate state structure of type B.
[0079] Among them, the pipeline data conversion scheme of the present application can be executed by the pipeline data conversion system. The pipeline data conversion system can be a software system, which can be an independent software or integrated into other software. The software system can be provided to the user in the form of a software package and deployed by the user. Or the software system can be provided to the user in the form of a cloud service. Specifically, the software system can expose an API for pipeline data conversion, wherein the API can collaborate with the back-end agent (agent) to achieve pipeline data conversion. The above software system can be deployed in a computing device cluster, and the computing device cluster executes the program code of the software to execute the pipeline data conversion method of the present application. In some possible implementations, the pipeline data conversion system can be a hardware system, such as a computing device cluster with pipeline data conversion capability, and when the computing device cluster is running, the pipeline data conversion method of the present application is executed.
[0080] Specifically, the pipeline data conversion system can receive the first PAC structure input by the user and the desired output PAC type, extract the task information of the input pipeline task and the arrangement relationship of the input pipeline task according to the first PAC structure, and then the pipeline data conversion system maps the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task, wherein the intermediate general task supports mutual conversion with different types of tasks, and converts the intermediate general task according to the desired output PAC type to obtain the target task, and then the pipeline data conversion system can assemble the target task according to the arrangement relationship to obtain a second PAC structure. The type of the second PAC structure matches the desired output PAC type. The first PAC structure and the second PAC structure are applicable to different cloud architectures, for example, they are applicable to PAC formats defined by different cloud vendors.
[0081] The method defines a set of universal pipeline PAC standards, such as a loosely coupled abstract data structure, which is used to establish a mapping relationship with the PAC structure defined by other cloud vendors. In this way, when converting the PAC structures of different cloud vendors, the universal pipeline PAC standard can be used as an intermediate transition layer for conversion to convert the PAC structure input by the user into a transitional intermediate state, and then according to the needs, such as the desired output PAC type, the intermediate state is converted into the desired output PAC structure. The method provides users with an automated PAC structure conversion solution, and in feasible PAC data conversion scenarios, it ensures the accuracy and integrity of the data as much as possible, can assist users in smoothly and quickly converting pipeline data, avoids manual operation, manual conversion, and manual verification by users, saves users a lot of time and energy, and realizes rapid cloud or cross-cloud pipeline data conversion, and provides convenience for data migration and conversion for users who use pipeline services across clouds, thereby improving the user experience.
[0082] In order to make the technical solution of the present application clearer and easier to understand, the architecture of the pipeline data conversion system of the present application is first introduced.
[0083] See also Figure 1 The schematic diagram of the architecture of a pipeline data conversion system is shown. The input of the pipeline data conversion system 100 can be a PAC structure input by a user and a desired PAC type of output, and the output of the pipeline data conversion system 100 can be a desired PAC structure of output. For the sake of distinction, the application may refer to the PAC structure input by the user as a first PAC structure, and the output PAC structure as a second PAC structure. In some possible implementations, the input of the pipeline data conversion system 100 may also include an input PAC type, specifically a type of the first PAC structure.
[0084] The pipeline data conversion system 100 may include an identification component 102, a conversion engine component 104, an assembler component 106, and an output component 108. In some possible implementations, the identification component 102 also has a verification function, and is also called a verification and identification component, denoted as an identifier. The conversion engine component 104 may also be denoted as a Conversion Engine, and the assembler component 106 may also be denoted as a Fabricator. Similar to the identification component 102, the output component 108 may also support verification, and therefore may also be called an output verifier component, denoted as an output verifier.
[0085] Specifically, the user inputs a complete first PAC structure to be converted and provides the desired output PAC type. This enters the PAC conversion process. Optionally, the user can also provide the input PAC type for verification. Figure 1 As shown in the direction of the arrow, the PAC conversion process goes through four main steps: input verification / recognition, conversion, assembly, and output verification. These steps are carried out by the corresponding processing components. The functions of each component are introduced below.
[0086] The identification component 102 is used to receive the first PAC structure input by the user and the desired output PAC type, and extract the task information of the input pipeline task and the arrangement relationship of the input pipeline task according to the first PAC structure. Among them, the pipeline task can refer to the job unit (job) in the pipeline, and the arrangement relationship can be carried or represented by DAG. In some possible implementations, the user can also provide the input PAC type, and accordingly, the identification component 102 is also used to check the legality of the input first PAC structure according to the input PAC type. For example, the identification component 102 can obtain the latest PAC specification matching the PAC type according to the input PAC type, and check whether the first PAC structure complies with the latest PAC specification. When the identification component 102 is used to pass the legality verification, it extracts the task information of the input pipeline task and the arrangement relationship of the input pipeline task according to the first PAC structure. It should be noted that when the legality verification fails, the identification component 102 can output a conversion interrupt.
[0087] The conversion engine component 104 is used to map the task information of the input pipeline task to the loosely coupled abstract data structure to obtain the intermediate general task, which supports mutual conversion with different types of tasks, and then converts the intermediate general task according to the PAC type of the expected output to obtain the target task. Among them, the conversion engine component 104 can convert the input pipeline task according to the parsing result of the recognition component 102, such as the task information of the input pipeline task. In some possible implementations, the conversion engine component 104 can also map the arrangement relationship of the input pipeline task to the intermediate PAC structure. The intermediate PAC structure includes the intermediate general task and the mapped arrangement relationship. Accordingly, the conversion engine component 104 can be used to convert the intermediate PAC structure according to the PAC type of the expected output to obtain the target task.
[0088] Among them, the conversion engine component 104 is used to use the PAC conversion model or a customized conversion rule set to convert the input pipeline task according to the strategy or configuration selection. Among them, the PAC conversion model can be a large language model (LLM) based on pipeline data. The large language model is an artificial intelligence model that can understand and generate human language. In this application, the large language model is trained using a large amount of pipeline data and parameters to achieve the conversion of pipeline tasks. In some examples, when the strategy is cost priority, the conversion engine component 104 can use the PAC conversion model to convert the input pipeline task. When the strategy is quality priority, the conversion engine component 104 can use a customized conversion rule set to convert the input pipeline task.
[0089] Task conversion based on a conversion rule set can be implemented through a custom conversion rule engine. The conversion rule engine can parse the first PAC structure, map the task information of the input pipeline task to a loosely coupled abstract data structure, and obtain an intermediate general task. Then the conversion rule engine identifies the task keywords of the intermediate general task, retrieves the keyword mapping set according to the task keywords of the intermediate general task and the desired output PAC type, and obtains the conversion strategy corresponding to the task keywords of the intermediate general task. Among them, when searching for keywords, the feature keyword set can be detected first, and when the feature keyword set is hit, the keyword mapping set (or keyword mapping relationship set) can be retrieved. The conversion rule engine performs task conversion according to the conversion strategy and can output the target task.
[0090] The task conversion based on the PAC conversion model can be realized through the PAC conversion large model. Specifically, the conversion engine component 104 prepares the input PAC prompt, calls the PAC conversion large model for reasoning, and then obtains the target task output by the PAC conversion large model.
[0091] The assembler component 106 is used to assemble the target task according to the arrangement relationship to obtain a second PAC structure. The type of the second PAC structure matches the type of the desired output PAC. Specifically, the assembler component 106 can assemble the target task output by the conversion component 104 into the PAC data format desired by the user according to the arrangement relationship in the form of DAG according to the output PAC type.
[0092] The output component 108 is used to output the assembled second PAC structure. Further, the output component 108 is also used to perform a legality check or verification on the assembled second PAC structure to ensure the integrity and validity of the converted data and avoid generating an unexecutable PAC structure.
[0093] The implementation details of each component are as follows Figure 2 As shown, the structure of each component and the connection relationship between the components are illustrated below.
[0094] The identification component 102 may include an input PAC identifier, a rule set verifier, a job extractor, and an orchestration relationship extractor (such as a DAGextractor). The complete first PAC structure, the input PAC type, and the desired output PAC type input by the user may be input into the above-mentioned identification component 102, wherein the input PAC type and the first PAC structure may be subjected to a validity check or verification by the rule set verifier after passing through the input PAC identifier. When the validity verification passes, the first PAC structure enters the job extractor and the orchestration relationship extractor respectively, thereby outputting the task information of the input pipeline task and the DAG orchestration relationship of the input pipeline task.
[0095] The conversion engine component 104 may include a DAG regulator, a job conversion strategy selector, a customized rulestrategy converter, and a pipeline job LLM converter. The DAG regulator is used to adjust the orchestration relationship, for example, mapping the DAG orchestration relationship to an intermediate state structure. The job conversion strategy selector is used to select whether to use rules for task conversion or to use a PVC conversion model for conversion. When choosing to use rules for task conversion, the task information of the input pipeline task can be input into the customized rule strategy converter for conversion. When choosing to use the PVC conversion model for task conversion, the task information of the input pipeline task can be input into the pipeline job LLM converter.
[0096] The assembler component 106 may include an integrator, such as a comprehensive integrator. The comprehensive integrator is used to integrate the target task output by the conversion engine component 104 with the orchestration relationship to form a second PVC structure. In some possible implementations, the assembler component 106 also includes an orchestration relationship generator (such as a DAG generator), a task generator (job generator) and a third-party task adapter (third-party job adapter). Among them, the orchestration relationship generator is used to generate an orchestration relationship that matches the output PVC type. The task generator is used to generate a pipeline task of the desired output PVC type. The third-party task adapter is used to adapt the third-party task.
[0097] The output component 108 may include a circular dependency checker, a job format checker, or a third-party dependency checker. The circular dependency checker is used to check whether the arrangement relationship of the pipeline tasks in the second PVC structure includes a circular dependency, the job format checker is used to check whether the pipeline tasks in the second PVC structure comply with the specification, and the third-party dependency checker is used to check whether the third-party dependency is legal.
[0098] When the output component 108 checks the second PVC structure, it can check it in conjunction with the PAC knowledge base, where the PAC knowledge base is also called a cross cloud pipeline domain knowledge data set or a PAC knowledge data set.
[0099] Among them, the knowledge in the PAC knowledge base can be used to construct training data, and the training data can be used to train the PAC conversion model. The pipeline data conversion system 100 also supports updating domain knowledge updates and model iteration updates. Specifically, the pipeline data conversion system 100 can also include a monitoring component, which monitors the PAC standard description document published by the cloud vendor. When PAC change information is monitored, the PAC change information may include text change information and / or structural change information, and the change can be reported. The pipeline data conversion system 100 can determine the type and impact of the PAC change information based on the text change information and / or structural change information. When the type and impact of the PAC change information indicate that the change is valid, the PAC change information is updated to the PAC knowledge base in units of tasks. For example, the pipeline data conversion system 100 can present the above-mentioned PAC change information to the user for manual review and determine whether it needs to be updated. When the user determines that the change is valid and needs to be updated, the pipeline data conversion system 100 automatically updates the updated PAC knowledge to the PAC knowledge base. Among them, the pipeline data conversion system 100 can also trigger the training iteration of the PAC conversion model to ensure that the PAC conversion model can keep in sync with the latest PAC structure of the target cloud vendor and include the latest PAC knowledge.
[0100] Alternatively, the conversion engine component 104 can perform retrieval-augmented generation (RAG) on the PAC conversion model in conjunction with the PAC knowledge base. RAG combines the capabilities of retrieval and generation, introduces external knowledge into the text sequence generation task, and enables the PAC conversion model to dynamically retrieve relevant knowledge from the PAC knowledge base when generating a response, and input the retrieved knowledge and the task information of the input pipeline task into the PAC conversion model for conversion to obtain the target task.
[0101] Based on the aforementioned pipeline data conversion system 100, the present application further provides a data conversion method. The pipeline data conversion method of the present application is introduced below in conjunction with the accompanying drawings.
[0102] See also Figure 3 A flow chart of a pipeline data conversion method is shown, the method comprising the following steps:
[0103] S302 : The pipeline data conversion system 100 receives a first PAC structure input by a user and a desired output PAC type.
[0104] The first PAC structure is a complete PAC structure to be converted, which may include a complete PAC code. In a cross-cloud migration scenario, when a user intends to migrate the service data of cloud vendor A to cloud vendor B, the user may input the first PAC structure and the desired output PAC type to convert it into a second PAC structure compatible with cloud vendor B. The desired output PAC type may be a PAC type (or PAC format, PAC type) defined by cloud vendor B.
[0105] In some possible implementations, the pipeline data conversion system 100 can also receive the PAC type of the first PAC structure input by the user, that is, the input PAC type. Accordingly, the pipeline data conversion system 100 can also perform a validity check on the input first PAC structure according to the input PAC type. Among them, the pipeline data conversion system 100 can first retrieve the PAC knowledge set according to the PAC type input by the user, obtain a subset related to the input PAC type, and perform a validity check on the input first PAC structure according to the PAC knowledge in the subset. When the validity check passes, the subsequent steps can be continued; when the validity check fails, the conversion process can be terminated.
[0106] S304. The pipeline data conversion system 100 extracts the task information of the input pipeline task and the arrangement relationship of the input pipeline task according to the first PAC structure.
[0107] Specifically, the pipeline data conversion system 100 can input the first PAC structure into the DAG extractor and the task extractor respectively, extract the arrangement relationship of the input pipeline task through the DAG extractor, and extract the task information of the input pipeline task through the task extractor. The task information may include the task name, the stage where the task is located, the pipeline where the task is located, and the steps included in the task.
[0108] S306 , the pipeline data conversion system 100 maps the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task.
[0109] Intermediate general tasks support mutual conversion with different types of tasks. Specifically, this application abstracts and models pipeline tasks in pipeline services of different cloud vendors to obtain loosely coupled abstract data structures. The loosely coupled abstract data structures are applicable to different cloud vendors and can carry different cloud vendor data formats and different types of tasks as intermediate states in the conversion process for subsequent task conversion.
[0110] The loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data and step layer data. The pipeline data conversion system 100 maps the task information of the input pipeline task to the corresponding level of the loosely coupled abstract data structure to obtain the intermediate general task.
[0111] Considering the different hierarchical relationships of different cloud vendors, for example, some cloud vendors have a 4-layer structure, and some cloud vendors have a 3-layer structure, so the orchestration relationship can be adaptively adjusted. In specific implementation, the pipeline data conversion system 100 can also map the orchestration relationship of the input pipeline task to the intermediate PAC structure, and the intermediate PAC structure includes the intermediate general task and the mapped orchestration relationship.
[0112] The process of obtaining the intermediate PAC structure is described in detail below.
[0113] The pipeline data conversion system 100 can use the DAG data structure and combine the PAC division standard (Pipeline / Stages / Jobs / Steps) of the four-layer pipeline stage customized by this system to abstract and model the orchestration relationship (dependency) of the pipeline tasks, and map the structure and orchestration relationship of the pipeline in the first PAC structure to the intermediate transition state of the conversion layer. At the same time, after extracting the task information of the input pipeline task in the first PAC structure, the pipeline data conversion system 100 can mark the input pipeline task in the first PAC structure input by the user according to the task information knowledge of the pipeline task entered in the PAC knowledge set, determine the type of the input pipeline task, and map it to the PAC structure customized by this system to obtain the intermediate state PAC structure.
[0114] S308 , the pipeline data conversion system 100 converts the intermediate general task according to the desired output PAC type to obtain the target task.
[0115] The converter component of the pipeline data conversion system 100 provides a variety of processing sub-links for task conversion. Among them, the various processing sub-links may include a sub-link for task conversion using a custom rule conversion engine and a sub-link for task conversion using a PAC conversion model. The pipeline data conversion system 100 can select a sub-link for task conversion according to the policy configured by the system. For example, when the policy is cost priority, the pipeline data conversion system 100 can choose a sub-link for task conversion using a PAC conversion model. For another example, when the policy is quality priority, the pipeline data conversion system 100 can choose a sub-link for task conversion using a custom rule conversion engine.
[0116] The following is an example of the processing of the sub-link of task conversion using the custom rule conversion engine and the sub-link of task conversion using the PAC conversion model.
[0117] In some possible implementations, the pipeline data conversion system 100 can identify the task keywords of the intermediate general task, and then retrieve a keyword mapping set based on the task keywords of the intermediate general task and the desired output PAC type to obtain a conversion strategy corresponding to the task keywords of the intermediate general task, and then convert the task keywords of the intermediate general task according to the conversion strategy to obtain the target task.
[0118] like Figure 4 As shown, the converter component of the pipeline data conversion system 100 can input the extracted intermediate general tasks and the PAC types that the user expects to output into the custom rule conversion engine. The custom rule conversion engine may include a custom rule set, and the custom rule set includes a feature keyword set and a keyword mapping set. Among them, the feature keyword set includes task keywords of different pipeline tasks and feature keywords of PAC structures of different systems, and the keyword mapping set includes the mapping relationship of keywords and the corresponding conversion strategy. For example, the keyword mapping set may include the mapping relationship between the pipeline task keywords of the pipeline service of cloud vendor A and the pipeline task keywords of the pipeline service of cloud vendor B, as well as the conversion strategy from the pipeline task keywords of cloud vendor A to the pipeline task keywords of cloud vendor B.
[0119] The custom rule conversion engine first parses the incoming intermediate general tasks, for example, it traverses the incoming intermediate general tasks, thereby parsing each intermediate general task and obtaining the task keywords of each intermediate general task. The custom rule conversion engine retrieves the characteristic keyword set in the manually entered custom rule set according to the task keywords of the intermediate general tasks. When the characteristic keyword set is hit, subsequent operations can be performed. For example, the custom rule conversion engine can retrieve the keyword mapping set according to the keywords in the task and the type of PAC expected by the user, and find the corresponding keyword of the keyword in the corresponding PAC system. The custom rule conversion engine can determine the conversion strategy of the intermediate general task based on the mapping relationship between the type of intermediate general task and the keyword, and then convert the pipeline task in the PAC data format expected by the user, and assign values to each attribute or field of the task.
[0120] It should be noted that Figure 4The example of converting the intermediate general task into the target task by using the custom rule conversion engine is used. In actual application, the input pipeline task is mapped to a loosely coupled abstract data structure, and the intermediate general task can also be obtained through the custom rule conversion engine. In other words, the converter component can also input the task information of the input pipeline task and the desired output PAC type into the custom rule conversion engine.
[0121] In some other possible implementations, the pipeline data conversion system 100 may input the task information of the pipeline task into the PAC conversion model to obtain an intermediate PAC structure, wherein the intermediate PAC structure includes an intermediate general task. Further, the pipeline data conversion system 100 inputs the intermediate general task and the desired output PAC type into the PAC conversion model to obtain a target task.
[0122] like Figure 5 As shown, the converter component includes a large model conversion engine, and the large model conversion engine includes a PAC conversion model. The PAC conversion model supports the mutual conversion between the PAC structure of the cloud vendor and the general intermediate state PAC structure, or the mutual conversion between the PAC type and PAC format of the cloud vendor and the general PAC type and PAC format. Specifically, the converter component can input the task information of the input pipeline task, the input PAC type, and the expected output PAC type into the large model conversion engine. The large model conversion engine first inputs the input PAC type and the task information of the input pipeline task into the PAC conversion model for conversion to obtain the task information of the intermediate state general task. Then the large model conversion engine inputs the expected output PAC type and the task information of the intermediate state general task into the PAC conversion model to obtain the target task. The target task can be a pipeline task of the target PAC system. In this way, the PAC2PAC conversion function can be realized through the intermediate state transfer.
[0123] In the middle stage of the conversion process, this system uses the DAG data structure to carry the orchestration relationship of the pipeline tasks, and uses the abstract loosely coupled data structure to carry the specific pipeline tasks. The complete input PAC structure is disassembled from these two aspects, and the input PAC structure is mapped to the abstract intermediate PAC structure customized by the cost system. The processed intermediate PAC structure and the desired output PAC type are passed to the converter component, and the converter component converts the intermediate PAC structure (including the intermediate general task) into the pipeline task of the PAC type that the user expects to output, specifically the aforementioned target task.
[0124] S310 , the pipeline data conversion system 100 assembles the target tasks according to the orchestration relationship to obtain a second PAC structure.
[0125] The type of the second PAC structure matches the type of the desired output PAC. Specifically, the pipeline data conversion system 100 can refer to the arrangement relationship of the input pipeline tasks to arrange the target task and obtain a complete second PAC structure. In some possible implementations, the pipeline data conversion system 100 can also arrange the target task through the converted arrangement relationship to obtain a complete second PAC structure. It should be noted that when the pipeline data conversion system 100 is assembling tasks, it can also retrieve the PAC knowledge base according to the type of the desired output PAC, and combine the retrieved knowledge to assemble the target task to obtain the second PAC structure. This can improve assembly efficiency and assembly accuracy.
[0126] In some possible implementations, the pipeline data conversion system 100 can also verify the assembled second PAC structure in combination with the PAC knowledge base. Specifically, the pipeline data conversion system 100 checks whether the assembled second PAC structure complies with the format specification, whether the arrangement relationship conflicts, and whether the data of the pipeline task is complete and valid. The pipeline data conversion system 100 can return the second PAC structure to the user after passing the verification to avoid outputting an erroneous and invalid second PAC structure.
[0127] The first PAC structure and the second PAC structure are applicable to different cloud architectures. For example, the first PAC structure is applicable to the PAC format defined by the first cloud vendor, and the second PAC structure is applicable to the PAC format defined by the second cloud vendor. In this way, PAC migration between different cloud vendors can be achieved.
[0128] Based on the above description, the pipeline data conversion method of the present application defines a set of universal PAC standards, such as a loosely coupled abstract data structure, which may include four-layer data structures such as Pipeline / Stages / Jobs / Steps, etc. By mapping the third-party pipeline to the PAC standard system defined by this system, it provides abstract modeling and summarization capabilities for different third-party pipelines. When converting the PAC structures of different cloud vendors, the above universal PAC standard can be used as an intermediate transition layer for conversion to convert the PAC structure input by the user into a transitional intermediate state, and then according to the desired output PAC type, the intermediate state is converted into the desired output PAC structure, thereby providing users with an automated PAC structure conversion solution, providing convenience for data migration and conversion for users who use pipeline services across clouds, and improving the user experience.
[0129] Considering that the PAC structure of cloud vendors can continue to evolve, the pipeline data conversion system 100 also supports PAC knowledge base updates and model adaptive updates.
[0130] See also Figure 6 A flow chart of a PAC knowledge base update and a model adaptive update is shown. The pipeline data conversion system 100 can monitor the documents of the cloud vendor. When PAC change information is monitored from the documents, the PAC knowledge base can be updated according to the PAC change information, thereby realizing the update of domain knowledge. For example, the pipeline data conversion system 100 can extract updated knowledge according to the PAC change information, thereby updating the PAC knowledge base. It should be noted that the pipeline data conversion system 100 can also trigger the update of the PAC conversion model according to the PAC change information. Specifically, after updating the PAC knowledge base, the pipeline data conversion system 100 can obtain training data through data processing, and the pipeline data conversion system 100 can fine-tune the language model according to the training data, such as supervised fine-tuning (FT), thereby obtaining a PAC conversion model.
[0131] The process of domain knowledge updating and model iterative updating is described in detail below.
[0132] The purpose of domain knowledge update is to obtain the latest PAC change information of cloud vendors, so that the conversion results of this system are synchronized with the documents of cloud vendors to prevent incompatibility or inefficiency. To this end, the pipeline data conversion system 100 can continuously monitor, analyze, test and verify the changes of cloud vendors, and timely feedback and correct the conversion results of this system.
[0133] See also Figure 7 The flowchart of a domain knowledge update shown in the figure shows that in order to capture the dynamic changes of the PAC field structure of the target cloud vendor in real time, the system is designed with a data monitoring function during the implementation process. By starting the crawler task at a fixed time, the PAC text information and PAC structure information are obtained from the PAC description documents published by each cloud vendor through identification and extraction, and compared with the last monitoring results to remove duplications to determine whether any changes have occurred. For the discovered change results (such as text change information and / or structure change information), the system can enter the manual analysis and review stage. After confirmation and marking, the system can determine the type and impact of the change, determine whether it is a valid PAC update, and store the updated task information in the database in units of pipeline tasks (such as PAC-JOB) to obtain the iterated PAC knowledge base. It should be noted that the above changes may include but are not limited to the addition, deletion, and modification of tasks.
[0134] Among them, the PAC knowledge base stores the task information of the pipeline tasks of the supported cloud vendors and the structural information of the tasks. In order to provide PAC2PAC training data for the PAC conversion model, it is usually necessary to process the data and construct it into the PAC2PAC format to generate structured data, specifically {cloud vendor task-intermediate general task} data pairs, which are used for iterative updates of the model. It should be noted that the mapping of {cloud vendor task-intermediate general task} is mutual. If the task P of cloud vendor A can be mapped to the intermediate general task X, then the intermediate general task X can be mapped to the task P of cloud vendor A. It is expressed as: A+P→X, A+X→P.
[0135] Next, see Figure 8 The flowchart of a model iteration update is shown in FIG. 1 . After the domain knowledge is updated, the automatic fine-tuning training of the PAC conversion model can be triggered to realize the dynamic update of the model and adapt to the new data distribution and task requirements. First, the pipeline data conversion system 100 uses the structured data {cloud vendor job-intermediate task} data pair as training data to construct a prompt input prompt to guide the PAC conversion model to generate the desired output. Then, the training data is used as the data set input to supervise and fine-tune the PAC conversion model to learn the task characteristics of PAC conversion.
[0136] In order to make the technical solution of the present application clearer and easier to understand, the pipeline data conversion method of the present application is explained below in conjunction with an application scenario.
[0137] See also Fig. 9 FIG. 1 is a schematic diagram of an application scenario of a pipeline data conversion method shown in FIG. 1 . In this scenario, after desensitization, the input PAC structure can be expressed as follows:
[0138]
[0139]
[0140] The above PVC structure first defines the correspondence between flow, stage, and job in the pipeline, and then defines the attribute fields of the job.
[0141] The above PVC structure can be verified by the input verifier. When the verification is passed, it can enter the DAG extractor and the task extractor to extract the DAG arrangement relationship and the task information of the input pipeline task respectively. Among them, the task information of the input pipeline task and the DAG arrangement relationship can be mapped to form intermediate state data, such as an intermediate state structure. The intermediate state structure can be input into the converter component for conversion to obtain the target task. The target task can be assembled by the assembler and verified by the output verification component to output a complete PAC structure that matches the PAC type expected by the user. The desensitized code is as follows:
[0142]
[0143]
[0144]
[0145] Compared with the input PVC structure, the output PVC structure reintegrates the pipeline stages and adopts the pipeline tasks of the corresponding PVC type.
[0146] Based on the aforementioned pipeline data conversion method, the present application also provides a pipeline data conversion system. Fig.10 As shown, the pipeline data conversion system 100 includes:
[0147] The identification module 1002 is used to receive a first pipeline code PAC structure input by a user and a desired output PAC type, and extract task information of an input pipeline task and an arrangement relationship of the input pipeline task according to the first PAC structure;
[0148] The conversion module 1004 is used to map the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task, and the intermediate general task supports mutual conversion with different types of tasks;
[0149] The conversion module 1004 is further used to convert the intermediate general task to obtain a target task according to the desired output PAC type;
[0150] The assembling module 1006 is used to assemble the target task according to the orchestration relationship to obtain a second PAC structure, the type of the second PAC structure matches the type of the desired output PAC, and the first PAC structure and the second PAC structure are applicable to different cloud architectures.
[0151] Among them, the identification module 1002, the conversion module 1004, and the assembly module 1006 can be implemented by software or by hardware.
[0152] When implemented by software, the identification module 1002, the conversion module 1004, and the assembly module 1006 may be applications running on a computer device. For example, the identification module 1002 may be the aforementioned identification component 102, the conversion module 1004 may be the aforementioned conversion engine component 104, and the assembly module 1006 may be the aforementioned organizer component 106. The above-mentioned application may also be virtualized through a virtualization service to be provided to users for use. Virtualization services may include virtual machine (VM) services, bare metal server (BMS) services, and container services. Among them, the VM service may be a service that virtualizes a virtual machine (VM) resource pool on multiple physical hosts through virtualization technology to provide users with VMs for use on demand. The BMS service is a service that virtualizes a BMS resource pool on multiple physical hosts to provide users with BMS for use on demand. The container service is a service that virtualizes a container resource pool on multiple physical hosts to provide users with containers for use on demand. VM is a simulated virtual computer, that is, a logical computer. BMS is a high-performance computing service that can be elastically scaled. Its computing performance is no different from that of traditional physical machines, and it has the characteristics of secure physical isolation. Container is a kernel virtualization technology that can provide lightweight virtualization to achieve the purpose of isolating user space, processes and resources. It should be understood that the VM service, BMS service and container service in the above virtualization services are only used as specific examples. In actual applications, virtualization services can also be other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0153] When implemented by hardware, the identification module 1002, the conversion module 1004, and the assembly module 1006 may include at least one computing device, such as a server, etc. Alternatively, the identification module 1002, the conversion module 1004, and the assembly module 1006 may also be implemented by using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0154] In some possible implementations, the conversion module 1004 is further configured to:
[0155] Mapping the arrangement relationship of the input pipeline task to an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task and the mapped arrangement relationship;
[0156] The conversion module 1004 is specifically used for:
[0157] The intermediate PAC structure is transformed according to the desired output PAC type to obtain the target task.
[0158] In some possible implementations, the conversion module 1004 is specifically used to:
[0159] Identify task keywords of the intermediate general task;
[0160] Retrieving a keyword mapping set according to the task keyword of the intermediate general task and the PAC type of the desired output, and obtaining a conversion strategy corresponding to the task keyword of the intermediate general task;
[0161] The task keywords of the intermediate general task are transformed according to the transformation strategy to obtain the target task.
[0162] In some possible implementations, the conversion module 1004 is specifically used to:
[0163] Inputting the task information of the pipeline task into a PAC conversion model to obtain an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task;
[0164] The conversion module 1004 is specifically used for:
[0165] The intermediate general task and the desired output PAC type are input into the PAC conversion model to obtain the target task.
[0166] In some possible implementations, the system 100 further includes:
[0167] The training module 1008 is used to obtain training data, wherein the training data includes data pairs of different types of pipeline tasks and intermediate general tasks. Based on the training data, the language model is supervised and fine-tuned to obtain the PAC conversion model.
[0168] The training module 1008 may be implemented by software or hardware.
[0169] When implemented by software, the training module 1008 may be an application program running on a computer device. The training module 1008 may also be virtualized through a virtualization service, such as a VM service, a BMS service, or a container service, to provide it to a user. When implemented by hardware, the training module 1008 may include at least one computing device, such as a server, etc. Alternatively, the training module 1008 may also be a device implemented by an application-specific integrated circuit ASIC, or a programmable logic device PLD, etc.
[0170] In some possible implementations, the system 100 further includes:
[0171] The update module 1009 is used to update the PAC knowledge base or the PAC conversion model according to the PAC change information when monitoring the PAC change information of the cloud vendor, wherein the knowledge in the PAC knowledge base is used to construct the training data of the PAC conversion model, or to perform retrieval enhancement on the PAC conversion model to generate RAG.
[0172] The update module 1009 may be implemented by software or hardware.
[0173] When implemented by software, the update module 1009 may be an application program running on a computer device. The update module 1009 may also be virtualized through a virtualization service, such as a VM service, a BMS service, or a container service, to provide it to a user. When implemented by hardware, the update module 1009 may include at least one computing device, such as a server, etc. Alternatively, the update module 1009 may also be a device implemented by an application specific integrated circuit ASIC, or a programmable logic device PLD, etc.
[0174] In some possible implementations, the PAC change information includes text change information and / or structure change information, and the updating module 1009 is specifically configured to:
[0175] Determine the type and impact of the PAC change information according to the text change information and / or the structure change information;
[0176] When the type and impact representation of the PAC change information are changed effectively, the PAC change information is updated to the PAC knowledge base in units of tasks.
[0177] In some possible implementations, the loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data, and step layer data.
[0178] The present application also provides a computing device 1100. Fig.11As shown, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate through the bus 1102. The computing device 1100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1100.
[0179] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (eg, the memory 1106, the processor 1104, and the communication interface 1108).
[0180] The processor 1104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0181] The memory 1106 may include a volatile memory, such as a random access memory (RAM). The memory 1106 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the aforementioned pipeline data conversion method. Specifically, the memory 1106 stores instructions for the pipeline data conversion system 100 to execute the pipeline data conversion method.
[0182] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.
[0183] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0184] like Fig.12 As shown, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same pipeline data conversion system 100 for executing instructions of the pipeline data conversion method.
[0185] In some possible implementations, one or more computing devices 1100 in the computing device cluster may also be used to execute some instructions of the pipeline data conversion system 100 for executing the pipeline data conversion method. In other words, a combination of one or more computing devices 1100 may jointly execute instructions of the pipeline data conversion system 100 for executing the pipeline data conversion method.
[0186] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster may store different instructions for executing partial functions of the pipeline data conversion system 100 .
[0187] Fig.13 A possible implementation is shown. Fig.13 As shown, two computing devices 1100A and 1100B are connected via a communication interface 1108. The memory in the computing device 1100A stores instructions for executing the functions of the identification module 1002. The memory in the computing device 1100B stores instructions for executing the functions of the conversion module 1004 and the assembly module 1006. In other words, the memories 1106 of the computing devices 1100A and 1100B jointly store instructions for the pipeline data conversion system 100 to execute the pipeline data conversion method. Among them, the memory in the computing device 1100A can also store instructions for executing the functions of the training module 1008 and the update module 1009.
[0188] Fig.13 The connection mode between the computing device clusters shown may be considered to be that the pipeline data conversion method provided by the present application requires a large amount of computing power to perform pipeline task conversion. Therefore, it is considered that the functions implemented by the conversion module 1004 and the assembly module 1006 are handed over to the computing device 1100B for execution.
[0189] It should be understood that Fig.13The functions of the computing device 1100A shown in FIG. 1100A may also be completed by multiple computing devices 1100. Similarly, the functions of the computing device 1100B may also be completed by multiple computing devices 1100.
[0190] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.14 A possible implementation is shown. Fig.14 As shown, two computing devices 1100C and 1100D are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1106 in the computing device 1100C stores instructions for executing the functions of the identification module 1002. At the same time, the memory 1106 in the computing device 1100D stores instructions for executing the functions of the conversion module 1004 and the assembly module 1006. Similarly, the memory 1106 in the computing device 1100C also stores instructions for executing the functions of the training module 1008 and the update module 1009.
[0191] Fig.14 The connection method between the computing device clusters shown can be based on the consideration that the pipeline data conversion method provided in this application requires a large amount of computing power to perform pipeline task conversion, so it is considered to hand over the functions implemented by the conversion module 1004 and the assembly module 1006 to the computing device 1100D for execution.
[0192] It should be understood that Fig.14 The functions of the computing device 1100C shown in FIG. 1100A may also be completed by multiple computing devices 1100. Similarly, the functions of the computing device 1100D may also be completed by multiple computing devices 1100.
[0193] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned pipeline data conversion system for executing the pipeline data conversion method.
[0194] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computer device, the at least one computer device executes the above-mentioned pipeline data conversion method.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pipeline data conversion method, characterized in that: The method comprises: Receive a first pipeline code PAC structure input by a user and a desired output PAC type; Extracting task information of an input pipeline task and an arrangement relationship of the input pipeline task according to the first PAC structure; Mapping the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task, wherein the intermediate general task supports mutual conversion with different types of tasks; According to the desired output PAC type, the intermediate general task is transformed to obtain a target task; The target tasks are assembled according to the orchestration relationship to obtain a second PAC structure, the type of the second PAC structure matches the type of the desired output PAC, and the first PAC structure and the second PAC structure are applicable to different cloud architectures.
2. The method according to claim 1, characterized in that: The method further comprises: Mapping the orchestration relationship of the input pipeline task to an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task and the mapped orchestration relationship; The converting the intermediate general task according to the desired output PAC type to obtain the target task includes: The intermediate PAC structure is transformed according to the desired output PAC type to obtain the target task.
3. The method according to claim 1 or 2, characterized in that: The converting the intermediate general task according to the desired output PAC type to obtain the target task includes: Identify task keywords of the intermediate general task; Retrieving a keyword mapping set according to the task keyword of the intermediate general task and the PAC type of the desired output, and obtaining a conversion strategy corresponding to the task keyword of the intermediate general task; The task keywords of the intermediate general task are transformed according to the transformation strategy to obtain the target task.
4. The method according to claim 1 or 2, characterized in that: Mapping the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task includes: Inputting the task information of the pipeline task into a PAC conversion model to obtain an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task; The converting the intermediate general task according to the desired output PAC type to obtain the target task includes: The intermediate general task and the desired output PAC type are input into the PAC conversion model to obtain the target task.
5. The method according to claim 4, characterized in that The PAC conversion model was constructed as follows: Acquire training data, wherein the training data includes data pairs of different types of pipeline tasks and intermediate general tasks; Based on the training data, supervised fine-tuning is performed on the language model to obtain the PAC conversion model.
6. The method according to claim 4 or 5, characterized in that: The method further comprises: When PAC change information from the cloud vendor is monitored, the PAC knowledge base or the PAC conversion model is updated according to the PAC change information, wherein the knowledge in the PAC knowledge base is used to construct training data for the PAC conversion model, or to perform retrieval enhancement on the PAC conversion model to generate RAG.
7. The method according to claim 6, characterized in that The PAC change information includes text change information and / or structure change information, and updating the PAC knowledge base according to the PAC change information includes: Determine the type and impact of the PAC change information according to the text change information and / or the structure change information; When the type and impact representation of the PAC change information are changed effectively, the PAC change information is updated to the PAC knowledge base in units of tasks.
8. The method according to any one of claims 1 to 7, characterized in that: The loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data and step layer data.
9. A pipeline data conversion system, characterized in that: The system comprises: An identification module, configured to receive a first pipeline code PAC structure input by a user and a desired output PAC type, and extract task information of an input pipeline task and an arrangement relationship of the input pipeline task according to the first PAC structure; A conversion module, used for mapping the task information of the input pipeline task to a loosely coupled abstract data structure to obtain an intermediate general task, wherein the intermediate general task supports mutual conversion with different types of tasks; The conversion module is further used to convert the intermediate general task to obtain the target task according to the desired output PAC type; An assembling module is used to assemble the target task according to the orchestration relationship to obtain a second PAC structure, the type of the second PAC structure matches the type of the desired output PAC, and the first PAC structure and the second PAC structure are applicable to different cloud architectures.
10. The system according to claim 9, characterized in that The conversion module is also used for: Mapping the orchestration relationship of the input pipeline task to an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task and the mapped orchestration relationship; The conversion module is specifically used for: The intermediate PAC structure is transformed according to the desired output PAC type to obtain the target task.
11. The system according to claim 9 or 10, characterized in that: The conversion module is specifically used for: Identify task keywords of the intermediate general task; Retrieving a keyword mapping set according to the task keyword of the intermediate general task and the PAC type of the desired output, and obtaining a conversion strategy corresponding to the task keyword of the intermediate general task; The task keywords of the intermediate general task are transformed according to the transformation strategy to obtain the target task.
12. The system according to claim 9 or 10, characterized in that The conversion module is specifically used for: Inputting the task information of the pipeline task into a PAC conversion model to obtain an intermediate PAC structure, wherein the intermediate PAC structure includes the intermediate general task; The conversion module is specifically used for: The intermediate general task and the desired output PAC type are input into the PAC conversion model to obtain the target task.
13. The system according to claim 12, characterized in that The system further comprises: A training module is used to obtain training data, wherein the training data includes data pairs of different types of pipeline tasks and intermediate general tasks. Based on the training data, the language model is supervised and fine-tuned to obtain the PAC conversion model.
14. The system according to claim 12 or 13, characterized in that The system further comprises: An update module is used to update the PAC knowledge base or the PAC conversion model according to the PAC change information when monitoring the PAC change information of the cloud vendor, wherein the knowledge in the PAC knowledge base is used to construct the training data of the PAC conversion model, or to perform retrieval enhancement on the PAC conversion model to generate RAG.
15. The system according to claim 14, characterized in that The PAC change information includes text change information and / or structure change information, and the update module is specifically used for: Determine the type and impact of the PAC change information according to the text change information and / or the structure change information; When the type and impact representation of the PAC change information are changed effectively, the PAC change information is updated to the PAC knowledge base in units of tasks.
16. The system according to any one of claims 9 to 15, characterized in that The loosely coupled abstract data structure includes pipeline layer data, stage layer data, task layer data and step layer data.
17. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, and the at least one computing device includes at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions so that the computing device cluster performs the pipeline data conversion method as described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: Comprising computer-readable instructions; the computer-readable instructions are used to implement the pipeline data conversion method described in any one of claims 1 to 8.
19. A computer program product, characterized in that Comprising computer-readable instructions; the computer-readable instructions are used to implement the pipeline data conversion method described in any one of claims 1 to 8.