Data pipeline generation method and apparatus based on cloud computing technology

Through the data pipeline generation method based on cloud computing technology, the problem description input by the user is used to generate a consistent data pipeline, which solves the problem of cumbersome operation of traditional data pipeline products and improves the generation efficiency and user experience.

WO2025102618A1PCT designated stage expired Publication Date: 2025-05-22HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/091137
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-05-06
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In traditional data pipeline products, users need tedious operational steps to create data pipelines, and the learning cost and usage threshold are relatively high, making it difficult for users to quickly and efficiently generate the required data pipelines.

Method used

The data pipeline generation method based on cloud computing technology is adopted to generate characterization vectors through the problem description input by users, search the pipeline templates in the template library, and generate data pipelines that match the problem description based on similarity, simplifying the user's operation process.

Benefits of technology

It improves the generation efficiency of data pipelines, reduces the difficulty of generation, and allows users to do not need to enter detailed operation instructions or click actions. It is recommended to users' most relevant and useful data pipelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024091137_22052025_PF_FP_ABST
    Figure CN2024091137_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A data pipeline generation method based on a cloud computing technology, applied to a cloud computing platform. The cloud computing platform runs on an infrastructure, the infrastructure comprises a plurality of data centers arranged in different areas, and each data center comprises a plurality of servers. The method comprises: on the basis of a question description input by a user, generating a first representation vector used for representing the question description, the question description being used for describing content of a data pipeline that a user desires to use; on the basis of the similarity between the first representation vector and a representation vector of each pipeline template in a template library, retrieving at least one pipeline template from the template library; and on the basis of the retrieved pipeline template, generating a data pipeline conforming to the question description. Therefore, when the user inputs the question description, the most related and most useful data pipeline can be automatically recommended to the user, so that the user does not need to input a detailed operation instruction or a click action and the like, thereby improving the generation efficiency of data pipelines, and reducing the generation difficulty of data pipelines.
Need to check novelty before this filing date? Find Prior Art

Description

A data pipeline generation method and device based on cloud computing technology

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 16, 2023, with application number 202311536346.9 and application name “A data pipeline creation method”, and the Chinese patent application filed with the State Intellectual Property Office of China on January 4, 2024, with application number 202410012950.X and application name “A data pipeline generation method and device based on cloud computing technology”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of information technology (IT), and in particular to a data pipeline generation method and device based on cloud computing technology. Background Art

[0003] A data pipeline is an automated process for extracting, transforming, and loading data from raw data sources into a target data warehouse or data lake. It is a critical tool used by data engineers and data scientists to process and manage large amounts of data. Automating data pipelines can improve the efficiency and accuracy of data processing, reducing manual errors and duplication of work. However, traditional data pipeline products require users to create data pipelines through cumbersome steps, resulting in a high learning curve and a high barrier to entry. Therefore, streamlining the data pipeline creation process is a pressing technical challenge.

[0004] Summary of the Invention

[0005] The present application provides a data pipeline generation method, apparatus, computing device cluster, computer storage medium and computer product based on cloud computing technology, which can simplify the process of creating a data pipeline.

[0006] In a first aspect, the present application provides a data pipeline generation method based on cloud computing technology, which is applied to a cloud computing platform. The cloud computing platform runs on an infrastructure, and the infrastructure includes multiple data centers located in different regions, each data center including multiple servers. The method includes: the cloud computing platform generates a first representation vector for representing the problem description based on a problem description input by a user, wherein the problem description is used to describe the content of the data pipeline that the user desires to use; the cloud computing platform retrieves at least one pipeline template from a template library based on the similarity between the first representation vector and the representation vectors of the pipeline templates in the template library; and the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template.

[0007] In this way, when a user has a data pipeline requirement, after the user enters a problem description, the most relevant and useful data pipeline will be recommended to the user based on the problem description entered by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., which improves the generation efficiency of the data pipeline and reduces the difficulty of generating the data pipeline.

[0008] In one possible implementation, the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline templates, including: when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to a target threshold, the cloud computing platform generates a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; when the maximum value is less than the target threshold, the cloud computing platform retrieves at least one job template from the template library based on the similarity between the first representation vector and the representation vector of the job template in the template library, and generates a data pipeline that matches the problem description based on the retrieved pipeline templates and the job templates. A job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template. In this way, the most relevant and useful data pipeline can be recommended to users in different situations.

[0009] In one possible implementation, the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates. This includes: the cloud computing platform, based on business rules, filters out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates, respectively; and the cloud computing platform generates a data pipeline that matches the problem description based on the filtered pipeline templates and job templates. This ensures that the generated data pipeline matches the business described in the problem, thereby recommending the most relevant and useful data pipeline to the user.

[0010] In one possible implementation, the cloud computing platform generates a first representation vector for representing the problem description based on a problem description input by a user. This includes: the cloud computing platform converting the problem description into at least one task; and generating the first representation vector based on keywords included in the at least one task. In this way, the problem description can be converted into tasks, keywords can be extracted from the tasks, and finally, the first representation vector for representing the problem description can be generated from the keywords.

[0011] In one possible implementation, the cloud computing platform generates a first representation vector based on the user's input question description. This includes: the cloud computing platform obtains the user's interest representation vector, which is used to represent the user's preference for data pipelines; and the cloud computing platform generates the first representation vector based on the interest representation vector and the question description. This ensures that the templates retrieved subsequently match the user's preferences, and the desired data pipeline can be recommended to the user.

[0012] In one possible implementation, before the cloud computing platform generates a first representation vector for representing the problem description based on the user-input problem description, the method further includes: the cloud computing platform receiving a seed job template imported by the user; the cloud computing platform generating a new job template based on the seed job template and the sampled problem description by changing the prompt word; and the cloud computing platform storing the new job template in the template library after verifying the legitimacy of the new job template. In this way, job templates can be automatically generated, improving the efficiency of job template generation.

[0013] In one possible implementation, before the cloud computing platform generates a first representation vector for representing the problem description based on the user-input problem description, the method further includes: the cloud computing platform receiving a seed pipeline template imported by the user; the cloud computing platform generating a new pipeline template based on the seed pipeline template and a job template sampled from a template library by changing the prompt word; and after verifying the legality of the new pipeline template, the cloud computing platform storing the new pipeline template in the template library. In this way, pipeline templates can be automatically generated, improving the efficiency of pipeline template generation and resolving the problem of insufficient pipeline templates.

[0014] In the second aspect, the present application provides a data pipeline generation device based on cloud computing technology, which is deployed on a cloud computing platform. The cloud computing platform runs on an infrastructure, and the infrastructure includes multiple data centers located in different areas, and each data center includes multiple servers. The device includes: a vector representation module, a template retrieval module, and a pipeline generation module. Among them, the vector representation module is used to generate a first representation vector for representing the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use. The template retrieval module is used to retrieve at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library. The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template.

[0015] In one possible implementation, when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value.

[0016] If the maximum value is less than the target threshold, the template retrieval module is used to retrieve at least one job template from the template library based on the similarity between the first representation vector and the representation vectors of the job templates in the template library. The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates. A job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template.

[0017] In one possible implementation, when the pipeline generation module generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, it is specifically used to: based on business rules, filter out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates respectively; and generate a data pipeline that matches the problem description based on the filtered pipeline templates and job templates.

[0018] In one possible implementation, when the vector representation module generates a first representation vector of a problem description based on a problem description input by a user, it is specifically used to: convert the problem description into at least one task; and generate the first representation vector based on keywords contained in the at least one task.

[0019] In one possible implementation, when the vector representation module generates a first representation vector of a problem description based on a problem description input by a user, it is specifically used to: obtain the user's interest representation vector, where the interest representation vector is used to represent the user's preference for the data pipeline; and generate a first representation vector based on the interest representation vector and the problem description.

[0020] In one possible implementation, the device also includes: a job template generation module, which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the problem description obtained by sampling, and by changing the prompt words; and store the new job template in the template library after verifying that the new job template is legal.

[0021] In one possible implementation, the device also includes: a pipeline template generation module, which is used to: receive a seed pipeline template imported by a user; generate a new pipeline template based on the seed pipeline template and the job template sampled from the template library, and by changing the prompt words; and store the new pipeline template in the template library after verifying that the new pipeline template is legal.

[0022] In a third aspect, the present application provides a computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0024] In a fifth aspect, the present application provides a computer program product comprising instructions that, when executed by a computing device cluster, cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0025] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The following is a brief introduction to the drawings required for describing the embodiments or prior art.

[0027] FIG1 is a schematic diagram of the architecture of a data pipeline generation system provided in an embodiment of the present application;

[0028] FIG2 is a schematic diagram of an interaction between a tenant and a cloud computing platform provided in an embodiment of the present application;

[0029] FIG3 is a flow chart of a data pipeline generation method based on cloud computing technology provided in an embodiment of the present application;

[0030] FIG4 is a schematic diagram of steps for generating a data pipeline that matches a problem description input by a user based on a retrieved pipeline template provided by an embodiment of the present application;

[0031] FIG5 is a schematic diagram of a process for progressively generating a job template and a pipeline template according to an embodiment of the present application;

[0032] FIG6 is a schematic diagram of the structure of a data pipeline generation device based on cloud computing technology provided in an embodiment of the present application;

[0033] FIG7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0034] FIG8 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0035] FIG9 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0037] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.

[0038] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0039] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.

[0040] Typically, components required for data integration, data cleaning, data processing, and data analysis in a data pipeline are laid out on a front-end graphical user interface. Within the GUI, users can drag and drop component nodes on the pipeline canvas, configure component parameters, and link component nodes based on dependencies to ultimately generate data pipeline jobs. While this approach can create a data pipeline, it requires users to select component nodes, program them, configure node parameters, and construct relationships. This leads to a high learning threshold, cumbersome steps, and high cost of use.

[0041] In view of this, an embodiment of the present application provides a data pipeline generation method based on cloud computing technology, which can recommend the most relevant and useful data pipelines to users based on the problem description input by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., thereby improving the generation efficiency of the data pipeline and reducing the difficulty of generating the data pipeline. Exemplarily, a data pipeline can refer to a series of data processing steps connected in a specific order, which can be used to convert raw data into useful information or derived data. At least one processing node can be included in the data pipeline. Among them, a processing node refers to a component or module that performs a specific task or operation. These processing nodes can be responsible for processing the input data, such as: cleaning, conversion, analysis, aggregation and other operations. Each processing node can perform a certain function to transfer data from one state to the next.

[0042] For example, Figure 1 shows a schematic diagram of the architecture of a data pipeline generation system according to an embodiment of the present application. As shown in Figure 1, the data pipeline generation system 100 may include: a task decomposition component 110, a task keyword representation component 120, an interest representation component 130, a user representation component 140, a template retrieval component 150, a template library 160, a template screening component 170, a pipeline orchestration component 180, and a parsing component 190.

[0043] The task decomposition component 110 is mainly used to obtain the problem description input by the user in natural language, such as: the order of data flow construction, etc., and to convert the user's problem description into at least one task. Among them, a task can be at least one processing node on the data pipeline. Exemplarily, the problem description can be used to describe the content of the data pipeline that the user expects to use. For example, the problem description can be: build a data pipeline in the following order: 1. First execute the two data entry tasks sdi_nps_question_result_detai and sdi_csbi_product_sale_catalog, 2. Then execute the monthly NSS value statistical data preparation task for each cloud service, 3. Finally, execute the two data instruction monitoring tasks of NSS value uniqueness verification and NSS value maximum value verification. For this problem description, it can be, but is not limited to, converted into: a data migration task, a data statistics task, and an NSS value verification task.

[0044] The task keyword characterization component 120 is mainly used to extract keywords contained in each task and characterize the extracted keywords to obtain a question characterization vector.

[0045] The interest representation component 130 is mainly used to vectorize information related to the user's interests, such as the data pipeline created by the user, the historical pipeline used by the user, and the script written by the user, to obtain the user's interest representation vector. The interest representation vector can be used to represent the user's preference for the data pipeline. In some embodiments, when no information related to the user's interests can be collected, the interest representation component 130 can, but is not limited to, randomly generate an interest representation vector, or select the most popular interest representation vector as the current user's interest representation vector, or set the interest representation vector to empty, and so on.

[0046] The user characterization component 140 is mainly used to characterize the question characterization vector and the interest characterization vector to obtain a user characterization vector. For example, when the interest characterization vector is not used, the question characterization vector can be understood as a user characterization vector.

[0047] The template retrieval component 150 is primarily used to retrieve job templates and pipeline templates from the template library 160, which contains job templates and pipeline templates, using user representation vectors based on, for example, the cosine similarity algorithm. A job template describes a job, such as the problem to be solved, while a pipeline template describes a data pipeline. For example, a job can be an independent task or step in a data pipeline, which can be understood as an execution unit in the data pipeline. A pipeline can consist of at least one job. In this embodiment, the template retrieval component 150 may include a job template retrieval unit 151 and a pipeline template retrieval unit 152. The job template retrieval unit 151 may be used to retrieve job templates from the template library 160 using user representation vectors based on, for example, the cosine similarity algorithm. The pipeline template retrieval unit 152 may be used to retrieve pipeline templates from the template library 160 using user representation vectors based on, for example, the cosine similarity algorithm. In some embodiments, the job template retrieval unit 151 may retrieve at least one job template from the template library 160. The pipeline template retrieval unit 152 can retrieve at least one pipeline template from the template library 160. In some embodiments, the template retrieval component 150 can first retrieve pipeline templates from the template library 160. When the similarity between the representation vector of the retrieved pipeline template and the user representation vector is greater than or equal to a certain similarity threshold, the pipeline template associated with the highest calculated similarity can be selected as an available template and transmitted to the parsing component 190, ending the template retrieval process. When the similarity between the representation vector of the retrieved pipeline template and the user representation vector is less than a certain similarity threshold, the top n pipeline templates associated with the calculated similarities can be selected. Furthermore, based on the similarity between the user representation vector and the representation vector of the job templates in the template library 160, the top m job templates associated with the calculated similarities can be selected. The template retrieval unit 150 can then transmit the retrieved top n pipeline templates and top m job templates to the template screening component 170.

[0048] Template library 160 is primarily used to store pre-prepared job templates and pipeline templates. This library uses a vector database to store the representation vectors of each job template and each pipeline template. Template library 160 includes a job template library 161 and a pipeline template library 162. Job template library 161 stores pre-prepared job templates. Pipeline template library 162 stores pre-prepared pipeline templates.

[0049] The template screening component 170 is mainly used to screen the topn pipeline templates and topm job templates retrieved by the template retrieval component 150 based on preconfigured business rules (such as those that must meet specific fields, etc.) to obtain job templates and pipeline templates that meet the business rules. The template screening component 170 may include: a job template screening unit 171 and a pipeline template screening unit 172. The job template screening unit 171 can be used to screen the job templates retrieved by the template retrieval component 150 based on business rules. The pipeline template screening unit 172 can be used to screen the pipeline templates retrieved by the template retrieval component 150 based on business rules.

[0050] The pipeline orchestration component 180 is mainly used to orchestrate the job templates and pipeline templates screened by the template screening component 170 to generate a data pipeline that matches the user's problem description. Exemplarily, the data pipeline generated by the pipeline orchestration component 180 can be in JSON format.

[0051] The parsing component 190 is mainly used to convert the data pipeline received from the template retrieval component 150 or the data pipeline generated by the pipeline orchestration component 180 into a pipeline description file in a readable format, and to call the pipeline creation API to generate the final data pipeline and output the data pipeline, such as presenting it to the user through a graphical user interface.

[0052] It is understandable that each component or unit in the data pipeline generation system 100 can be, but is not limited to, a large language model (LLM), a convolutional neural network (CNN), or a deep neural network (DNN), or other neural networks, and can also be one or more network layers in LLM, CNN or DNN.

[0053] The above is an introduction to the data pipeline generation system 100 provided in the embodiment of the present application. Among them, the above-mentioned data pipeline generation system 100 can be configured on a cloud computing platform. For example, it is deployed on at least one virtual machine or container instance, so that the cloud computing platform can provide data pipeline generation services. Of course, the data pipeline generation system 100 can also be configured on a node other than the cloud computing platform. For example, it can be deployed in at least one data center, or deployed on at least one server. The specific details can be determined according to the actual situation and are not limited here. Among them, the cloud computing platform can provide pages related to public cloud services for tenants to remotely access public cloud services. In this embodiment, tenants (also referred to as "users") can purchase the data pipeline generation service that can be provided by the data pipeline generation system 100 in advance on the cloud computing platform. For ease of understanding, the interaction between tenants and the cloud computing platform is described below. As shown in Figure 2, the interaction between a tenant and the cloud computing platform primarily involves logging into the cloud computing platform 200 through a client webpage, selecting and purchasing a cloud service related to the data pipeline generation system 100 (i.e., the data pipeline generation service) within the cloud computing platform 200. After purchase, the tenant can then generate a data pipeline on the cloud computing platform 200 based on the functionality provided by the data pipeline generation service. The cloud computing platform 200 primarily manages the infrastructure for running the data pipeline generation service. For example, the infrastructure for running the data pipeline generation service may include multiple data centers located in different regions, each of which includes multiple servers. Data centers may provide basic resources for the data pipeline generation service, such as computing resources and storage resources. Therefore, when purchasing and using the data pipeline generation service, the tenant primarily pays for the resources used. When using the data pipeline generation service, the tenant can enter a problem description through the configuration interface, application program interface (API), or other tenant interaction interface provided by the cloud computing platform 200. The cloud computing platform 200 then generates a data pipeline that matches the problem description.

[0054] The above is an introduction to the data pipeline generation system 100 provided in the embodiment of the present application, as well as the interaction between tenants and the cloud computing platform. Based on the above content, the following introduces a data pipeline generation method based on cloud computing technology provided in the embodiment of the present application.

[0055] For example, FIG3 shows a flow chart of a data pipeline generation method based on cloud computing technology provided in an embodiment of the present application. This method can be applied to the cloud computing platform 200 described in FIG2 . As shown in FIG3 , the data pipeline generation method based on cloud computing technology can include the following steps:

[0056] S301 : The cloud computing platform 200 generates a first representation vector of the problem description based on the problem description input by the user.

[0057] In this embodiment, after the user enters a problem description, the cloud computing platform 200 can convert the problem description into tasks, extract keywords from each task, and perform vector representation on the extracted keywords to generate a first representation vector for the problem description. The first representation vector in this case can be understood as the user representation vector obtained from the problem representation vector. Of course, the cloud computing platform 200 can also directly perform vector representation on the problem description and use the representation result as the first representation vector, or first convert the problem description into tasks, then perform vector representation on the tasks, and use the representation result as the first representation vector.

[0058] In addition, with the user's authorization and legal permission, the cloud computing platform 200 can also collect the user's historical usage data, such as: data pipelines that have been used, data pipelines or scripts that have been built by themselves, and other data. Then, the cloud computing platform 200 can perform vector representation on the historical usage data it has collected to obtain the user's interest representation vector, which reflects the user's preference for the data pipeline. Finally, the cloud computing platform 200 can generate a first representation vector based on the interest representation vector and the problem description input by the user. The first representation vector at this time can be understood as the user representation vector obtained from the problem representation vector and the interest representation vector. Exemplarily, when the first representation vector is generated by keywords in the task converted from the problem description, the cloud computing platform 200 can combine the interest representation vector and the keywords in the task for vector representation to generate a first representation vector.

[0059] S302: The cloud computing platform 200 retrieves at least one pipeline template from the template library based on the similarity between the first characterization vector and the characterization vectors of the pipeline templates in the template library.

[0060] In this embodiment, after obtaining the first characterization vector, the cloud computing platform 200 can calculate the similarity between the first characterization vector and the characterization vectors of each pipeline template stored in the template library by using a cosine similarity algorithm, etc. Then, based on the calculated similarity, at least one pipeline template is retrieved from the template library. For example, the cloud computing platform 200 can sort the calculated similarities from large to small and use the pipeline templates associated with the top 10 or other number of similarities as the required pipeline templates. Exemplarily, the characterization vector of the pipeline template can be used to describe the content of the pipeline template. Exemplarily, the pipeline template can be understood as a template of a data pipeline.

[0061] S303: The cloud computing platform 200 generates a data pipeline that matches the problem description input by the user based on the retrieved pipeline template.

[0062] In this embodiment, after retrieving the pipeline template, the cloud computing platform 200 can generate a data pipeline that matches the problem description entered by the user based on the retrieved pipeline template. For example, when multiple pipeline templates are retrieved, the cloud computing platform 200 can randomly select one of them as the data pipeline that matches the problem description entered by the user.

[0063] As a possible implementation, as shown in FIG4 , the cloud computing platform 200 generates a data pipeline that matches the problem description input by the user based on the retrieved pipeline template, which may include the following steps:

[0064] S401, the cloud computing platform 200 determines whether the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold. Among them, the cloud computing platform 200 can sort the similarities between the representation vector of the retrieved pipeline template and the first representation vector from large to small (or small to large), and select the largest similarity therefrom, that is, select the maximum value among these similarities. Then, the cloud computing platform 200 can compare the maximum value with the target threshold. When the maximum value is greater than or equal to the template threshold, it indicates that the pipeline template associated with the maximum value can be directly used to process the problem description input by the user, and therefore, S402 can be executed. When the maximum value is less than the template threshold, it indicates that the pipeline template associated with the maximum value cannot be directly used to process the problem description input by the user, and therefore, S403 can be executed.

[0065] S402: The cloud computing platform 200 generates a data pipeline that matches the problem description entered by the user based on the pipeline template associated with the maximum value. The cloud computing platform 200 may convert the pipeline template associated with the maximum value into a readable pipeline description file and call a pipeline creation API to generate the final data pipeline. In this way, a data pipeline that matches the problem description entered by the user is obtained.

[0066] S403. The cloud computing platform 200 retrieves at least one job template from the template library based on the similarity between the first characterization vector and the characterization vectors of the job templates in the template library. The cloud computing platform 200 may calculate the similarity between the first characterization vector and the characterization vectors of the job templates stored in the template library using a cosine similarity algorithm, etc. Then, based on the calculated similarity, at least one job template is retrieved from the template library. For example, the cloud computing platform 200 may sort the calculated similarities from large to small, and use the job templates associated with the top 10 or other number of similarities as the required job templates. Exemplarily, a pipeline template may include at least one job template. A job template may be an execution unit in a pipeline template. Exemplarily, the characterization vector of a job template may be used to describe the content of the job template.

[0067] S404: The cloud computing platform 200 generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates. The cloud computing platform 200 may rearrange the retrieved pipeline templates and job templates using LLM, etc., to obtain a data pipeline. Furthermore, the cloud computing platform 200 may convert the rearranged data pipeline into a readable pipeline description file and call a pipeline creation API to generate the final data pipeline, thereby obtaining a data pipeline that matches the problem description entered by the user. In some embodiments, given that the retrieved pipeline templates and job templates may differ from the user-entered problem description in terms of application domain, if all pipeline templates and job templates are directly rearranged, the resulting data pipeline may not match the user-entered problem description. Therefore, to avoid this situation, the cloud computing platform 200 may first select at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates based on user-configured or pre-set business rules. Then, based on the filtered pipeline templates and job templates, a data pipeline that matches the problem description entered by the user is generated.

[0068] In this way, when a user has a data pipeline requirement, after the user enters a problem description, the most relevant and useful data pipeline will be recommended to the user based on the problem description entered by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., which improves the generation efficiency of the data pipeline and reduces the difficulty of generating the data pipeline.

[0069] In some embodiments, considering that the number of job templates and pipeline templates carried in the template library is insufficient, the quality of the subsequently generated data pipeline may be poor. Therefore, in this embodiment, when preparing job templates and pipeline templates, a high-quality expert template can be used as a seed to progressively generate the job template first and then the pipeline template, thereby improving the generation efficiency of job templates and pipeline templates and reducing the labor cost of experts writing templates. Specifically, as shown in Figure 5, the following steps can be included:

[0070] S501: The cloud computing platform 200 receives a seed job template imported by a user from a seed job template library and saves the seed job template in the job template library. In this embodiment, a user can import a seed job template from the seed job template library into the job template library on the cloud computing platform 200 at the front end of the cloud computing platform 200. The seed job template can be, but is not limited to, a job template written by a domain expert. After successful import, the user can trigger a physical or virtual button on the front end of the cloud computing platform 200 to indicate the start of training, causing the cloud computing platform 200 to initiate job template generation, i.e., execute the subsequent job template generation process.

[0071] S502: The cloud computing platform 200 randomly samples questions from a data source and randomly samples job templates from a job template library, and transmits the sampled questions and job templates to a job generator.

[0072] S503: The cloud computing platform 200 processes the sampled questions and the job template through a job generator to generate a new job template. The job generator may change the prompt during the processing.

[0073] S504: The cloud computing platform 200 verifies the new job template, such as whether the new job template can be run, and feeds back the verification results to the job generator so that the job generator can change the prompt according to the feedback results, for example, changing the constraints. In addition, the cloud computing platform 200 can save the verified job templates to the job template library. In addition, after the cloud computing platform 200 completes the verification, manual verification methods can be introduced, such as manually determining whether the dependencies between the various components in the job template are appropriate, and feeding back the results of the manual verification to the job generator so that the job generator can change the prompt according to the manual feedback results, thereby further improving the generation quality of the job template. Repeating S502 to S504 can accumulate a large number of job templates. After the accumulated job templates meet the requirements, the pipeline template generation process can be executed.

[0074] S505: The cloud computing platform 200 receives the seed pipeline template imported by the user from the seed pipeline template library and stores the seed pipeline template in the pipeline template library. In this embodiment, the user can import the seed pipeline template in the seed pipeline template library into the pipeline template library on the cloud computing platform 200 at the front end of the cloud computing platform 200. The seed pipeline template can be, but is not limited to, a pipeline template written by a domain expert. After the import is successful, the user can trigger a physical key or virtual key at the front end of the cloud computing platform 200 to indicate the start of training, so that the cloud computing platform 200 starts generating the pipeline template, i.e., executing the subsequent pipeline template generation process. In some embodiments, the user can also import the seed operation template and the seed pipeline template into the cloud computing platform 200 at the same time. In this case, after the accumulated operation templates meet the requirements, the subsequent pipeline generation process can be automatically executed, or it can be executed after the user's permission. The specific method can be determined according to the actual situation and is not limited here.

[0075] S506 , the cloud computing platform 200 randomly samples a job template from the job template library, and randomly samples a pipeline template from the pipeline template library, and transmits the sampled job template and pipeline template to the pipeline generator.

[0076] S507: The cloud computing platform 200 processes the sampled job template and pipeline template through the pipeline generator to generate a new pipeline template. The pipeline generator may change the prompt during the processing.

[0077] S508, the cloud computing platform 200 verifies the new pipeline template, such as whether the new pipeline template can be run, and feeds back the verification results to the pipeline generator so that the pipeline generator can change the prompt according to the feedback results, for example, change the constraint conditions, etc. In addition, the cloud computing platform 200 can save the pipeline template that has passed the verification to the pipeline template library. In addition, after the cloud computing platform 200 completes the verification, it can also introduce a manual verification method, such as: manually judging whether the pipeline template meets expectations, etc., and feeding back the results of the manual inspection to the pipeline generator, so that the pipeline generator can change the prompt according to the manual feedback results, thereby further improving the generation quality of the pipeline template. Repeating S506 to S508 can accumulate a large number of pipeline templates.

[0078] In this embodiment, after the accumulated job templates meet expectations, the LLM-based job description generator processes each job template in the job template library to generate a text description for each job template. The NN-based job characterizer then performs vector representation on the text description of each job template to obtain a representation vector for each job template. The representation vectors of each job template can then be indexed using a vector database, allowing the corresponding job template to be found by searching the representation vector.

[0079] Furthermore, once the accumulated pipeline templates meet expectations, the LLM-based pipeline description generator can be used to process each pipeline template in the pipeline template library to generate a text description for each pipeline template. The NN-based pipeline characterizer then performs vector representation on each pipeline template's text description to obtain a representation vector for each pipeline template. The representation vectors for each pipeline template can then be indexed using a vector database, allowing the corresponding pipeline template to be found by searching the representation vector.

[0080] It should be understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In addition, the various embodiments described above can be combined according to actual circumstances, and the combined solutions are still within the scope of protection of this application.

[0081] The above is an introduction to the data pipeline generation method based on cloud computing technology provided by an embodiment of the present application. Next, based on the method in the above embodiment, a data pipeline generation device based on cloud computing technology provided by an embodiment of the present application is introduced.

[0082] For example, FIG6 shows a schematic diagram of the structure of a data pipeline generation device based on cloud computing technology provided in an embodiment of the present application. For example, the data pipeline generation device based on cloud computing technology can be, but is not limited to, deployed on a cloud computing platform. The cloud computing platform runs on an infrastructure comprising multiple data centers located in different regions, each data center comprising multiple servers. As shown in FIG6 , the data pipeline generation device 600 based on cloud computing technology comprises: a vector representation template 601, a template retrieval module 602, and a pipeline generation module 603. The vector representation module 601 is configured to generate a first representation vector of a problem description input by a user, wherein the problem description describes the content of the data pipeline the user desires to use. The template retrieval module 602 is configured to retrieve at least one pipeline template from a template library based on the similarity between the first representation vector and the representation vectors of pipeline templates in the template library. The pipeline generation module 603 is configured to generate a data pipeline that matches the problem description based on the retrieved pipeline template.

[0083] In some embodiments, when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module 603 is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value.

[0084] If the maximum similarity between the representation vector of the retrieved pipeline template and the first representation vector is less than a target threshold, the template retrieval module is used to retrieve at least one job template from the template library based on the similarity between the first representation vector and the representation vectors of the job templates in the template library. The pipeline generation module 603 is used to generate a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates. A job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template.

[0085] In some embodiments, when the pipeline generation module 603 generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, it is specifically used to: based on business rules, filter out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates respectively; based on the filtered pipeline templates and job templates, generate a data pipeline that matches the problem description.

[0086] In some embodiments, when generating a first representation vector for a problem description based on a user-input problem description, the vector representation module 601 is specifically configured to: convert the problem description into at least one task; and generate the first representation vector based on keywords included in the at least one task. A task is at least one processing node in a data pipeline.

[0087] In some embodiments, when the vector representation module 601 generates a first representation vector of a problem description based on a problem description input by a user, it is specifically used to: obtain the user's interest representation vector, where the interest representation vector is used to represent the user's preference for the data pipeline; and generate a first representation vector based on the interest representation vector and the problem description.

[0088] In some embodiments, the device also includes: a job template generation module (not shown in the figure), which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the problem description obtained by sampling, and by changing the prompt words; and store the new job template in the template library after verifying that the new job template is legal.

[0089] In some embodiments, the device also includes: a pipeline template generation module (not shown in the figure), which is used to: receive a seed pipeline template imported by a user; generate a new pipeline template based on the seed pipeline template and the job template sampled from the template library, and by changing the prompt words; and store the new pipeline template in the template library after verifying that the new pipeline template is legal.

[0090] In some embodiments, the vector representation template 601, template retrieval module 602, and pipeline generation module 603 shown in FIG6 can all be implemented via software or hardware. For example, the implementation of vector representation template 601 will be described below using vector representation template 601 as an example. Similarly, the implementation of template retrieval module 602 and pipeline generation module 603 can refer to the implementation of vector representation template 601.

[0091] As an example of a software functional unit, the vector representation template 601 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the vector representation template 601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0092] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0093] As an example of a hardware functional unit, a module may include at least one computing device, such as a server. Alternatively, the vector representation template 601 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0094] The multiple computing devices included in vector representation template 601 can be distributed in the same region or in different regions. The multiple computing devices included in vector representation template 601 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in vector representation template 601 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0095] It should be noted that, in other embodiments, the vector representation template 601 can be used to execute any step in the data pipeline generation method based on cloud computing technology described in the above embodiment, the template retrieval module 602 can also be used to execute any step in the data pipeline generation method based on cloud computing technology described in the above embodiment, and the pipeline generation module 603 can also be used to execute any step in the data pipeline generation method based on cloud computing technology described in the above embodiment. In addition, any two of the vector representation template 601, the template retrieval module 602, and the pipeline generation module 603 can also be combined together to be responsible for executing any step in the data pipeline generation method based on cloud computing technology described in the above embodiment. In addition, the steps that the vector representation template 601, the template retrieval module 602, and the pipeline generation module 603 are responsible for implementing can also be specified as needed. By respectively implementing different steps in the data pipeline generation method based on cloud computing technology described in the above embodiment, the full functions of the data pipeline generation device 600 based on cloud computing technology shown in Figure 6 can be realized.

[0096] This application also provides a computing device 700. As shown in Figure 7, computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. Processor 704, memory 706, and communication interface 708 communicate with each other via bus 702. Computing device 700 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 700.

[0097] Bus 702 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 704 may include a path for transmitting information between various components of computing device 700 (e.g., memory 706, processor 704, and communication interface 708).

[0098] The processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0099] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0100] Memory 706 stores executable program code, which processor 704 executes to implement the functions of vector representation template 601, template retrieval module 602, and pipeline generation module 603 shown in FIG6 , thereby implementing the cloud computing-based data pipeline generation method described in the above embodiment. Specifically, memory 706 stores instructions for executing the cloud computing-based data pipeline generation method described in the above embodiment.

[0101] Alternatively, the memory 706 stores executable code, and the processor 704 executes the executable code to implement the functions of the data pipeline generation device 600 based on cloud computing technology shown in Figure 6, thereby implementing the data pipeline generation method based on cloud computing technology described in the above embodiment. In other words, the memory 706 stores instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.

[0102] The communication interface 703 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.

[0103] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0104] As shown in Figure 8, the computing device cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.

[0105] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment. In other words, the combination of one or more computing devices 700 can jointly execute instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.

[0106] It should be noted that the memory 706 in different computing devices 700 in the computing device cluster can store different instructions, each used to execute part of the functions of the cloud computing-based data pipeline generation apparatus 600 shown in FIG6 . In other words, the instructions stored in the memory 706 in different computing devices 700 can implement the functions of one or more of the vector representation template 601, the template retrieval module 602, and the pipeline generation module 603.

[0107] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG9 shows a possible implementation. As shown in FIG9 , two computing devices 700A and 700B are connected via a network. Specifically, the connection to the network is made via a communication interface in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the vector representation template 701. At the same time, the memory 706 in the computing device 700B stores instructions for executing the functions of the template retrieval module 602 and the pipeline generation module 603.

[0108] It should be understood that the functionality of the computing device 700A shown in FIG9 may also be implemented by multiple computing devices 700. Similarly, the functionality of the computing device 700B may also be implemented by multiple computing devices 700.

[0109] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection relationship between computing device clusters described in Figures 8 and 9 . However, the memory 706 in one or more computing devices 700 in this computing device cluster can store the same instructions for executing the methods described in the above embodiments.

[0110] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the aforementioned data pipeline generation method based on cloud computing technology. In other words, the combination of one or more computing devices 700 can jointly execute instructions for executing the aforementioned data pipeline generation method based on cloud computing technology.

[0111] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method in the above embodiment; or, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the above embodiment. Exemplarily, the computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0112] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product containing instructions, which, when executed by a computing device, enables the computing device to execute the method in the above embodiment, or, when executed by a computing device cluster, enables the computing device cluster to execute the method in the above embodiment.

[0113] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0114] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).

[0116] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data pipeline generation method based on cloud computing technology, characterized in that: Applied to a cloud computing platform, the cloud computing platform runs on an infrastructure, the infrastructure includes a plurality of data centers arranged in different regions, each data center includes a plurality of servers; The method comprises: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use; The cloud computing platform retrieves at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library; The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template.

2. The method according to claim 1, characterized in that The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template, including: When the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the cloud computing platform generates a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; When the maximum value is less than the target threshold, the cloud computing platform retrieves at least one job template from the template library based on the similarity between the first characterization vector and the characterization vector of the job template in the template library, and generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, wherein one of the job templates is an execution unit in one of the pipeline templates, and one of the pipeline templates includes at least one of the job templates.

3. The method according to claim 2, characterized in that The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, including: The cloud computing platform selects at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates based on the business rules; The cloud computing platform generates a data pipeline that matches the problem description based on the screened pipeline templates and job templates.

4. The method according to any one of claims 1 to 3, characterized in that: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, including: The cloud computing platform converts the problem description into at least one task, wherein one of the tasks is at least one processing node on the data pipeline; The cloud computing platform generates the first representation vector based on keywords included in the at least one task.

5. The method according to any one of claims 1 to 4, characterized in that: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, including: The cloud computing platform obtains an interest representation vector of the user, where the interest representation vector is used to represent the user's preference for the data pipeline; The cloud computing platform generates the first representation vector based on the interest representation vector and the problem description.

6. The method according to any one of claims 1 to 5, characterized in that: Before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: The cloud computing platform receives a seed job template imported by a user; The cloud computing platform generates a new job template based on the seed job template and the sampled problem description and by changing the prompt word; The cloud computing platform stores the new job template in the template library after verifying that the new job template is legal.

7. The method according to any one of claims 1 to 6, characterized in that: Before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: The cloud computing platform receives a seed pipeline template imported by a user; The cloud computing platform generates a new pipeline template based on the seed pipeline template and the job template sampled from the template library and by changing the prompt words; The cloud computing platform stores the new pipeline template in the template library after verifying that the new pipeline template is legal.

8. A data pipeline generation device based on cloud computing technology, characterized in that: Deployed on a cloud computing platform, the cloud computing platform runs on an infrastructure, the infrastructure includes multiple data centers located in different regions, each data center includes multiple servers; The device comprises: A vector characterization module, used to generate a first characterization vector for characterizing the problem description based on a problem description input by a user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use; A template retrieval module, configured to retrieve at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library; The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template.

9. The device according to claim 8, characterized in that In the case where the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; In the case where the maximum value is less than the target threshold, the template retrieval template is used to retrieve at least one job template from the template library based on the similarity between the first characterization vector and the characterization vectors of the job templates in the template library; The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, wherein one job template is an execution unit in one pipeline template, and one pipeline template includes at least one job template.

10. The device according to claim 9, characterized in that When the pipeline generation module generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, it is specifically used to: Based on the business rules, at least one assembly line template and at least one job template are respectively selected from the retrieved assembly line templates and job templates; Based on the screened pipeline templates and job templates, a data pipeline that matches the problem description is generated.

11. The device according to any one of claims 8 to 10, characterized in that: When the vector representation module generates a first representation vector for representing the problem description based on the problem description input by the user, it is specifically used to: Convert the problem description into at least one task; The first representation vector is generated based on keywords included in the at least one task.

12. The device according to claim 11, characterized in that When the vector representation module generates a first representation vector for representing the problem description based on the problem description input by the user, it is specifically used to: Obtaining an interest representation vector of the user, where the interest representation vector is used to represent the user's preference for the data pipeline; The first representation vector is generated based on the interest representation vector and the problem description.

13. The device according to any one of claims 8 to 12, characterized in that: The device also includes: a job template generation module, which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the sampled problem description and by changing the prompt word; When the new job template is verified to be legal, the new job template is stored in the template library.

14. The device according to any one of claims 8 to 13, characterized in that: The device also includes: a pipeline template generation module, which is used to: Receive the seed pipeline template imported by the user; Based on the seed pipeline template and the job template sampled from the template library, and by changing the prompt words, a new pipeline template is generated; When the new pipeline template is verified to be legal, the new pipeline template is stored in the template library.

15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1-7.

16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7, wherein the computing device cluster comprises at least one computing device.

17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7, wherein the computing device cluster includes at least one computing device.

Citation Information

Patent Citations

  • Feedback information processing method and device, equipment and storage medium

    CN113722577A

  • Cloud manufacturing service process recommendation method based on deep learning

    CN115048490A

  • Continuous integration assembly line generation method and device, server, medium and product

    CN115617381A

  • System and method for automatic generation of BI models using data introspection and curation

    WO2021168331A1