Data pipeline generation method and device based on cloud computing technology
Through the data pipeline generation method based on cloud computing technology, the problem description entered by the user is automatically retrieved and generated data pipelines, which solves the cumbersome problems in the creation process in the existing technology, improves the generation efficiency and reduces the difficulty.
Patent Information
- Application Number
- CN202410012950.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-01-04
- Publication Date
- 2025-05-16
AI Technical Summary
The creation process of existing data pipeline products is cumbersome, and the learning cost and usage threshold are high, making it difficult for users to quickly generate the required data pipelines.
The data pipeline generation method based on cloud computing technology is adopted to generate a representation vector through the problem description input by the user, search the pipeline templates in the template library, and automatically generate a data pipeline that matches the problem description.
It simplifies the data pipeline creation process, improves generation efficiency, reduces generation difficulty, and allows users to avoid entering detailed operation instructions or clicking actions.
Smart Images

Figure CN120011430A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 16, 2023, with application number 202311536346.9 and application name “A Data Pipeline Creation Method”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of information technology (IT) technology, and in particular to a data pipeline generation method and device based on cloud computing technology. Background Art
[0003] Data pipeline is an automated process of extracting, transforming, and loading data from the original data source to the target data warehouse or data lake. It is a key tool used by data engineers and data scientists to process and manage large amounts of data. The automation of data pipeline can improve the efficiency and accuracy of data processing and reduce manual errors and duplication of work. However, in traditional data pipeline products, users need to follow cumbersome steps to create data pipelines, and the learning cost and usage threshold are relatively high. Therefore, how to simplify the process of creating data pipelines is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides a data pipeline generation method, device, computing device cluster, computer storage medium and computer product based on cloud computing technology, which can simplify the process of creating a data pipeline.
[0005] In a first aspect, the present application provides a data pipeline generation method based on cloud computing technology, which is applied to a cloud computing platform, and the cloud computing platform runs on an infrastructure, and the infrastructure includes multiple data centers set up in different regions, and each data center includes multiple servers. The method includes: the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use; the cloud computing platform retrieves at least one pipeline template from the template library based on the similarity between the first characterization vector and the characterization vector of the pipeline template in the template library; the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template.
[0006] In this way, when a user has a data pipeline demand, after the user enters a problem description, the most relevant and useful data pipeline is recommended to the user based on the problem description entered by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., which improves the generation efficiency of the data pipeline and reduces the difficulty of generating the data pipeline.
[0007] In a possible implementation, the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template, including: when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the cloud computing platform generates a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; when the maximum value is less than the target threshold, the cloud computing platform retrieves at least one job template from the template library based on the similarity between the first representation vector and the representation vector of the job template in the template library, and generates a data pipeline that matches the problem description based on the retrieved pipeline template and the job template. Wherein, a job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template. In this way, the most relevant and useful data pipeline can be recommended to users in different situations.
[0008] In a possible implementation, the cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, including: the cloud computing platform selects at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates based on business rules; the cloud computing platform generates a data pipeline that matches the problem description based on the selected pipeline templates and job templates. In this way, it can be ensured that the generated data pipeline matches the business corresponding to the problem description, so as to recommend the most relevant and useful data pipeline to the user.
[0009] In a possible implementation, the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, including: the cloud computing platform converts the problem description into at least one task; the cloud computing platform generates the first characterization vector based on the keywords included in the at least one task. In this way, the problem description can be converted into tasks, keywords can be extracted from the tasks, and finally the first characterization vector for characterizing the problem description can be generated from the keywords.
[0010] In a possible implementation, the cloud computing platform generates a first representation vector for representing the problem description based on the problem description input by the user, including: the cloud computing platform obtains the user's interest representation vector, the interest representation vector is used to represent the user's preference for the data pipeline; the cloud computing platform generates the first representation vector based on the interest representation vector and the problem description. In this way, the template retrieved subsequently can be consistent with the user's preference, and the desired data pipeline can be recommended to the user.
[0011] In a possible implementation, before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: the cloud computing platform receives a seed job template imported by the user; the cloud computing platform generates a new job template based on the seed job template and the sampled problem description and by changing the prompt words; the cloud computing platform stores the new job template in the template library after verifying that the new job template is legal. In this way, the job template can be automatically generated, which improves the efficiency of generating the job template.
[0012] In a possible implementation, before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: the cloud computing platform receives the seed pipeline template imported by the user; the cloud computing platform generates a new pipeline template based on the seed pipeline template and the job template sampled from the template library and by changing the prompt words; the cloud computing platform stores the new pipeline template in the template library after verifying that the new pipeline template is legal. In this way, the pipeline template can be automatically generated, the efficiency of generating the pipeline template is improved, and the problem of insufficient pipeline templates is solved.
[0013] In a second aspect, the present application provides a data pipeline generation device based on cloud computing technology, which is deployed on a cloud computing platform. The cloud computing platform runs on an infrastructure, and the infrastructure includes multiple data centers arranged in different areas, and each data center includes multiple servers. The device includes: a vector representation module, a template retrieval module and a pipeline generation module. Among them, the vector representation module is used to generate a first representation vector for representing the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use. The template retrieval module is used to retrieve at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library. The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template.
[0014] In one possible implementation, when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value.
[0015] When the maximum value is less than the target threshold, the template retrieval template is used to retrieve at least one job template from the template library based on the similarity between the first representation vector and the representation vector of the job template in the template library; the pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template and the job template. A job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template.
[0016] In one possible implementation, when the pipeline generation module generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, it is specifically used to: based on business rules, filter out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates respectively; and generate a data pipeline that matches the problem description based on the filtered pipeline templates and job templates.
[0017] In a possible implementation, when the vector representation module generates a first representation vector of a problem description based on a problem description input by a user, the vector representation module is specifically used to: convert the problem description into at least one task; and generate a first representation vector based on keywords included in the at least one task.
[0018] In one possible implementation, when the vector representation module generates a first representation vector of a problem description based on a problem description input by a user, it is specifically used to: obtain a user's interest representation vector, where the interest representation vector is used to represent the user's preference for the data pipeline; and generate a first representation vector based on the interest representation vector and the problem description.
[0019] In one possible implementation, the device also includes: a job template generation module, which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the problem description obtained by sampling, and by changing the prompt words; and store the new job template in the template library after verifying that the new job template is legal.
[0020] In one possible implementation, the device also includes: a pipeline template generation module, which is used to: receive a seed pipeline template imported by a user; generate a new pipeline template based on the seed pipeline template and a job template sampled from a template library, and by changing prompt words; and store the new pipeline template in the template library after verifying that the new pipeline template is legal.
[0021] In a third aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect. Exemplarily, the computing device cluster may include one or more computing devices.
[0023] In a fifth aspect, the present application provides a computer program product including instructions, which, when executed by a computing device cluster, enables the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect. Exemplarily, the computing device cluster may include one or more computing devices.
[0024] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of the architecture of a data pipeline generation system provided in an embodiment of the present application;
[0026] Figure 2 It is a schematic diagram of an interaction form between a tenant and a cloud computing platform provided in an embodiment of the present application;
[0027] Figure 3 It is a flow chart of a data pipeline generation method based on cloud computing technology provided in an embodiment of the present application;
[0028] Figure 4 It is a schematic diagram of steps for generating a data pipeline that matches a problem description input by a user based on a retrieved pipeline template provided by an embodiment of the present application;
[0029] Figure 5 It is a schematic diagram of a process of progressively generating a job template and a pipeline template provided in an embodiment of the present application;
[0030] Figure 6 It is a structural schematic diagram of a data pipeline generation device based on cloud computing technology provided in an embodiment of the present application;
[0031] Figure 7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0032] Figure 8 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0033] Fig. 9 It is a structural diagram of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0035] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first response message and a second response message are used to distinguish different response messages rather than to describe a specific order of the response messages.
[0036] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0037] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0038] Generally, the components such as data integration, data cleaning, data processing, and data analysis required in the data pipeline can be laid out on the front-end graphical user interface. On the graphical user interface, users can drag and drop component nodes on the pipeline canvas, configure component parameters, associate component nodes according to dependencies, and finally generate data pipeline jobs. Although this method can create a data pipeline, it requires users to select component nodes, write component nodes, configure node parameters, construct blood relationships, etc., resulting in a high learning threshold for users, cumbersome operation steps, and high cost of use.
[0039] In view of this, the embodiment of the present application provides a data pipeline generation method based on cloud computing technology, which can recommend the most relevant and useful data pipeline to the user based on the problem description input by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., thereby improving the generation efficiency of the data pipeline and reducing the difficulty of generating the data pipeline. Exemplarily, a data pipeline can refer to a series of data processing steps connected in a specific order, which can be used to convert raw data into useful information or derived data. At least one processing node may be included in the data pipeline. Among them, a processing node refers to a component or module that performs a specific task or operation. These processing nodes can be responsible for processing the input data, such as: cleaning, conversion, analysis, aggregation and other operations. Each processing node can perform a certain function to transfer data from one state to the next state.
[0040] For example, Figure 1 FIG. 1 is a schematic diagram showing the architecture of a data pipeline generation system according to an embodiment of the present application. Figure 1 As shown, the data pipeline generation system 100 may include: a task decomposition component 110, a task keyword representation component 120, an interest representation component 130, a user representation component 140, a template retrieval component 150, a template library 160, a template screening component 170, a pipeline orchestration component 180 and a parsing component 190.
[0041] The task decomposition component 110 is mainly used to obtain the problem description input by the user using natural language, such as: data flow construction order, etc., and convert the user's problem description into at least one task. Among them, a task can be at least one processing node on the data pipeline. Exemplarily, the problem description can be used to describe the content of the data pipeline that the user expects to use. For example, the problem description can be: build a data pipeline in the following order: 1. First execute the two data entry tasks sdi_nps_question_result_detai and sdi_csbi_product_sale_catalog, 2. Then execute the monthly NSS value statistical data preparation task for each cloud service, 3. Finally, execute the two data instruction monitoring tasks of NSS value uniqueness verification and NSS value maximum value verification. For this problem description, it can be, but is not limited to, converted into: data migration tasks, data statistics tasks, and NSS value verification tasks.
[0042] The task keyword characterization component 120 is mainly used to extract keywords contained in each task, and characterize the extracted keywords to obtain a problem characterization vector.
[0043] The interest representation component 130 is mainly used to vectorize information related to the user's interests, such as the data pipeline created by the user, the historical pipeline used by the user, and the script written by the user, so as to obtain the user's interest representation vector. The interest representation vector can be used to represent the user's preference for the data pipeline. In some embodiments, when no information related to the user's interests can be collected, the interest representation component 130 can, but is not limited to, randomly generate an interest representation vector, or select the most popular interest representation vector as the current user's interest representation vector, or set the interest representation vector to empty, and so on.
[0044] The user characterization component 140 is mainly used to characterize the question characterization vector and the interest characterization vector to obtain a user characterization vector. Exemplarily, when the interest characterization vector is not used, the question characterization vector can be understood as a user characterization vector.
[0045] The template retrieval component 150 is mainly used to retrieve job templates and pipeline templates from the template library 160 carrying job templates and pipeline templates based on the cosine similarity algorithm, etc., using the user characterization vector. Among them, the job template is a description of the job, such as what problem to solve, etc., and the pipeline template is a description of the data pipeline. Exemplarily, a job can be an independent task or step in the data pipeline, which can be understood as an execution unit in the data pipeline. A pipeline can be composed of at least one job. In this embodiment, the template retrieval component 150 may include: a job template retrieval unit 151 and a pipeline template retrieval unit 152. The job template retrieval unit 151 can be used to retrieve the job template from the template library 160 based on the cosine similarity algorithm, etc., using the user characterization vector. The pipeline template retrieval unit 152 can be used to retrieve the pipeline template from the template library 160 based on the cosine similarity algorithm, etc., using the user characterization vector. In some embodiments, the job template retrieval unit 151 can retrieve at least one job template from the template library 160. The pipeline template retrieval unit 152 can retrieve at least one pipeline template from the template library 160. In some embodiments, the template retrieval component 150 can first retrieve the pipeline template from the template library 160. When there is a similarity between the representation vector of the retrieved pipeline template and the user representation vector that is greater than or equal to a certain similarity threshold, the pipeline template associated with the highest similarity among the calculated similarities can be used as an available template, and transmitted to the parsing component 190, and the template retrieval process ends. When the similarity between the representation vector of the retrieved pipeline template and the user representation vector is less than a certain similarity threshold, the pipeline templates associated with the top n similarities can be selected from the calculated similarities; and, based on the similarity between the user representation vector and the representation vector of the job template in the template library 160, the job templates associated with the top m similarities can be selected from the calculated similarities. Then, the template retrieval unit 150 can transmit the retrieved topn pipeline templates and top m job templates to the template screening component 170.
[0046] The template library 160 is mainly used to carry pre-prepared job templates and pipeline templates. In the template library 160, the characterization vectors of each job template and the characterization vectors of each pipeline template can be carried through the vector database. The template library 160 includes: a job template library 161 and a pipeline template library 162. The job template library 161 can be used to carry pre-prepared job templates. The pipeline template library 162 can be used to carry pre-prepared pipeline templates.
[0047] The template screening component 170 is mainly used to screen the top n pipeline templates and top m job templates retrieved by the template retrieval component 150 based on preconfigured business rules (such as the need to meet specific fields, etc.) to obtain job templates and pipeline templates that meet the business rules. The template screening component 170 may include: a job template screening unit 171 and a pipeline template screening unit 172. The job template screening unit 171 can be used to screen the job templates retrieved by the template retrieval component 150 based on business rules. The pipeline template screening unit 172 can be used to screen the pipeline templates retrieved by the template retrieval component 150 based on business rules.
[0048] The pipeline arrangement component 180 is mainly used to arrange the job templates and pipeline templates screened by the template screening component 170 to generate a data pipeline that matches the user's problem description. Exemplarily, the data pipeline generated by the pipeline arrangement component 180 can be in JSON format.
[0049] The parsing component 190 is mainly used to convert the data pipeline received from the template retrieval component 150 or the data pipeline generated by the pipeline orchestration component 180 into a pipeline description file in a readable format, and to call the pipeline creation API to generate the final data pipeline and output the data pipeline, such as presenting it to the user through a graphical user interface.
[0050] It can be understood that each component or unit in the data pipeline generation system 100 can be, but is not limited to, a large language model (LLM), a convolutional neural network (CNN), or a deep neural network (DNN) and other neural networks, and can also be one or more network layers in LLM, CNN or DNN.
[0051] The above is an introduction to the data pipeline generation system 100 provided in the embodiment of the present application. Among them, the above-mentioned data pipeline generation system 100 can be configured on a cloud computing platform. For example, it is deployed on at least one virtual machine or container instance, so that the cloud computing platform can provide data pipeline generation services. Of course, the data pipeline generation system 100 can also be configured on nodes other than the cloud computing platform. For example, it can be deployed in at least one data center, or deployed on at least one server. The specific situation can be determined according to the actual situation and is not limited here. Among them, the cloud computing platform can provide pages related to public cloud services for tenants to remotely access public cloud services. In this embodiment, tenants (also referred to as "users") can purchase the data pipeline generation service that can be provided by the data pipeline generation system 100 on the cloud computing platform in advance. For ease of understanding, the interaction between tenants and the cloud computing platform is described below. As Figure 2 As shown, the interaction between the tenant and the cloud computing platform mainly includes: the tenant logs in to the cloud computing platform 200 through the client web page, selects and purchases the cloud service (i.e., the data pipeline generation service) related to the data pipeline generation system 100 in the cloud computing platform 200, and after the purchase, the tenant can generate a data pipeline on the cloud computing platform 200 based on the functions provided by the data pipeline generation service. Among them, the cloud computing platform 200 is mainly used to manage the infrastructure for running the data pipeline generation service. Exemplarily, the infrastructure for running the data pipeline generation service may include multiple data centers set up in different regions, and each data center includes multiple servers. The data center can provide basic resources for the data pipeline generation service, such as computing resources, storage resources, etc. Therefore, when the tenant purchases and uses the data pipeline generation service, it mainly pays for the resources used. When using the data pipeline generation service, the tenant can enter the problem description through the configuration interface, application program interface (API) or the interface for interacting with the tenant provided by the cloud computing platform 200, and the cloud computing platform 200 can generate a data pipeline that matches the problem description according to the problem description entered by the tenant.
[0052] The above is an introduction to the data pipeline generation system 100 provided in the embodiment of the present application, and the interaction between the tenant and the cloud computing platform. Based on the above content, a data pipeline generation method based on cloud computing technology provided in the embodiment of the present application is introduced below.
[0053] For example, Figure 3 The flowchart of a data pipeline generation method based on cloud computing technology provided by an embodiment of the present application is shown. Figure 2 The cloud computing platform 200 described in Figure 3As shown, the data pipeline generation method based on cloud computing technology may include the following steps:
[0054] S301. The cloud computing platform 200 generates a first representation vector of the problem description based on the problem description input by the user.
[0055] In this embodiment, after the user inputs the problem description, the cloud computing platform 200 can convert the problem description into tasks, extract keywords from each task, and perform vector representation on the extracted keywords to generate a first representation vector of the problem description. The first representation vector at this time can be understood as the user representation vector obtained from the problem representation vector. Of course, the cloud computing platform 200 can also directly perform vector representation on the problem description and use the representation result as the first representation vector, or first convert the problem description into tasks, then perform vector representation on the tasks, and use the representation result as the first representation vector.
[0056] In addition, with the authorization of the user and the permission of the law, the cloud computing platform 200 can also collect the user's historical usage data, such as: data pipelines that have been used, data pipelines or scripts that have been self-built, and other data. Then, the cloud computing platform 200 can perform vector representation on the historical usage data it has collected to obtain the user's interest representation vector, which reflects the user's preference for the data pipeline. Finally, the cloud computing platform 200 can generate a first representation vector based on the interest representation vector and the problem description input by the user. The first representation vector at this time can be understood as the user representation vector obtained from the problem representation vector and the interest representation vector. Exemplarily, when the first representation vector is generated by keywords in the task converted from the problem description, the cloud computing platform 200 can combine the interest representation vector and the keywords in the task for vector representation to generate a first representation vector.
[0057] S302: The cloud computing platform 200 retrieves at least one pipeline template from the template library based on the similarity between the first characterization vector and the characterization vector of the pipeline template in the template library.
[0058] In this embodiment, after obtaining the first characterization vector, the cloud computing platform 200 can calculate the similarity between the first characterization vector and the characterization vectors of each pipeline template stored in the template library by using a cosine similarity algorithm or the like. Then, based on the calculated similarity, at least one pipeline template is retrieved from the template library. For example, the cloud computing platform 200 can sort the calculated similarities from large to small, and use the pipeline templates associated with the first 10 or other number of similarities as the required pipeline templates. Exemplarily, the characterization vector of the pipeline template can be used to describe the content of the pipeline template. Exemplarily, the pipeline template can be understood as a template of a data pipeline.
[0059] S303. The cloud computing platform 200 generates a data pipeline that matches the problem description input by the user based on the retrieved pipeline template.
[0060] In this embodiment, after retrieving the pipeline template, the cloud computing platform 200 can generate a data pipeline that matches the problem description input by the user based on the retrieved pipeline template. For example, when multiple pipeline templates are retrieved, the cloud computing platform 200 can randomly select a pipeline template from them as the data pipeline that matches the problem description input by the user.
[0061] As a possible implementation, Figure 4 As shown, the cloud computing platform 200 generates a data pipeline that matches the problem description input by the user based on the retrieved pipeline template, which may include the following steps:
[0062] S401, the cloud computing platform 200 determines whether the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold. Among them, the cloud computing platform 200 can sort the similarities between the representation vector of the retrieved pipeline template and the first representation vector from large to small (or small to large), and select the largest similarity therefrom, that is, select the maximum value among these similarities. Then, the cloud computing platform 200 can compare the maximum value with the target threshold. When the maximum value is greater than or equal to the template threshold, it indicates that the pipeline template associated with the maximum value can be directly used to process the problem description input by the user, and therefore, S402 can be executed. When the maximum value is less than the template threshold, it indicates that the pipeline template associated with the maximum value cannot be directly used to process the problem description input by the user, and therefore, S403 can be executed.
[0063] S402, the cloud computing platform 200 generates a data pipeline that matches the problem description input by the user based on the pipeline template associated with the maximum value. The cloud computing platform 200 can convert the pipeline template associated with the maximum value into a pipeline description file in a readable format, and call the pipeline creation API to generate the final data pipeline, so that the data pipeline that matches the problem description input by the user is obtained.
[0064] S403, the cloud computing platform 200 retrieves at least one job template from the template library based on the similarity between the first characterization vector and the characterization vector of the job template in the template library. The cloud computing platform 200 may calculate the similarity between the first characterization vector and the characterization vectors of each job template stored in the template library by using a cosine similarity algorithm or the like. Then, based on the calculated similarity, at least one job template is retrieved from the template library. For example, the cloud computing platform 200 may sort the calculated similarities from large to small, and use the job templates associated with the first 10 or other number of similarities as the required job templates. Exemplarily, a pipeline template may include at least one job template. A job template may be an execution unit in a pipeline template. Exemplarily, the characterization vector of a job template may be used to describe the content of the job template.
[0065] S404, the cloud computing platform 200 generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template. The cloud computing platform 200 can rearrange the retrieved pipeline template and job template through LLM, etc. to obtain a data pipeline. Further, the cloud computing platform 200 can convert the rearranged data pipeline into a pipeline description file in a readable format, and call the pipeline creation API to generate the final data pipeline, so that a data pipeline that matches the problem description input by the user is obtained. In some embodiments, considering that the retrieved pipeline template and job template may be different from the problem description input by the user in terms of application field, if all pipeline templates and job templates are directly rearranged, the matching degree between the obtained data pipeline and the problem description input by the user may be poor. Therefore, in order to avoid the above situation, the cloud computing platform 200 can first filter out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates based on user configuration or preset business rules. Then, based on the filtered pipeline templates and job templates, a data pipeline that matches the problem description input by the user is generated.
[0066] In this way, when a user has a data pipeline demand, after the user enters a problem description, the most relevant and useful data pipeline is recommended to the user based on the problem description entered by the user, so that the user does not need to enter detailed operation instructions or click actions, etc., which improves the generation efficiency of the data pipeline and reduces the difficulty of generating the data pipeline.
[0067] In some embodiments, when the number of job templates and pipeline templates carried in the template library is insufficient, the quality of the data pipeline generated subsequently may be poor. Therefore, in this embodiment, for preparing job templates and pipeline templates, high-quality expert templates can be used as seeds to progressively generate job templates first and then pipeline templates, thereby improving the generation efficiency of job templates and pipeline templates and reducing the labor cost of experts writing templates. Specifically, Figure 5 As shown, the following steps may be included:
[0068] S501, the cloud computing platform 200 receives the seed job template imported by the user from the seed job template library, and saves the seed job template in the job template library. In this embodiment, the user can import the seed job template in the seed job template library into the job template library on the cloud computing platform 200 at the front end of the cloud computing platform 200. Among them, the seed job template can be, but is not limited to, a job template written by a domain expert. After the import is successful, the user can trigger a physical button or a virtual button at the front end of the cloud computing platform 200 for indicating the start of training, so that the cloud computing platform 200 starts to generate the job template, that is, executes the subsequent job template generation process.
[0069] S502: The cloud computing platform 200 randomly samples questions from a data source, and randomly samples job templates from a job template library, and transmits the sampled questions and job templates to a job generator.
[0070] S503: The cloud computing platform 200 processes the sampled questions and the job templates through the job generator to generate a new job template. The job generator may change the prompt during the processing.
[0071] S504, the cloud computing platform 200 verifies the new job template, such as whether the new job template can be run, and feeds back the verification results to the job generator, so that the job generator can change the prompt according to the feedback results, for example, change the constraints, etc. In addition, the cloud computing platform 200 can save the verified job templates to the job template library. In addition, after the cloud computing platform 200 completes the verification, it can also introduce manual verification methods, such as: manually judging whether the dependencies between the components in the job template are appropriate, and feeding back the results of the manual inspection to the job generator, so that the job generator can change the prompt according to the manual feedback results, thereby further improving the generation quality of the job template. Repeating S502 to S504 can accumulate a large number of job templates. After the accumulated job templates meet the requirements, the pipeline template generation process can be executed.
[0072] S505, the cloud computing platform 200 receives the seed pipeline template imported by the user from the seed pipeline template library, and saves the seed pipeline template in the pipeline template library. In the present embodiment, the user can import the seed pipeline template in the seed pipeline template library into the pipeline template library on the cloud computing platform 200 at the front end of the cloud computing platform 200. Among them, the seed pipeline template can be, but is not limited to, a pipeline template written by a domain expert. After the import is successful, the user can trigger a physical key or a virtual key for indicating the start of training at the front end of the cloud computing platform 200, so that the cloud computing platform 200 starts to generate the pipeline template, that is, execute the subsequent pipeline template generation process. In some embodiments, the user can also import the seed operation template and the seed pipeline template to the cloud computing platform 200 at the same time. At this time, after the accumulated operation template meets the demand, the subsequent pipeline generation process can be automatically executed, or it can be executed after the user's permission. It can be determined according to the actual situation and is not limited here.
[0073] S506. The cloud computing platform 200 randomly samples a job template from the job template library, and randomly samples a pipeline template from the pipeline template library, and transmits the sampled job template and pipeline template to the pipeline generator.
[0074] S507: The cloud computing platform 200 processes the sampled job template and pipeline template through the pipeline generator to generate a new pipeline template. The pipeline generator may change the prompt during the processing.
[0075] S508, the cloud computing platform 200 verifies the new pipeline template, such as whether the new pipeline template can be run, and feeds back the verification results to the pipeline generator, so that the pipeline generator can change the prompt according to the feedback results, for example, change the constraints, etc. In addition, the cloud computing platform 200 can save the pipeline templates that have passed the verification to the pipeline template library. In addition, after the cloud computing platform 200 completes the verification, it can also introduce manual verification methods, such as: manually judging whether the pipeline template meets expectations, etc., and feeding back the results of the manual inspection to the pipeline generator, so that the pipeline generator can change the prompt according to the manual feedback results, thereby further improving the generation quality of the pipeline template. Repeating S506 to S508 can accumulate a large number of pipeline templates.
[0076] In this embodiment, after the accumulated job templates meet expectations, each job template in the job template library can be processed by the LLM-based job description generator to generate a text description of each job template. The text description of each job template is vectorized by the NN-based job characterizer to obtain a characterization vector of each job template. Then, the characterization vector of each job template can be indexed by the vector database, so that the corresponding job template can be found by searching the characterization vector.
[0077] In addition, after the accumulated pipeline templates meet expectations, the LLM-based pipeline description generator can be used to process each pipeline template in the pipeline template library to generate a text description of each pipeline template. The NN-based pipeline characterizer can be used to vectorize the text description of each pipeline template to obtain a characterization vector of each pipeline template. Then, the characterization vector of each pipeline template can be indexed through the vector database, so that the corresponding pipeline template can be found by searching the characterization vector.
[0078] It is understandable that the order of execution of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In addition, the various embodiments described above can be combined according to actual conditions, and the combined solutions are still within the scope of protection of the present application.
[0079] The above is an introduction to the data pipeline generation method based on cloud computing technology provided by the embodiment of the present application. Next, based on the method in the above embodiment, a data pipeline generation device based on cloud computing technology provided by the embodiment of the present application is introduced.
[0080] For example, Figure 6 The schematic diagram of the structure of a data pipeline generation device based on cloud computing technology provided by an embodiment of the present application is shown. Exemplarily, the data pipeline generation device based on cloud computing technology can be deployed on a cloud computing platform, but is not limited to it. The cloud computing platform runs on an infrastructure, and the infrastructure includes multiple data centers set up in different regions, and each data center includes multiple servers. Figure 6As shown, the data pipeline generation device 600 based on cloud computing technology includes: a vector representation template 601, a template retrieval module 602 and a pipeline generation module 603. The vector representation module 601 is used to generate a first representation vector of the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use. The template retrieval module 602 is used to retrieve at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library. The pipeline generation module 603 is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template.
[0081] In some embodiments, when the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module 603 is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value.
[0082] When the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is less than the target threshold, the template retrieval template is used to retrieve at least one job template from the template library based on the similarity between the first representation vector and the representation vector of the job template in the template library. The pipeline generation module 603 is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template and the job template. , wherein a job template is an execution unit in a pipeline template, and a pipeline template includes at least one job template.
[0083] In some embodiments, when the pipeline generation module 603 generates a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, it is specifically used to: based on business rules, filter out at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates respectively; based on the filtered pipeline templates and job templates, generate a data pipeline that matches the problem description.
[0084] In some embodiments, when generating a first representation vector of a problem description based on a problem description input by a user, the vector representation module 601 is specifically configured to: convert the problem description into at least one task; and generate a first representation vector based on keywords included in the at least one task. A task is at least one processing node on a data pipeline.
[0085] In some embodiments, when the vector representation module 601 generates a first representation vector of a problem description based on a problem description input by a user, it is specifically used to: obtain a user's interest representation vector, where the interest representation vector is used to represent the user's preference for the data pipeline; and generate a first representation vector based on the interest representation vector and the problem description.
[0086] In some embodiments, the device also includes: a job template generation module (not shown in the figure), which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the problem description obtained by sampling, and by changing the prompt words; and store the new job template in the template library after verifying that the new job template is legal.
[0087] In some embodiments, the device also includes: a pipeline template generation module (not shown in the figure), which is used to: receive a seed pipeline template imported by a user; generate a new pipeline template based on the seed pipeline template and the job template sampled from the template library, and by changing the prompt words; and store the new pipeline template in the template library after verifying that the new pipeline template is legal.
[0088] In some embodiments, Figure 6 The vector representation template 601, template retrieval module 602 and pipeline generation module 603 shown in the figure can be implemented by software or by hardware. Exemplarily, the implementation of the vector representation template 601 is introduced below by taking the vector representation template 601 as an example. Similarly, the implementation of the template retrieval module 602 and the pipeline generation module 603 can refer to the implementation of the vector representation template 601.
[0089] As an example of a software functional unit, the vector representation template 601 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the vector representation template 601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Generally, a region may include multiple AZs.
[0090] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0091] As an example of a hardware functional unit, the vector representation template 601 may include at least one computing device, such as a server, etc. Alternatively, the vector representation template 601 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0092] The multiple computing devices included in the vector representation template 601 can be distributed in the same region or in different regions. The multiple computing devices included in the vector representation template 601 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the vector representation template 601 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0093] It should be noted that, in other embodiments, the vector representation template 601 can be used to execute any step of the data pipeline generation method based on cloud computing technology described in the above embodiment, the template retrieval module 602 can also be used to execute any step of the data pipeline generation method based on cloud computing technology described in the above embodiment, and the pipeline generation module 603 can also be used to execute any step of the data pipeline generation method based on cloud computing technology described in the above embodiment. In addition, any two of the vector representation template 601, the template retrieval module 602 and the pipeline generation module 603 can also be combined together to be responsible for executing any step of the data pipeline generation method based on cloud computing technology described in the above embodiment. In addition, the steps that the vector representation template 601, the template retrieval module 602 and the pipeline generation module 603 are responsible for implementing can also be specified as needed, and the different steps of the data pipeline generation method based on cloud computing technology described in the above embodiment can be implemented by the vector representation template 601, the template retrieval module 602 and the pipeline generation module 603 respectively. Figure 6 The data pipeline generation device 600 based on cloud computing technology has all the functions shown.
[0094] The present application also provides a computing device 700. Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.
[0095] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 704 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 704 may include a path for transmitting information between various components of the computing device 700 (eg, the memory 706, the processor 704, and the communication interface 708).
[0096] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0097] The memory 706 may include a volatile memory, such as a random access memory (RAM). The processor 704 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0098] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the aforementioned Figure 6 The functions of the vector representation template 601, the template retrieval module 602 and the pipeline generation module 603 shown in the above embodiment are implemented to realize the data pipeline generation method based on cloud computing technology described in the above embodiment. That is, the memory 706 stores instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.
[0099] Alternatively, the memory 706 stores executable codes, and the processor 704 executes the executable codes to respectively implement the aforementioned Figure 6 The functions of the data pipeline generation device 600 based on cloud computing technology shown in the embodiment are implemented to realize the data pipeline generation method based on cloud computing technology described in the above embodiment. That is, the memory 706 stores instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.
[0100] The communication interface 703 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.
[0101] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0102] like Figure 8 As shown, the computing device cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.
[0103] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also respectively store some instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment. In other words, the combination of one or more computing devices 700 can jointly execute the instructions for executing the data pipeline generation method based on cloud computing technology described in the above embodiment.
[0104] It should be noted that the memory 706 in different computing devices 700 in the computing device cluster may store different instructions, which are respectively used to execute the aforementioned Figure 6 The data pipeline generation device 600 based on cloud computing technology shown in FIG. 7 shows some functions. That is, the instructions stored in the memory 706 in different computing devices 700 can implement the functions of one or more modules in the vector representation template 601, the template retrieval module 602 and the pipeline generation module 603.
[0105] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig. 9 A possible implementation is shown. Fig. 9 As shown, two computing devices 700A and 700B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the vector representation template 701. At the same time, the memory 706 in the computing device 700B stores instructions for executing the functions of the template retrieval module 602 and the pipeline generation module 603.
[0106] It should be understood that Fig. 9 The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.
[0107] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 8 and Fig. 9The connection mode of the computing device cluster is different in that the memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the method in the above embodiment.
[0108] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also respectively store partial instructions for executing the aforementioned data pipeline generation method based on cloud computing technology. In other words, the combination of one or more computing devices 700 can jointly execute instructions for executing the aforementioned data pipeline generation method based on cloud computing technology.
[0109] Based on the method in the above embodiment, the embodiment of the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method in the above embodiment; or, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the above embodiment. Exemplarily, the computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc.
[0110] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product containing instructions, which, when executed by a computing device, enables the computing device to execute the method in the above embodiment, or, when executed by a computing device cluster, enables the computing device cluster to execute the method in the above embodiment.
[0111] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0112] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0113] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0114] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data pipeline generation method based on cloud computing technology, characterized in that: Applied to a cloud computing platform, the cloud computing platform runs on an infrastructure, the infrastructure includes a plurality of data centers arranged in different regions, each data center includes a plurality of servers; The method comprises: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use; The cloud computing platform retrieves at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library; The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template.
2. The method according to claim 1, characterized in that The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template, including: When the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the cloud computing platform generates a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; When the maximum value is less than the target threshold, the cloud computing platform retrieves at least one job template from the template library based on the similarity between the first characterization vector and the characterization vector of the job template in the template library, and generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, wherein one of the job templates is an execution unit in one of the pipeline templates, and one of the pipeline templates includes at least one of the job templates.
3. The method according to claim 2, characterized in that The cloud computing platform generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, including: The cloud computing platform selects at least one pipeline template and at least one job template from the retrieved pipeline templates and job templates based on the business rules; The cloud computing platform generates a data pipeline that matches the problem description based on the screened pipeline templates and job templates.
4. The method according to any one of claims 1 to 3, characterized in that: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, including: The cloud computing platform converts the problem description into at least one task, wherein one task is at least one processing node on the data pipeline; The cloud computing platform generates the first representation vector based on keywords included in the at least one task.
5. The method according to any one of claims 1 to 4, characterized in that: The cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, including: The cloud computing platform obtains an interest representation vector of the user, where the interest representation vector is used to represent the user's preference for the data pipeline; The cloud computing platform generates the first representation vector based on the interest representation vector and the problem description.
6. The method according to any one of claims 1 to 5, characterized in that: Before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: The cloud computing platform receives a seed job template imported by a user; The cloud computing platform generates a new job template based on the seed job template and the sampled problem description and by changing the prompt word; The cloud computing platform stores the new job template in the template library after verifying that the new job template is legal.
7. The method according to any one of claims 1 to 6, characterized in that: Before the cloud computing platform generates a first characterization vector for characterizing the problem description based on the problem description input by the user, the method further includes: The cloud computing platform receives a seed pipeline template imported by a user; The cloud computing platform generates a new pipeline template based on the seed pipeline template and the job template sampled from the template library and by changing the prompt words; The cloud computing platform stores the new pipeline template in the template library after verifying that the new pipeline template is legal.
8. A data pipeline generation device based on cloud computing technology, characterized in that: Deployed on a cloud computing platform, the cloud computing platform runs on an infrastructure, the infrastructure includes multiple data centers located in different regions, each data center includes multiple servers; The device comprises: A vector characterization module, used to generate a first characterization vector for characterizing the problem description based on a problem description input by a user, wherein the problem description is used to describe the content of the data pipeline that the user expects to use; A template retrieval module, configured to retrieve at least one pipeline template from the template library based on the similarity between the first representation vector and the representation vector of the pipeline template in the template library; The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline template.
9. The device according to claim 8, characterized in that In the case where the maximum value of the similarity between the representation vector of the retrieved pipeline template and the first representation vector is greater than or equal to the target threshold, the pipeline generation module is used to generate a data pipeline that matches the problem description based on the pipeline template associated with the maximum value; In the case where the maximum value is less than the target threshold, the template retrieval template is used to retrieve at least one job template from the template library based on the similarity between the first characterization vector and the characterization vectors of the job templates in the template library; The pipeline generation module is used to generate a data pipeline that matches the problem description based on the retrieved pipeline templates and job templates, wherein one job template is an execution unit in one pipeline template, and one pipeline template includes at least one job template.
10. The device according to claim 9, characterized in that When the pipeline generation module generates a data pipeline that matches the problem description based on the retrieved pipeline template and job template, it is specifically used to: Based on the business rules, at least one assembly line template and at least one job template are respectively selected from the retrieved assembly line templates and job templates; Based on the screened pipeline templates and job templates, a data pipeline that matches the problem description is generated.
11. The device according to any one of claims 8 to 10, characterized in that: When the vector representation module generates a first representation vector for representing the problem description based on the problem description input by the user, it is specifically used to: Convert the problem description into at least one task; The first representation vector is generated based on keywords included in the at least one task.
12. The device according to claim 11, characterized in that When the vector representation module generates a first representation vector for representing the problem description based on the problem description input by the user, it is specifically used to: Obtaining an interest representation vector of the user, where the interest representation vector is used to represent the user's preference for the data pipeline; The first representation vector is generated based on the interest representation vector and the problem description.
13. The device according to any one of claims 8 to 12, characterized in that: The device also includes: a job template generation module, which is used to: receive a seed job template imported by a user; generate a new job template based on the seed job template and the sampled problem description and by changing the prompt word; When the new job template is verified to be legal, the new job template is stored in the template library.
14. The device according to any one of claims 8 to 13, characterized in that: The device also includes: a pipeline template generation module, which is used to: Receive the seed pipeline template imported by the user; Based on the seed pipeline template and the job template sampled from the template library, and by changing the prompt words, a new pipeline template is generated; When the new pipeline template is verified to be legal, the new pipeline template is stored in the template library.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1-7.
16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7, wherein the computing device cluster comprises at least one computing device.
17. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7, wherein the computing device cluster includes at least one computing device.