Data processing method and device, electronic equipment and nonvolatile storage medium
By obtaining user demand text and using metadata vector library to adjust prompt words, and combining with large language models to analyze target prompt words, automated data extraction conversion loading processing is realized, solving the problem of low data synchronization efficiency caused by manual configuration parameters in existing ETL solutions.
Patent Information
- Application Number
- CN202510221785.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing ETL implementation solution requires manual configuration of a large number of parameters, resulting in low data synchronization efficiency.
By obtaining the user demand text, determining the initial prompt word, adjusting the prompt word using the metadata vector library, generating task scripts, and analyzing the target prompt word through a large language model to realize data extraction, transformation and loading.
It reduces the need for manual configuration parameters, improves the efficiency of data synchronization, and simplifies the ETL processing process.
Smart Images

Figure CN120067197A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, electronic device, and non-volatile storage medium. Background Art
[0002] The MSS (management support system) service realizes the comprehensive sharing and efficient utilization of enterprise data by integrating enterprise big data, including management (M domain), operation (O domain), and business (B domain), thereby improving the management efficiency and decision-making ability of the enterprise. In order for the enterprise to better integrate and manage various types of business data, eliminate information islands, and ensure the accuracy and consistency of data, before effectively mining high-value data, it is first necessary to fuse and converge the original data (ODS (Operational Data Store) layer data fusion), and then realize the efficient and stable extraction, transformation, and loading (Extract-Transform-Load, ETL) of the data originally scattered in various systems of the enterprise to build an enterprise-level data warehouse and data lake. ETL has become a bridge and transfer station for the flow of internal and external data of the enterprise and plays an important role in scenarios such as data transmission, backup, transformation, and aggregation.
[0003] The ETL implementation solution in the related art is realized by a data synchronization tool or an open-source tool through manual configuration of synchronization conditions, synchronization scripts, and scheduled scheduling. It has relatively high requirements for configuration personnel. In addition to understanding the data model, a large number of parameters need to be configured, or business synchronization scripts need to be written, etc. There are technical problems such as complex ETL data synchronization process and low efficiency.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a data processing method, apparatus, electronic device, and non-volatile storage medium to at least solve the technical problem of low data synchronization efficiency caused by the need to rely on manual configuration of a large number of parameters in the solution for extracting, transforming, and loading data in the related art.
[0006] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining a user requirement text and determining an initial prompt corresponding to the user requirement text, where the user requirement text is used to represent the requirement information of the user for data extraction, transformation, and loading processing, and the initial prompt includes: initial metadata information; determining standard metadata information corresponding to the initial metadata information in the metadata vector library, where the standard metadata information is used to represent the schema information corresponding to the metadata in the data domain; adjusting the initial prompt according to the standard metadata information to obtain a target prompt; analyzing the target prompt using a large language model to obtain a task script, and performing data extraction, transformation, and loading processing on the data in the data domain by executing the task script, where the task script is used to implement the requirements corresponding to the user requirement text.
[0007] Optionally, the initial prompt further includes: source data source, target data source, cleaning and transformation conditions, scheduling instructions; determining the initial prompt corresponding to the user requirement text includes: performing semantic analysis on the user requirement text using a large language model to identify the initial metadata information included in the user requirement text, where the initial metadata information includes at least one of the following: data source name, table name, field name; determining the source data source and target data source corresponding to the extraction, transformation, and loading processing in the user requirement text, where the extraction, transformation, and loading processing is used for data synchronization, the source data source is used to represent the source of the data that needs to be synchronized, and the target data source is used to represent the target location of the data synchronization; determining the cleaning and transformation conditions and scheduling instructions corresponding to the extraction, transformation, and loading processing in the user requirement text, where the cleaning and transformation conditions are used to represent the limiting conditions during the extraction, transformation, and loading processing of the data, and the cleaning and transformation conditions include at least one of the following: data filtering conditions, field transformation conditions, and the scheduling instructions are used to represent the time and period for which data synchronization needs to be performed.
[0008] Optionally, determining the standard metadata information corresponding to the initial metadata information in the metadata vector library includes: performing word segmentation processing on the initial metadata information to obtain multiple word segments, and determining the word frequency corresponding to each word segment, where the word frequency is used to represent the occurrence frequency of the word segment in the initial metadata information; converting the initial metadata information into a corresponding word embedding vector according to the word frequency corresponding to the word segment to obtain a first vector; determining the similarity parameter between the first vector corresponding to the initial metadata information and the second vector corresponding to the standard metadata information in the metadata vector library, where the standard metadata information includes at least one of the following: table name, field name, primary key, field data type, field data precision and length, partition, bucket, index, character set, text delimiter, storage format, storage path; determining the standard metadata information corresponding to the second vector with the highest similarity parameter in the metadata vector library as the standard metadata information corresponding to the initial metadata information.
[0009] Optionally, adjusting the initial prompt according to the standard metadata information to obtain the target prompt includes: determining the first standard metadata information corresponding to the source data source and the second standard metadata information corresponding to the target data source in the standard metadata information; sending the first standard metadata information and the second standard metadata information to the front-end interaction interface for display; in the case of detecting a confirmation instruction triggered on the front-end interaction interface, filling the first standard metadata information, the second standard metadata information, as well as the cleaning and transformation conditions and the scheduling instruction into the corresponding placeholder positions in the prompt template to obtain the target prompt, where the confirmation instruction is used to indicate that the user has confirmed that the information of the first standard metadata information and the second standard metadata information is correct, and the prompt template is used to instruct the large language model to generate a task script based on the provided various types of information.
[0010] Optionally, analyzing the target prompt using the large language model to obtain the task script includes: using the large language model to generate a structured query statement corresponding to the target prompt, where the structured query statement is used to synchronize data from the source data source to the target data source according to the cleaning and transformation conditions; determining the task synchronization type corresponding to the target prompt and generating a synchronization script corresponding to the task synchronization type, where the task synchronization type includes: offline synchronization, real-time synchronization; integrating the structured query statement and the synchronization script to obtain the task script.
[0011] Optionally, the large language model is trained through the proximal policy optimization algorithm. During the training process of the large language model, the loss function used for model parameter adjustment is determined based on the reward value parameter and the divergence parameter, where the reward value parameter is used to represent the reward obtained after taking the target action in the target state, and the divergence parameter is used to represent the difference degree between the optimized policy and the old policy for the probability of taking the target action. The optimized policy is the policy adopted by the large language model after model parameter update, and the old policy is the policy adopted by the large language model before model parameter update.
[0012] Optionally, the method further includes: using a scheduling engine to execute the task script according to the synchronization time and period corresponding to the scheduling instruction; obtaining the task execution parameters corresponding to the synchronization task of the task script and generating an intelligent report based on the task execution parameters and the user requirement text, where the task execution parameters include: at least one of the following: task start time, end time, amount of data processed, task status, and the intelligent report is used to represent the task execution situation and the completion situation of the user requirements; sending the intelligent report to the front-end interaction interface for display.
[0013] According to another aspect of the embodiments of the present application, there is also provided a data processing device, including: a preliminary analysis module, configured to obtain a user requirement text and determine an initial prompt word corresponding to the user requirement text, where the user requirement text is used to represent requirement information of the user for data extraction, transformation, and loading processing, and the initial prompt word includes: initial metadata information; a metadata retrieval module, configured to determine standard metadata information corresponding to the initial metadata information in a metadata vector library, where the standard metadata information is used to represent pattern information corresponding to metadata in a data domain; a precise prompt word construction module, configured to adjust the initial prompt word according to the standard metadata information to obtain a target prompt word; a task generation and execution module, configured to analyze the target prompt word by using a large language model to obtain a task script, and perform data extraction, transformation, and loading processing on the data in the data domain by executing the task script, where the task script is used to implement the requirements corresponding to the user requirement text.
[0014] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run a program stored in the memory, and when the program runs, it executes a data processing method.
[0015] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored computer program, and the device where the non-volatile storage medium is located executes a data processing method by running the computer program.
[0016] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of a data processing method.
[0017] In the embodiments of the present application, user requirement text is obtained, and an initial prompt word corresponding to the user requirement text is determined. The user requirement text is used to represent the requirement information of the user for data extraction, transformation, and loading processing. The initial prompt word includes: initial metadata information; standard metadata information corresponding to the initial metadata information in the metadata vector library is determined, where the standard metadata information is used to represent the schema information corresponding to the metadata in the data domain; the initial prompt word is adjusted according to the standard metadata information to obtain a target prompt word; a large language model is used to analyze the target prompt word to obtain a task script, and by executing the task script, data extraction, transformation, and loading processing are performed on the data in the data domain. The task script is used to implement the requirement corresponding to the user requirement text. By means of an artificial intelligence large language model and through an artificial questioning method to propose data synchronization requirements, combining natural language understanding ability and combined retrieval ability of metadata vectors, the purpose of generating a synchronization task script across multiple heterogeneous data sources and executing the synchronization task through a task scheduling engine to complete the transportation process of data from the source end to the target end is achieved. Furthermore, the technical problem of low data synchronization efficiency caused by the need to rely on manual configuration of a large number of parameters in the data extraction, transformation, and loading processing solution in the related technology is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a data processing method according to an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of a data processing method flow according to an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of a technical route of a data synchronization method based on a large language model according to an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of a method flow of extraction, transformation, and loading for data integration based on a large language model according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of a human-machine dialogue interface according to an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of a scenario for outputting standard metadata information in a human-machine dialogue interface according to an embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a Deepseek model structure provided according to an embodiment of the present application;
[0026] Figure 8 is a schematic diagram of a scenario for self-service query results provided according to an embodiment of the present application;
[0027] Figure 9 is a schematic diagram of a scenario for historical conversation query and historical task query provided according to an embodiment of the present application;
[0028] Figure 10 is a schematic diagram of the structure of a data processing device provided according to an embodiment of the present application. Detailed implementation manners
[0029] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] For the convenience of those skilled in the art to better understand the embodiments of the present application, some technical terms or noun explanations involved in the embodiments of the present application are as follows:
[0032] MBO: It refers to the three major data domains in the big data field of the telecommunications industry. Among them, the M domain, that is, the data domain of the management support system, also known as the management domain, abbreviated as MSS. The M domain includes location information, such as the movement trajectories of people, map information, etc.; the O domain, that is, the data domain of the operation support system, also known as the operation domain, abbreviated as OSS. The O domain includes network data, such as signaling, alarms, faults, network resources, etc.; the B domain, that is, the data domain of the business support system, also known as the business domain, abbreviated as BSS. The B domain includes user data and business data, such as users' consumption habits, terminal information, ARPU grouping, business content, business audience, etc.
[0033] ETL: An abbreviation for Extract-Transform-Load, used to describe the process of extracting data from the source end, transforming it, and loading it to the destination end. The term ETL is more commonly used in data warehouses, but its object is not limited to data warehouses.
[0034] ODS layer (Operational Data Store): An important part of the data warehouse Q architecture, which is located between the data source system and the data mart of the data warehouse, mainly used to store the raw data extracted from each business system.
[0035] LLM (Large Language Model): It is a neural network model based on deep learning. By training a large amount of text data, it can perform various natural language processing (NLP) tasks, such as text generation, text translation, and text question answering.
[0036] NLP (Natural Language Processing): An important direction in the fields of computer science and artificial intelligence, aiming to study various theories and methods for achieving effective communication between humans and computers using natural language.
[0037] DataX: An offline data synchronization tool that supports data synchronization between multiple heterogeneous data sources. It adopts a Framework+plugin architecture, including Reader and Writer plugins for data reading and writing. DataX defines data synchronization tasks by configuring JSON files and is applicable to scenarios such as data warehouse construction, data backup, and data migration.
[0038] Flink: It is a framework and distributed processing engine for stateful computing of unbounded and bounded data streams. Flink can provide two types of functions, stream processing and batch processing, based on the same engine, and the data processed can be real-time data or historical data stored in a database.
[0039] PPO (Proximal Policy Optimization) algorithm: It is a policy gradient method in reinforcement learning. Its goal is to optimize a "surrogate" objective function using stochastic gradient ascent to improve the policy. The PPO algorithm is characterized by being able to perform multiple small-batch updates, rather than performing a gradient update for each data sample like the standard policy gradient method.
[0040] In the related art, the ETL implementation solution is achieved through data synchronization tools or open-source tools by manually configuring synchronization conditions, synchronization scripts, and scheduled tasks. Therefore, the requirements for configurators are relatively high. In addition to understanding the data model, a large number of parameters need to be configured, or business synchronization scripts need to be written, etc. The data processing department of an enterprise manages a large number of ETL tasks and process configurations, and often requires the support of technical personnel and operation and maintenance personnel to complete the tasks.
[0041] To solve the above problems and reduce the learning, development, and maintenance costs of ETL, relevant solutions are provided in the embodiments of this application. Based on the artificial intelligence large language model (LLM), by manually asking questions to put forward data synchronization requirements, combined with the powerful natural language understanding ability of the large model and the combined retrieval ability of metadata vectors, synchronous SQL scripts can be generated across multiple heterogeneous data sources and provided to the data synchronization engine to generate a scheduling task template, and then the synchronization task is executed through the task scheduling engine, thereby completing the transportation process of data from the source to the target end. The following will be described in detail.
[0042] According to the embodiments of this application, a method embodiment for data processing is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0043] The method embodiments provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or electronic device) for implementing the data processing method is shown. As Figure 1As shown, the computer terminal 10 (or electronic device) may include one or more processors 102 (shown as 102a, 102b, ……, 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown in, or have a different configuration from that Figure 1 shown.
[0044] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or electronic device). As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0045] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned data processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0046] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0047] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or electronic device).
[0048] Under the above operating environment, an embodiment of the present application provides a data processing method. Figure 2 It is a schematic diagram of a data processing method flow provided according to an embodiment of the present application, as Figure 2 shown. The method includes the following steps:
[0049] Step S202, obtain the user requirement text and determine the initial prompt word corresponding to the user requirement text, where the user requirement text is used to represent the requirement information of the user for data extraction, transformation, and loading processing, and the initial prompt word includes: initial metadata information;
[0050] Step S204, determine the standard metadata information corresponding to the initial metadata information in the metadata vector library, where the standard metadata information is used to represent the schema information corresponding to the metadata in the data domain;
[0051] Step S206, adjust the initial prompt word according to the standard metadata information to obtain the target prompt word;
[0052] Step S208, analyze the target prompt word using a large language model to obtain a task script, and by executing the task script, perform extraction, transformation, and loading processing on the data in the data domain, where the task script is used to implement the requirements corresponding to the user requirement text.
[0053] Through the above steps, by asking data synchronization requirements in an artificial question-and-answer manner, combining the natural language understanding ability and the combined retrieval ability of metadata vectors, the purpose of generating a synchronization task script across multiple heterogeneous data sources and executing the synchronization task through a task scheduling engine to complete the data transportation process from the source end to the target end is achieved, thereby solving the technical problem of low data synchronization efficiency caused by the need to rely on manual configuration of a large number of parameters in the data extraction, transformation, and loading processing solutions in the related art.
[0054] The data processing method in steps S202 to S208 of the embodiments of the present application will be further introduced below.
[0055] The embodiments of the present application are based on LLM natural language recognition instructions, enhance the recognition of data synchronization requirements through technical means, and rely on the big data base to realize an intelligent ETL synchronization system for cross-domain, cross-system, and cross-department data circulation and aggregation. Figure 3 It is a schematic diagram of the technical route of a data synchronization method based on a large language model provided by the embodiments of the present application. As Figure 3 shown, the embodiments of the present application use a large model natural language dialogue to achieve enterprise-level multi-source data ETL, reduce script development and process configuration, lower the learning cost and usage threshold, make full use of the natural language reasoning ability, code generation ability, and multi-round dialogue context understanding ability of LLM, and achieve zero synchronization configuration and zero parameter setting in the entire process of source and target data collection, cleaning, and transformation through multi-layer technical hidden layers, realizing a friendly interaction experience of accurately understanding user business needs and technical transparency. The following is a specific introduction.
[0056] Figure 4 It is a schematic diagram of the method flow of extraction, transformation, and loading for data integration based on a large language model provided by the embodiments of the present application. As Figure 4 shown, the specific steps are as follows.
[0057] First, obtain the user requirement text proposed by the user on the human-computer dialogue interface (front-end interaction interface). In this user requirement text, the user can describe basic information such as the source, target, and synchronization period of the preliminary ETL (extraction, transformation, and loading) process. If there are conditional restrictions on the source data, the user needs to add conditional information in the description.
[0058] In this embodiment, to make the user operation more convenient, when the user first enters the human-computer dialogue mode, the system can provide relevant case prompt operations. For example, as Figure 5 shown, the user enters the synchronization task dialogue under the system prompt: "I need to synchronize the data of the INF_IPDZGL_NMANZX_D table in the first quarter of this year to the Hive data warehouse, partitioned by day, and automatically synchronize at 0:30 every day."
[0059] After that, after the system obtains the user requirement text, it analyzes the user requirement text to determine the initial prompt word corresponding to the user requirement text. The specific steps are as follows.
[0060] In some embodiments of the present application, the initial prompt words further include: source data source, target data source, cleaning and transformation conditions, and scheduling instructions; determining the initial prompt words corresponding to the user requirement text includes the following steps: using a large language model to perform semantic analysis on the user requirement text to identify the initial metadata information included in the user requirement text, where the initial metadata information includes at least one of the following: data source name, table name, field name; determining the source data source and target data source corresponding to the extract, transform, and load (ETL) process in the user requirement text, where the ETL process is used for data synchronization, the source data source is used to represent the source of the data that needs to be synchronized, and the target data source is used to represent the target location of the data synchronization; determining the cleaning and transformation conditions and scheduling instructions corresponding to the ETL process in the user requirement text, where the cleaning and transformation conditions are used to represent the limiting conditions when performing the ETL process on the data, and the cleaning and transformation conditions include at least one of the following: data filtering conditions, field transformation conditions, and the scheduling instructions are used to represent the time and period for which data synchronization is required.
[0061] Specifically, in this embodiment, the user requirement text can be identified in four directions to determine the initial prompt words such as the initial metadata information, source data source, target data source, cleaning and transformation conditions, and scheduling instructions. Specifically, it includes: 1) identifying the initial metadata information, mainly including Chinese information such as data source name, table name, field name, etc. that belong to the category of metadata; 2) identifying the source data source and target data source. Among them, the user requirement text may not necessarily contain information such as the name and type of the clear source data source and target data source. In this case, the source and target data sources to which it belongs can be judged based on the metadata; 3) identifying the cleaning and transformation conditions. If the user needs to filter, transform, etc. the data in the source table during the ETL process, it needs to be clearly stated in the user requirement text. For example: deducting a certain type of data, converting the data of a certain field to text type, etc. In the case of no special prompt, the embodiments of the present application will perform deduplication processing on the target data based on the primary key and then perform synchronization; 4) identifying the scheduling instructions. The ETL offline synchronization task needs to customize the synchronization frequency. For example: synchronize at 0:30 every day, synchronize at 0:30 every Monday, etc.
[0062] After obtaining the initial prompt words through preliminary semantic understanding and splitting, the initial metadata information can be provided to the metadata vector library for retrieval and matching to determine the standard metadata information corresponding to the initial metadata information in the metadata vector library. The specific steps are as follows.
[0063] In some embodiments of the present application, determining the standard metadata information corresponding to the initial metadata information in the metadata vector library includes the following steps: performing word segmentation on the initial metadata information to obtain a plurality of segmented words, and determining the word frequency corresponding to each segmented word, where the word frequency is used to represent the occurrence frequency of the segmented word in the initial metadata information; converting the initial metadata information into a corresponding word embedding vector according to the word frequency corresponding to the segmented word to obtain a first vector; determining a similarity parameter between the first vector corresponding to the initial metadata information and the second vector corresponding to the standard metadata information in the metadata vector library, where the standard metadata information includes at least one of the following: table name, field name, primary key, field data type, field data precision and length, partition, bucket, index, character set, text delimiter, storage format, storage path; determining the standard metadata information corresponding to the second vector with the highest corresponding similarity parameter in the metadata vector library as the standard metadata information corresponding to the initial metadata information.
[0064] Specifically, the initial metadata information can be provided to the metadata vector library for vector similarity retrieval, and the metadata information of the most accurate match with the highest similarity pair is returned, that is, the standard metadata information, which includes information such as the English name of the table, the English and Chinese names of the fields, the primary key, the field data type, the field data precision and length, the partition, the bucket, the index, the character set, the text delimiter, the storage format, and the storage path. In this embodiment, Hive is an internal table by default, and is considered an external table if a specified mounting path is provided.
[0065] The embodiments of the present application achieve reverse metadata normalization management and automatic update, retrieve and identify metadata based on vectorization, establish a unique index for the metadata, accurately locate the data space where the metadata is located, and establish data standards and basis for data synchronization.
[0066] In the embodiments of the present application, the word2vector vector similarity algorithm can be used to calculate the similarity parameter. Among them, Word2vec is a model for generating word vectors, which can map words into a continuous vector space, so that words with similar semantics are also close in the vector space. Word vectors are an important technology in natural language processing, which can capture the semantic and syntactic relationships between words and provide strong support for tasks such as text analysis, sentiment analysis, and text classification.
[0067] In this embodiment, the Word2vec algorithm can be implemented through the Skip-gram model. Among them, in the Skip-gram model, the input layer accepts a target word (usually in the form of a one-hot encoding of the word) and converts it into a sparse one-hot vector. Each word corresponds to a position in the one-hot vector, the value at this position is 1, and the values at the remaining positions are 0; the role of the hidden layer is to map the input one-hot vector to a lower-dimensional word embedding space. In the Skip-gram model, the hidden layer actually does not use an activation function (such as ReLU or sigmoid), but directly generates a word embedding vector through matrix multiplication. The weight matrix of this hidden layer is the word embedding matrix of the model, and each row corresponds to the word vector of a word; the output layer uses a softmax function to map the output of the hidden layer to the probability distribution of each possible context word. The weight matrix of this output layer is a matrix of the same size as the word embedding matrix, which is used to calculate the occurrence probability of each word in the context. Through the softmax function, the model can obtain the predicted probability distribution of each context word.
[0068] Specifically, first perform one-hot encoding to form a V*1 vector for each word in the full text; for the entire vocabulary, it is a V*V matrix; then perform word embedding. In the Skip-gram model, the parameter matrix (W) of the hidden layer has a shape of (V*d), which maps each word to a (d)-dimensional space, so that each word corresponds one-to-one with a certain column of the matrix (W); during the training process of the skip-gram model, a matrix (W') with a shape of (V*d) can be initialized as the weight matrix, where each column corresponds to the word vector when a word is used as a context word. At the same time, there are already two (d)-dimensional matrices at this time: each row of the matrix (W) corresponds to the word vector when a word is used as a center word, and each column of the matrix (W”) corresponds to the word vector when a word is used as a context word. As the window moves, the model will calculate each word as the center word; then predict the context word through the center word. The word vector of the center word (with a shape of (1*d)) will perform an inner product operation with each column of the context word weight matrix (W'), so that the score of each context word can be obtained. Through these scores, the probability of the context word as the predicted word can be calculated. This process will be iteratively performed during the training process, continuously adjusting the weight matrix (W') to minimize the loss function, thereby gradually optimizing the model so that the model can more accurately predict the probability distribution of the context word.
[0069] After all the words in the vocabulary are trained, two word vectors for each word, namely Vi and Ui, can be obtained. Among them, Vi is the vector when the word is used as the center word, and Ui corresponds to the vector when the word is used as the context word. Generally, the Vi vector is selected as the final one and stored in the vector library. The final similarity calculation formula is as follows:
[0070]
[0071] The following is an example to illustrate the process of similarity calculation.
[0072] Suppose the original text is as follows:
[0073] Sentence A: I like watching TV and don't like watching movies.
[0074] Sentence B: I don't like watching TV and don't like watching movies either.
[0075] After word segmentation, we get:
[0076] Sentence A: I / like / watch / TV, don't / like / watch / movies.
[0077] Sentence B: I / don't / like / watch / TV, also / don't / like / watch / movies.
[0078] Calculate the word frequency:
[0079] Sentence A: I 1, like 2, watch 2, TV 1, movies 1, don't 1, also.
[0080] Sentence B: I 1, like 2, watch 2, TV 1, movies 1, don't 2, also 1.
[0081] Word frequency vector:
[0082] Sentence A: [1, 2, 2, 1, 1, 1, 0];
[0083] Sentence B: [1, 2, 2, 1, 1, 2, 1];
[0084] Calculate the vector similarity:
[0085] The similarity calculation formula for Sentence A and Sentence B is as follows:
[0086]
[0087] The closer the cosine value is to 1, the closer the included angle is to 0 degrees, that is, the more similar the two vectors are. Therefore, the cosine similarity between Sentence A and Sentence B above is very high, and their included angle is approximately 20.3 degrees.
[0088] Specifically, as Figure 6As shown, in the embodiments of the present application, accurate source table schema information (i.e., the first standard metadata information corresponding to the source data source) can be retrieved and matched through vector retrieval, and then confirmed with the user again; and, accurate target table schema information (i.e., the second standard metadata information corresponding to the target data source) can also be matched. If it does not exist, the target table will be automatically created and confirmed with the user again.
[0089] After determining the standard metadata information, the initial prompt can be adjusted according to the standard metadata information to obtain the target prompt. The specific steps are as follows.
[0090] In some embodiments of the present application, adjusting the initial prompt according to the standard metadata information to obtain the target prompt includes the following steps: determining the first standard metadata information corresponding to the source data source and the second standard metadata information corresponding to the target data source in the standard metadata information; sending the first standard metadata information and the second standard metadata information to the front-end interaction interface for display; when a confirmation instruction triggered on the front-end interaction interface is detected, filling the first standard metadata information, the second standard metadata information, the cleaning and transformation conditions, and the scheduling instruction into the corresponding placeholder positions in the prompt template to obtain the target prompt, where the confirmation instruction is used to indicate that the user has confirmed that the first standard metadata information and the second standard metadata information are correct, and the prompt template is used to instruct the large language model to generate a task script based on the provided various information.
[0091] Specifically, after obtaining accurate standard metadata information, source and target data sources, and cleaning and transformation conditions, the construction of the accurate prompt (i.e., the target prompt) can be started. The initial prompt can be modified according to the source data source type, the schema information of the source database (i.e., its corresponding standard metadata information), and the cleaning and transformation conditions, etc., to generate a target prompt containing accurate standard metadata information. The target prompt contains information such as the source data source, the target data source, and the synchronization query rule.
[0092] Subsequently, the target prompt can be used as the input to the large model to enable the large model to output relevant SQL scripts and synchronization scripts required for ETL. The specific steps are as follows.
[0093] In some embodiments of the present application, analyzing the target prompt using a large language model to obtain the task script includes: using the large language model to generate a structured query statement corresponding to the target prompt, where the structured query statement is used to synchronize data from the source data source to the target data source according to the cleaning and transformation conditions; determining the task synchronization type corresponding to the target prompt and generating a synchronization script corresponding to the task synchronization type, where the task synchronization type includes: offline synchronization, real-time synchronization; integrating the structured query statement and the synchronization script to obtain the task script.
[0094] Specifically, after the above steps, the large model will output synchronization statement SQL (i.e., Structured Query Language), the source Schema and the target Schema, and generate corresponding synchronization scripts according to the task synchronization type. For example, for offline synchronization, DataX-json is generated, and for real-time synchronization, Flink sql is generated. The system encapsulates this information into an ETL task template to provide a basis for task scheduling. Additionally, in model fine-tuning, through multiple rounds of conversations, the large model can be continuously corrected until the synchronization script information is accurately output. The accuracy of the model is evaluated and a reward mechanism is set up to promote the self-improvement and repair of the model.
[0095] For example, in this embodiment, this step can be executed between the system backend and the AI model layer. The accurate standard prompt words constructed from the accurate schema information of the source and the target are input into the large model to output synchronization SQL. According to the recognized source data source type, data source name, and connection method, the scheduling engine selects the default DataX offline synchronization plugin and automatically sets the synchronization period to 0:30 every day. Based on this information, a task execution template and a task ID are generated. This process requires the support of the backend service and takes a certain waiting time. According to the SQL and scheduling instructions output by the large model, the backend service generates Datax Json and task information. Among them, the querysql content included in the Json is as follows: "SELECT*FROM INF IPDZGL NMANZX DWHERE creation time>='2024-81-01'AND Creation time<='2024-03-31'", which is the ETL task script generated by the large model.
[0096] It should be noted that in this embodiment, the above large language model can be based on the Deepseek model, and its model structure is as Figure 7As shown below, it specifically includes: 1) Input processing: Word embedding: Convert the words in the input text into continuous vector representations. This step is usually implemented using pre-trained word embedding models (such as Word2Vec, GloVe) or a word embedding matrix obtained through training. Position encoding: If the model needs to capture the positional relationships of words in the sequence, position encoding is added to retain the sequence information. 2) Encoder-decoder architecture: Encoder part: Convert the input sequence into context-related hidden states. Although DeepSeek is decoder-only, generally the encoder will include multiple self-attention layers and feed-forward neural networks. Decoder part: The decoder uses the self-attention mechanism to generate the output and only focuses on the generated part when generating each word, ensuring that the model maintains the autoregressive property during both the training and inference phases. 3) Self-attention mechanism: The self-attention layer weights each position of the input and calculates its dependence on other positions. By calculating the weighted sum of each position in the input sequence, the model can capture long-range dependencies. Masked self-attention: In the decoding phase, the masked self-attention mechanism ensures that the model can only consider the words before the current word when generating the current word, to maintain the coherence and consistency of generation. 4) Feed-forward neural network: Usually, a feed-forward neural network follows each self-attention layer to further process the representation of each position. The feed-forward network typically includes two fully connected layers and an activation function (such as ReLU). 5) Output layer: Linear transformation: Map the output of the decoder to the size of the vocabulary to generate a vocabulary probability distribution. 6) Softmax layer: Calculate the probability of each word in the vocabulary through the Softmax function to select the most likely next word. During the model training process, the cross-entropy loss function is usually used to measure the gap between the generated sequence and the true sequence, that is, the loss function. The goal of the model is to minimize this loss function, thereby improving the generation accuracy. Optimization algorithms (such as Adam, AdamW) can be used to update the model parameters to improve the generation performance.
[0097] In addition, to further improve the accuracy of the large language model in the process of generating ETL-related scripts and packaging them into DataX scripts, the embodiment of this application adopts the Proximal Policy Optimization (PPO) algorithm. After model fine-tuning, reinforcement learning is carried out. Through the reward and punishment mechanism, the parameter weights of the model change, tending to the probability event of the optimal solution in the environment, specifically as follows.
[0098] In some embodiments of the present application, the large language model is trained by the proximal policy optimization algorithm. During the training process of the large language model, the loss function used for model parameter adjustment is determined based on the reward value parameter and the divergence parameter. Among them, the reward value parameter is used to represent the reward obtained after taking the target action in the target state, and the divergence parameter is used to represent the difference degree between the optimized policy and the old policy for the probability of taking the target action. The optimized policy is the policy adopted by the large language model after the model parameters are updated, and the old policy is the policy adopted by the large language model before the model parameters are updated.
[0099] Specifically, the traditional policy gradient method in the related art is very sensitive to the learning step size and often faces difficulties in selecting an appropriate step size. In addition, if the change between the new and old policies is too large, it will have a negative impact on the learning process. To solve the above problems, this embodiment adopts the PPO algorithm to introduce a new objective function, allowing for mini-batch updates in multiple training steps, thus solving the problem of difficult step size determination in the policy gradient algorithm. PPO makes the optimization process simpler and more efficient by using the KL divergence as a penalty term. The main steps are as follows:
[0100] 1) Set a reward function in the large model environment:
[0101] R(s, a) = Reward(s, a) - γ × Expected\Reward(s′, a′)
[0102] Where, R(s, a) represents the reward obtained after taking the target action actcion(a) in the target state stat(s); γ is the discount factor; s' is the next state where the agent is located after taking the action a; a' is the action taken in the state s'; Expected\Reward(s′, a′) is the expected reward after taking the action a' in the state s'. During the model training process, the total reward value r t (θ) of all states executed in a certain path of the model in the obtained environment can be used.
[0103] 2) Reward training: Collect a set of interaction data by executing the current policy in the environment. These data include the state, action, reward, and possibly the next state. In this embodiment, for each input prompt word and the large model output schema, the process of multiple state switches and reward acquisitions of the synchronous SQL and DataX scripts can be performed, and an immediate reward is obtained for each next state switch.
[0104] 3) Optimize the objective function: In the PPO algorithm, the objective function usually consists of two parts: the reward value (reward, i.e., the above-mentioned reward value parameter) and the KL divergence between the optimized policy distribution and the old policy distribution (KL divergence, i.e., the above-mentioned divergence parameter). KL can be calculated as the similarity value between the predicted content output between two states and the standard output. The larger the value, the greater the similarity, and the smaller the divergence, the closer the two probability events are. The objective optimization function of PPO is the loss function of the reward model for reinforcement learning, as shown in the following formula:
[0105]
[0106] where is the advantage function, is the probability ratio of the new and old policies, is the divergence, and the clip function restricts the change range of the probability ratio r t (θ), (θ) to prevent the update step from being too large.
[0107] Specifically, during the training process, the gradient ascent method can be used to update the policy parameter θ, where α is the learning rate. By adjusting the learning rate and controlling the learning rate of the large model, the training stability and convergence speed of the model can be adjusted. An appropriate learning rate can help the model quickly find the optimal solution, while too high a learning rate will lead to unstable training and even divergence of the loss function; on the contrary, too low a learning rate may result in a slow convergence speed and increased training time. Repeat the above steps with the new policy parameters until certain stopping criteria are met, such as the policy performance no longer improves or a certain number of iterations have been reached.
[0108] Through the above iterations, reward values are continuously generated in the training environment and then continuously adjusted and optimized through the PPO objective function. The objective optimization function is the loss function (Loss) of the model. When the value of the Loss function reaches the optimal value, the model reaches the best fitting degree. In this patent, through the manual annotation and reinforcement training method, the reward mechanism is used to prompt the model to self-adjust and optimize the parameters, and finally a privatized model that can automatically generate ETL three-element scripts is generated. That is, the model fully understands the technical characteristics of generating ETL scripts during the reward process, has the ability to understand the synchronization script characteristics of various data across platforms and data sources, and no longer requires separate development of acquisition adapters.
[0109] In the embodiments of the present application, a reward mechanism and a PPO batch small gradient optimization algorithm are constructed at the algorithm level. Through reinforcement learning, the model can accurately generate ETL scripts for various MSS-related business systems in this patent. That is, various data source information and task templates for synchronization scripts are generated according to different sources and targets. It can solve the cross-platform, cross-domain, and cross-source SQL syntax recognition and conversion between multi-source relational databases (such as MySQL, PG, etc.) and Hive data warehouses, and there is no need to develop corresponding acquisition adapters for different data source types. It can replace the traditional single combination of acquisition adapters, thereby reducing the R & D cost and software complexity.
[0110] After obtaining the task script generated by the large model, it can be distributedly executed and reported through the scheduling engine. The specific steps are as follows.
[0111] In some embodiments of the present application, the method further includes the following steps: using the scheduling engine to execute the task script according to the synchronization time and period corresponding to the scheduling instruction; obtaining the task execution parameters corresponding to the synchronization task of the task script, and generating an intelligent report based on the task execution parameters and the user requirement text, where the task execution parameters include at least one of the following: task start time, end time, amount of data processed, task status, and the intelligent report is used to characterize the task execution situation and the completion situation of the user requirements; sending the intelligent report to the front-end interaction interface for display.
[0112] Specifically, in this embodiment, the offline synchronization plugin can be defaulted to Datax, and the real-time synchronization component is Flink. In the offline case, the synchronization information is converted into Json and passed to the Datax component, and the DATAX synchronization task is executed on the execution machine. If it is scheduled according to a period, the scheduling engine automatically schedules the DATAX synchronization task according to the period, and outputs the task status and logs to the task monitoring module in real time. In the real-time case, the synchronization sql statement is sent to Flink sql, a Flink Job is created and run, and the task status and performance metrics are output to the task monitoring module to the scheduling engine.
[0113] In the embodiments of the present application, the AI automatically matches a suitable synchronization plugin. In the case where the underlying data synchronization plugin is not explicitly specified in the prompt, the synchronization plugin can be automatically matched according to the metadata characteristics, and the cross-source and cross-platform synchronization difference parameters are eliminated by presetting and optimizing the performance parameters according to the data scale. The performance parameters are automatically preset and optimized according to the data scale and characteristics to ensure the best performance of the system within the resource limit, reducing the attention to technical details and the tuning work.
[0114] In addition, intelligent reports can be integrated and output based on the monitoring metrics output by scheduled tasks or immediate tasks executed by the scheduler, associating with the original requirements of human-machine dialogue, source and target metadata, task ID, task name, task status, exception logs, etc. Self-service query can query reports according to questions, task names, task status, etc.; proactive reports can be sent when the task execution fails or the task is successfully executed.
[0115] For example, after the task is executed according to the scheduling time, the task execution status can be queried through the original dialogue, and the multi-round dialogue is coherent and effective. The self-service query results are as Figure 8 shown; historical dialogue query and historical task query can also be performed, as Figure 9 shown; in addition, users can also preview the data situation after synchronization.
[0116] Through the method of "self-service acquisition of task information + intelligent autonomous reporting", the system can monitor the task progress in real time, automatically detect anomalies and push them actively, answer questions about the task running status at any time, and provide efficient means for operation and maintenance personnel to handle and intervene in faults.
[0117] The solution of this application uses natural language processing (NLP) technology to allow users to operate through simple natural language instructions, enabling ordinary business personnel or ordinary operation and maintenance personnel to complete one-click data synchronization according to production needs in the form of "human-machine dialogue", without the need to pay attention to technical details such as differences between multiple data sources, script specifications, and task scheduling parameters. This method enables non-technical users to easily complete data synchronization tasks and greatly reduces the operation threshold; through AI technology, the cleaning and conversion rules between data are automatically recognized. For example, the modeling conversion from MySQL to hive, the conversion of data type from numeric to text, etc. are intelligently converted by automatically comparing the source and target data structures, reducing the need for manual configuration and rule writing, reducing human errors and improving the accuracy of conversion.
[0118] The NLP2SQL technology automatically converts natural language queries into SQL statements, making the queries and data processing in the data warehousing process more intuitive and efficient. The system does not need to develop multiple sets of plug-ins to match heterogeneous data sources, nor does it need to have the ability to write SQL. It can achieve complex data warehousing and migration, saving a large amount of labor costs. At the R & D level, originally, three people were required to separately perform data synchronization work in the M, B, and O domains of MSS, undertaking a large amount of configuration and script writing workload. After using this patent application, only one person can achieve data synchronization in the three domains, without configuration or script writing, and only need to verify whether the model output results are correct. Through the NLP2SQL technology, the queries and operations in the data migration process can be described and managed in natural language, simplifying the complex process of data migration and reducing the dependence on developers. The metadata vector retrieval technology can quickly identify and match the structures of source data and target data, improving the matching degree of data migration and ensuring the security of data operations. Through the metadata vector retrieval technology, it is possible to quickly match similar schemas, quickly locate and restore critical data after a disaster, shorten the system recovery time, and improve business continuity and disaster recovery capabilities.
[0119] According to an embodiment of the present application, an embodiment of a data processing device is also provided. Figure 10 is a structural schematic diagram of a data processing device provided according to an embodiment of the present application. As Figure 10 shown, the device includes:
[0120] A preliminary analysis module 100, configured to obtain a user requirement text and determine an initial prompt word corresponding to the user requirement text, where the user requirement text is used to represent requirement information of the user for extracting, transforming, and loading data, and the initial prompt word includes: initial metadata information;
[0121] A metadata retrieval module 102, configured to determine standard metadata information corresponding to the initial metadata information in a metadata vector library, where the standard metadata information is used to represent schema information corresponding to metadata in a data domain;
[0122] A precise prompt word construction module 104, configured to adjust the initial prompt word according to the standard metadata information to obtain a target prompt word;
[0123] A task generation and execution module 106, configured to analyze the target prompt word using a large language model to obtain a task script, and perform extraction, transformation, and loading processing on the data in the data domain by executing the task script, where the task script is used to implement the requirements corresponding to the user requirement text.
[0124] Optionally, the initial prompt also includes: source data source, target data source, cleaning and transformation conditions, scheduling instructions; determining the initial prompt corresponding to the user requirement text includes: performing semantic analysis on the user requirement text using a large language model to identify the initial metadata information included in the user requirement text, where the initial metadata information includes at least one of the following: data source name, table name, field name; determining the source data source and target data source corresponding to the extract, transform, and load process in the user requirement text, where the extract, transform, and load process is used for data synchronization, the source data source is used to represent the source of the data that needs to be synchronized, and the target data source is used to represent the target location of the data synchronization; determining the cleaning and transformation conditions and scheduling instructions corresponding to the extract, transform, and load process in the user requirement text, where the cleaning and transformation conditions are used to represent the limiting conditions when performing the extract, transform, and load process on the data, and the cleaning and transformation conditions include at least one of the following: data filtering conditions, field transformation conditions, and the scheduling instructions are used to represent the time and period for which data synchronization needs to be performed.
[0125] Optionally, determining the standard metadata information corresponding to the initial metadata information in the metadata vector library includes: performing word segmentation on the initial metadata information to obtain multiple word segments, and determining the word frequency corresponding to each word segment, where the word frequency is used to represent the occurrence frequency of the word segment in the initial metadata information; converting the initial metadata information into a corresponding word embedding vector according to the word frequency corresponding to the word segment to obtain a first vector; determining the similarity parameter between the first vector corresponding to the initial metadata information and the second vector corresponding to the standard metadata information in the metadata vector library, where the standard metadata information includes at least one of the following: table name, field name, primary key, field data type, field data precision and length, partition, bucket, index, character set, text delimiter, storage format, storage path; determining the standard metadata information corresponding to the second vector with the highest corresponding similarity parameter in the metadata vector library as the standard metadata information corresponding to the initial metadata information.
[0126] Optionally, adjusting the initial prompt according to the standard metadata information to obtain the target prompt includes: determining the first standard metadata information corresponding to the source data source and the second standard metadata information corresponding to the target data source in the standard metadata information; sending the first standard metadata information and the second standard metadata information to the front-end interaction interface for display; in the case of detecting a confirmation instruction triggered on the front-end interaction interface, filling the first standard metadata information, the second standard metadata information, and the cleaning and transformation conditions and scheduling instructions into the corresponding placeholder positions in the prompt template to obtain the target prompt, where the confirmation instruction is used to represent that the user has confirmed that the information of the first standard metadata information and the second standard metadata information is correct, and the prompt template is used to instruct the large language model to generate a task script based on the provided various types of information.
[0127] Optionally, a large language model is used to analyze the target prompt, and the obtained task script includes: using the large language model to generate a structured query statement corresponding to the target prompt, where the structured query statement is used to synchronize data from the source data source to the target data source according to the cleaning and transformation conditions; determining the task synchronization type corresponding to the target prompt, and generating a synchronization script corresponding to the task synchronization type, where the task synchronization type includes: offline synchronization, real-time synchronization; integrating the structured query statement and the synchronization script to obtain the task script.
[0128] Optionally, the large language model is trained by the proximal policy optimization algorithm. During the training process of the large language model, the loss function used for model parameter adjustment is determined based on the reward value parameter and the divergence parameter. The reward value parameter is used to represent the reward obtained after taking the target action in the target state, and the divergence parameter is used to represent the difference degree between the optimized policy and the old policy for the probability of taking the target action. The optimized policy is the policy adopted by the large language model after the model parameters are updated, and the old policy is the policy adopted by the large language model before the model parameters are updated.
[0129] Optionally, the data processing device is further configured to: use a scheduling engine to execute the task script according to the synchronization time and period corresponding to the scheduling instruction; obtain the task execution parameters corresponding to the synchronization task of the task script, and generate an intelligent report based on the task execution parameters and the user requirement text, where the task execution parameters include at least one of the following: task start time, end time, amount of data processed, task status, and the intelligent report is used to represent the task execution situation and the completion situation of the user requirements; send the intelligent report to the front-end interaction interface for display.
[0130] It should be noted that each module in the above data processing device can be a program module (for example, a set of program instructions that implement a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the presentation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0131] It should be noted that the data processing device provided in this embodiment can be used to execute Figure 2 the data processing method shown, therefore, the relevant explanations of the above data processing method also apply to the embodiments of the present application, and will not be repeated here.
[0132] The embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program. Wherein, the device where the non-volatile storage medium is located executes the following data processing method by running the computer program: obtaining a user requirement text and determining an initial prompt word corresponding to the user requirement text, where the user requirement text is used to represent the requirement information of the user for data extraction, transformation, and loading processing, and the initial prompt word includes: initial metadata information; determining the standard metadata information corresponding to the initial metadata information in the metadata vector library, where the standard metadata information is used to represent the pattern information corresponding to the metadata in the data domain; adjusting the initial prompt word according to the standard metadata information to obtain a target prompt word; analyzing the target prompt word by using a large language model to obtain a task script, and performing extraction, transformation, and loading processing on the data in the data domain by executing the task script, where the task script is used to implement the requirements corresponding to the user requirement text.
[0133] The embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of the data processing method described in each embodiment of the present application when executed by a processor: obtaining a user requirement text and determining an initial prompt word corresponding to the user requirement text, where the user requirement text is used to represent the requirement information of the user for data extraction, transformation, and loading processing, and the initial prompt word includes: initial metadata information; determining the standard metadata information corresponding to the initial metadata information in the metadata vector library, where the standard metadata information is used to represent the pattern information corresponding to the metadata in the data domain; adjusting the initial prompt word according to the standard metadata information to obtain a target prompt word; analyzing the target prompt word by using a large language model to obtain a task script, and performing extraction, transformation, and loading processing on the data in the data domain by executing the task script, where the task script is used to implement the requirements corresponding to the user requirement text.
[0134] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0135] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0136] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0137] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0139] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. And the aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0140] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A data processing method, characterized in that: include: Obtaining a user demand text and determining an initial prompt word corresponding to the user demand text, wherein the user demand text is used to represent the user's demand information for extracting, transforming and loading data, and the initial prompt word includes: initial metadata information; Determining standard metadata information corresponding to the initial metadata information in the metadata vector library, wherein the standard metadata information is used to represent the pattern information corresponding to the metadata in the data domain; According to the standard metadata information, the initial prompt word is adjusted to obtain a target prompt word; The target prompt word is analyzed by using a large language model to obtain a task script, and the data in the data domain is extracted, converted, loaded and processed by executing the task script, wherein the task script is used to realize the requirements corresponding to the user requirement text.
2. The data processing method according to claim 1, characterized in that: The initial prompt words also include: source data source, target data source, cleaning conversion conditions, and scheduling instructions; Determining the initial prompt words corresponding to the user requirement text includes: Using the large language model to perform semantic analysis on the user demand text to identify the initial metadata information contained in the user demand text, wherein the initial metadata information includes at least one of the following: data source name, table name, field name; Determine the source data source and the target data source corresponding to the extraction, transformation and loading process in the user requirement text, wherein the extraction, transformation and loading process is used for data synchronization, the source data source is used to represent the source of data that needs to be synchronized, and the target data source is used to represent the target location of data synchronization; Determine the cleaning conversion conditions and the scheduling instructions corresponding to the extraction, conversion and loading processing in the user requirement text, wherein the cleaning conversion conditions are used to represent the restriction conditions when the extraction, conversion and loading processing is performed on the data, and the cleaning conversion conditions include at least one of the following: data filtering conditions, field conversion conditions, and the scheduling instructions are used to represent the time and period when data synchronization is required.
3. The data processing method according to claim 2, characterized in that: Determining the standard metadata information corresponding to the initial metadata information in the metadata vector library includes: Performing word segmentation processing on the initial metadata information to obtain multiple word segments, and determining a word frequency corresponding to each word segment, wherein the word frequency is used to represent the frequency of occurrence of the word segment in the initial metadata information; According to the word frequency corresponding to the word segmentation, convert the initial metadata information into a corresponding word embedding vector to obtain a first vector; Determine a similarity parameter between the first vector corresponding to the initial metadata information and the second vector corresponding to the standard metadata information in the metadata vector library, wherein the standard metadata information includes at least one of the following: table name, field name, primary key, field data type, field data precision and length, partition, bucket, index, character set, text delimiter, storage format, and storage path; The standard metadata information corresponding to the second vector having the highest similarity parameter in the metadata vector library is determined as the standard metadata information corresponding to the initial metadata information.
4. The data processing method according to claim 3, characterized in that: According to the standard metadata information, the initial prompt word is adjusted to obtain the target prompt word including: Determine first standard metadata information corresponding to the source data source and second standard metadata information corresponding to the target data source in the standard metadata information; Sending the first standard metadata information and the second standard metadata information to a front-end interactive interface for display; When a confirmation instruction triggered on the front-end interactive interface is detected, the first standard metadata information, the second standard metadata information, the cleaning conversion condition and the scheduling instruction are filled into the corresponding placeholder positions in the prompt word template to obtain the target prompt word, wherein the confirmation instruction is used to indicate that the user has confirmed that the first standard metadata information and the second standard metadata information are correct, and the prompt word template is used to instruct the large language model to generate the task script based on the various types of information provided.
5. The data processing method according to claim 2, characterized in that: The target prompt word is analyzed using a large language model to obtain a task script including: Using the large language model, generating a structured query statement corresponding to the target prompt word, wherein the structured query statement is used to synchronize data from the source data source to the target data source according to the cleaning conversion condition; Determine the task synchronization type corresponding to the target prompt word, and generate a synchronization script corresponding to the task synchronization type, wherein the task synchronization type includes: offline synchronization and real-time synchronization; The structured query statement and the synchronization script are integrated to obtain the task script.
6. The data processing method according to claim 5, characterized in that: The large language model is obtained by training through a proximal strategy optimization algorithm. During the training process of the large language model, a loss function used to adjust the model parameters is determined based on a reward value parameter and a divergence parameter, wherein the reward value parameter is used to characterize the reward obtained after taking a target action under a target state, and the divergence parameter is used to characterize the degree of difference between the optimized strategy and the old strategy in the probability of taking the target action. The optimized strategy is the strategy adopted by the large language model after the model parameters are updated, and the old strategy is the strategy adopted by the large language model before the model parameters are updated.
7. The data processing method according to claim 2, characterized in that: The method further comprises: Using a scheduling engine, executing the task script according to the synchronization time and cycle corresponding to the scheduling instruction; Obtaining task execution parameters corresponding to the synchronization task corresponding to the task script, and generating an intelligent report based on the task execution parameters and the user requirement text, wherein the task execution parameters include: at least one of the following: task start time, end time, amount of data processed, task status, and the intelligent report is used to characterize the task execution status and the completion status of the user requirements; The intelligent report is sent to the front-end interactive interface for display.
8. A data processing device, characterized in that: include: A preliminary analysis module is used to obtain a user demand text and determine an initial prompt word corresponding to the user demand text, wherein the user demand text is used to represent the user's demand information for extracting, converting and loading data, and the initial prompt word includes: initial metadata information; A metadata retrieval module, used to determine standard metadata information corresponding to the initial metadata information in the metadata vector library, wherein the standard metadata information is used to represent the pattern information corresponding to the metadata in the data domain; An accurate prompt word construction module is used to adjust the initial prompt word according to the standard metadata information to obtain a target prompt word; The task generation and execution module is used to analyze the target prompt word using a large language model to obtain a task script, and extract, transform and load the data in the data domain by executing the task script, wherein the task script is used to realize the requirements corresponding to the user requirement text.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the data processing method according to any one of claims 1 to 7 when running.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the data processing method according to any one of claims 1 to 7 by running the computer program.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
ETL task intelligent generation method and device
CN121210551A
Timing log analysis task generation system and method, electronic equipment and storage medium
CN121480713A
Document information extraction method for realizing AI Agent by combining RPA, AI and LLM and related product
CN121542445A
Data processing method and device for large language model output distribution, equipment and medium
CN122528965A