Data processing method, apparatus, device, medium, and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECH (BEIJING) CO LTD
- Filing Date
- 2025-02-07
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]目前数据工程涉及的数据处理流程对人工依赖较高,例如,数据工程师需要编写大量代码,以实现数据工程中的数据采集、清洗、整合、存储与分析等功能,并且还需要熟练使用这些复杂的功能,从而导致数据处理效率和准确度较低的问题
[0026]通过本申请提供的技术方案,使得在数据工程领域中,基于智能体可以实现自动化数据处理,由于智能体具备高度的自主性,能够在不依赖人类干预的情况下进行决策和行动,因此,基于这种自主性使得智能体能够快速且准确地输出数据工程相关响应,从而可以提高数据处理效率和准确度。
Smart Images

Figure CN122528935A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular to a data processing method, apparatus, device, medium, and program product. Background Technology
[0002] Data engineering is a comprehensive process that encompasses multiple key stages, including data collection, cleaning, integration, storage, and analysis. Its aim is to transform complex and diverse data into data information with business value.
[0003] Currently, data engineering involves data processing workflows that rely heavily on manual labor. For example, data engineers need to write a lot of code to implement data collection, cleaning, integration, storage and analysis functions in data engineering, and they also need to be proficient in using these complex functions, which leads to problems of low data processing efficiency and accuracy. Summary of the Invention
[0004] This application provides a data processing method, apparatus, device, medium, and program product, which can improve data processing efficiency and accuracy.
[0005] In a first aspect, embodiments of this application provide a data processing method applied to an intelligent agent. The method includes: acquiring a current data engineering-related request; generating a target data engineering-related request based on the current data engineering-related request; retrieving the target data engineering-related request from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request, and generating prompt information based on the associated knowledge; decomposing the target data engineering-related request to obtain multiple sub-tasks; and calling a corresponding data engineering tool for each sub-task and prompt information to obtain a data engineering-related response.
[0006] Secondly, embodiments of this application provide a data processing apparatus, including: an acquisition module, a generation module, a retrieval module, a decomposition module, and a calling module; wherein, the acquisition module is used to acquire a current data engineering-related request; the generation module is used to generate a target data engineering-related request based on the current data engineering-related request; the retrieval module is used to retrieve the target data engineering-related request from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request; the generation module is also used to generate prompt information based on the associated knowledge; the decomposition module is used to decompose the target data engineering-related request to obtain multiple sub-tasks; and the calling module is used to call the corresponding data engineering tool for each sub-task and prompt information to obtain a data engineering-related response.
[0007] In some implementations, the generation module is specifically used to: determine historical data engineering related requests associated with the current data engineering related request; and generate a target data engineering related request based on the current data engineering related request and the historical data engineering related requests associated with the current data engineering related request.
[0008] In some implementations, the generation module is specifically used to: obtain at least one historical data engineering-related request within a target time window; wherein the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset time as the start time, or the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset duration; and based on the current data engineering-related request and at least one historical data engineering-related request, determine the historical data engineering-related request related to the current data engineering-related request.
[0009] In some implementations, the generation module is specifically used to: calculate the similarity between the current data engineering related request and at least one historical data engineering related request; and take the historical data engineering related request with the highest similarity among the at least one historical data engineering related requests as the historical data engineering related request related to the current data engineering related request.
[0010] In some implementations, the generation module is specifically used to: search in the knowledge graph for at least one historical data engineering related request that is related to the current data engineering related request.
[0011] In some implementations, the device further includes a rewriting module for rewriting the current data engineering request to conform to the input requirements of the agent's large language model before the generation module generates the target data engineering request based on the current data engineering request.
[0012] In some implementations, the generation module is specifically used to: combine the agent's role information and associated knowledge into prompt information; wherein, the agent's role information is data engineering expert role information.
[0013] In some implementations, the generation module is specifically used to: compose the agent's role information and associated knowledge into prompt information according to the prompt information template.
[0014] In some implementations, the acquisition module is specifically used to: obtain current data engineering-related requests through the agent's display interface.
[0015] In some implementations, the device further includes an interception module for intercepting the current data engineering request if the current data engineering request contains illegal data.
[0016] In some implementations, the device further includes a control module, which, after the calling module invokes the corresponding data engineering tool for each subtask and prompt, and obtains the data engineering-related response, controls the display interface of the intelligent agent to display the data engineering-related response.
[0017] In some implementations, the control module is specifically used to control the display interface of the intelligent agent to display the data engineering-related response if the data engineering-related response does not include illegal data.
[0018] In some implementations, the control module is also configured to: if the data engineering-related response includes illegal data, then the control agent's display interface does not display the data engineering-related response, or the control agent's display interface does not display illegal data.
[0019] In some implementations, the device further includes: a selection module for selecting a graphics card compatible with the target data engineering request before the calling module invokes the corresponding data engineering tool for each subtask and prompt to obtain a data engineering-related response; correspondingly, the calling module is specifically used to: invoke the corresponding data engineering tool for each subtask and prompt through the graphics card compatible with the target data engineering request to obtain a data engineering-related response.
[0020] In some implementations, the external knowledge base includes at least one of the following: database metadata, data lineage, data preview information, and definitions of business terms.
[0021] In some implementations, data engineering tools include at least one of the following: data acquisition tools, data cleaning tools, data transformation tools, data integration tools, database lookup tools, and code generation tools.
[0022] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the methods as described in the first aspect or its various implementations.
[0023] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the methods described in the first aspect or its various implementations.
[0024] Fifthly, embodiments of this application provide a computer program product including computer program instructions that cause a computer to perform the methods as described in the first aspect or its various implementations.
[0025] Sixthly, embodiments of this application provide a computer program that causes a computer to perform the methods as described in the first aspect or its various implementations.
[0026] The technical solution provided in this application enables automated data processing based on intelligent agents in the field of data engineering. Because intelligent agents have a high degree of autonomy, they can make decisions and take actions without human intervention. Therefore, based on this autonomy, intelligent agents can quickly and accurately output data engineering-related responses, thereby improving data processing efficiency and accuracy. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A schematic diagram illustrating the working principle of an intelligent agent;
[0029] Figure 2 This is a schematic diagram of a system architecture according to an embodiment of this application;
[0030] Figure 3 A flowchart illustrating a data processing method provided in an embodiment of this application;
[0031] Figure 4 A partial schematic diagram of the knowledge graph provided in the embodiments of this application;
[0032] Figure 5 A schematic diagram of a data processing apparatus 500 provided in an embodiment of this application;
[0033] Figure 6 A schematic diagram of a data processing system 600 provided in an embodiment of this application;
[0034] Figure 7 This is a schematic block diagram of the electronic device 700 provided in the embodiments of this application. Detailed Implementation
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0037] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0038] Before introducing the technical solution of this application, the relevant knowledge of the technical solution of this application will be explained below:
[0039] I. Data engineering is a comprehensive process encompassing multiple key stages such as data collection, cleaning, integration, storage, and analysis. Its aim is to transform complex and diverse data into valuable business information. Data engineering primarily involves the following product functions:
[0040] (1) Data acquisition function
[0041] Data source connectivity: The product should support connectivity to multiple data sources, including relational databases, non-relational databases, file storage systems, and application programming interfaces (APIs), in order to obtain data from multiple sources.
[0042] Real-time data acquisition: To meet the demand for real-time data, the product should provide real-time data acquisition capabilities, which can capture changes in the data stream and import them into the data processing system in a timely manner.
[0043] (2) Data cleaning function
[0044] Data validation: The product should be able to automatically validate the integrity, accuracy and consistency of data, and identify and mark missing values, outliers or duplicate data.
[0045] Data Conversion: Provides functions such as data format conversion, data type conversion, and data standardization to ensure data consistency and comparability.
[0046] Data deduplication: It can automatically identify and delete duplicate data to avoid bias in subsequent analysis.
[0047] (3) Data integration function
[0048] Data merging: Supports merging data from different data sources to form a unified data view.
[0049] Data mapping: Provides data field mapping functionality to ensure that the same or similar fields in different data sources can be correctly mapped.
[0050] (4) Data storage and management functions
[0051] Distributed storage: Utilizing technologies such as distributed file systems, data warehouses, or data lakes to achieve efficient storage and management of large-scale data.
[0052] Data Indexing and Querying: Provides efficient data indexing and querying functions for quick data retrieval and analysis.
[0053] Data backup and recovery: Ensure data security and availability by providing data backup and recovery functions to prevent data loss or damage.
[0054] (5) Data analysis and visualization functions
[0055] Data analysis tools: Built-in data analysis algorithms and models, supporting advanced analysis functions such as data mining and machine learning.
[0056] Data Visualization: Offers a rich set of data visualization tools and chart types to help users intuitively understand and analyze data.
[0057] (6) Other auxiliary functions
[0058] Scheduling and Detection: Provides task scheduling and detection functions to ensure the automation and reliability of data processing workflows.
[0059] Security and Access Management: Ensures data security and privacy protection by providing functions such as user authentication, access management, and data encryption.
[0060] II. An intelligent agent refers to an entity or program capable of perceiving its environment, making decisions, and executing actions to achieve specific goals. Among them, an intelligent agent uses a Large Language Model (LLM) as its brain; it can not only understand, perceive, plan, remember, and use tools, but also automatically perform complex tasks, demonstrating the ability to think and act independently. It achieves specific goals or solves specific problems through continuous self-learning and environmental adaptation.
[0061] Figure 1 A schematic diagram illustrating the working principle of an intelligent agent, such as... Figure 1 As shown, the agent can perform task planning, including reflexes (i.e., self-criticism), self-reflection, chained thinking, and breaking down tasks into multiple sub-tasks. When planning tasks or engaging in self-criticism, the agent can rely on its memory functions, including short-term and long-term memory. Short-term memory is used for all contextual learning, such as cue engineering, while long-term memory allows the agent to retain and recall information for extended periods, typically achieved through external vector storage and fast retrieval. After completing task planning, the agent can invoke tools to execute the task. These tools can include calendars, calculators, code interpreters, search engines, and more. The search tool is primarily used to supplement additional knowledge to better perform the task.
[0062] The technical problems to be solved, the inventive concept and the system architecture of the embodiments of this application will be described below:
[0063] As mentioned above, the current data processing workflows involved in data engineering rely heavily on manual labor. For example, data engineers need to write a lot of code to implement functions such as data collection, cleaning, integration, storage and analysis in data engineering, and they also need to be proficient in using these complex functions, which leads to problems with low data processing efficiency and accuracy.
[0064] To address the aforementioned technical problems, this application proposes an automated data processing approach based on intelligent agents in the field of data engineering. Because intelligent agents possess a high degree of autonomy, they can make decisions and take actions without human intervention. Therefore, this autonomy enables intelligent agents to quickly and accurately output data engineering-related responses, thereby improving data processing efficiency and accuracy.
[0065] In some possible implementations, the system architecture of embodiments of this application is as follows: Figure 2 As shown.
[0066] Figure 2 This is a schematic diagram of a system architecture involved in an embodiment of this application, such as... Figure 2As shown, the system architecture involves an intelligent agent 210, an external database 220, and a data engineering tool 230. The intelligent agent 210 is wirelessly connected to both the external database 220 and the data engineering tool 230.
[0067] In some implementations, the agent 210 can be understood as an agent in software or hardware form.
[0068] In this context, the software-based intelligent agent can be understood as a large language model, or as an intelligent agent built upon a large language model. This can be achieved by adding callback functions, which can call functions from data engineering tools, but are not limited to these. Alternatively, it can be combined with a Reasoning and Acting (ReAct) framework. This ReAct framework is used to coordinate the interaction between the LLM and external information, enabling the large language model to construct a complete series of actions based on logical reasoning (Reson) to achieve the desired goal.
[0069] Among them, the intelligent agent in hardware form is actually an intelligent agent that combines software and hardware. That is, the intelligent agent has a certain physical form and includes intelligent agents in software form. For example, the intelligent agent can be a terminal device or a server, but is not limited to these.
[0070] In some possible implementations, the terminal device can be a desktop computer, laptop computer, tablet computer, smartphone, smartwatch, virtual reality (VR) device, augmented reality (AR) device, smart robot, vehicle terminal, etc., but is not limited to these.
[0071] In some implementations, the server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0072] In some implementations, agent 210 can acquire current data engineering-related requests; generate target data engineering-related requests based on current data engineering-related requests; search for target data engineering-related requests in external knowledge bases related to data engineering to obtain associated knowledge of target data engineering-related requests, and generate prompt information based on associated knowledge; decompose target data engineering-related requests to obtain multiple sub-tasks; and for each sub-task and prompt information, call the corresponding data engineering tool to obtain data engineering-related responses.
[0073] In some possible implementations, the external knowledge base includes at least one of the following, but is not limited to: database metadata, data lineage, data preview information, and definitions of business terms.
[0074] In some possible implementations, data engineering tools include at least one of the following, but are not limited to: data acquisition tools, data cleaning tools, data transformation tools, data integration tools, database lookup tools, and code generation tools.
[0075] It should be noted that, Figure 2 This is merely a schematic diagram of a system architecture provided in this application embodiment; the system architecture involved in this application embodiment is not limited to... Figure 2 The system architecture shown, for example, Figure 2 The system architecture shown includes only one external database 220, but in reality, it can include multiple external databases 220.
[0076] The technical solution of this application will be described in detail below:
[0077] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The method is applied to an intelligent agent, which can be in software or hardware form. For details regarding these two types of intelligent agents, please refer to the above description; further details are omitted here. Figure 3 As shown, the method may include:
[0078] S310: Retrieve requests related to the current data engineering process;
[0079] It should be understood that the current data engineering-related requests refer to the data engineering-related requests currently received by the intelligent agent. These requests are also referred to as queries, questions, instructions, etc., and the embodiments of this application do not limit them.
[0080] In the embodiments of this application, "related to data engineering" can refer to at least one aspect of data engineering, such as at least one aspect of data acquisition, data cleaning, data transformation, data integration, and data analysis.
[0081] The following is an example illustrating current data engineering-related requests:
[0082] For example, the current data engineering-related request is used to request the merging of the user activity table for video platform A this week and the user activity table for video platform A last week.
[0083] For example, the current data engineering-related request is used to check whether video platform A's user activity table for this week contains a user identifier (ID) field.
[0084] In some possible implementations, the agent obtains current data engineering-related requests, including obtaining current data engineering-related requests through the agent's display interface.
[0085] It should be understood that users can input current data engineering-related requests on the agent's display interface so that the agent can obtain the current data engineering-related requests.
[0086] It should be understood that, in this implementation, the current data engineering-related requests are in text format.
[0087] In some implementations, the agent acquires current data engineering-related requests, including by acquiring current data engineering-related requests through the agent's voice acquisition unit.
[0088] In some implementations, the voice acquisition unit may include a microphone, audio circuitry, etc.
[0089] It should be understood that, in this feasible approach, the current data engineering-related request is a voice request.
[0090] It should be understood that the embodiments of this application do not limit the form or acquisition method of current data engineering-related requests.
[0091] In some implementations, after receiving a current data engineering-related request, the agent can detect whether the request includes illegal data. If the request includes illegal data, the agent can intercept it; that is, if the request includes illegal data, the agent will not execute steps S320 to S350. Conversely, if the request does not include illegal data, the agent will execute steps S320 to S350 normally.
[0092] In some implementations, the agent detects whether the current data engineering-related request includes illegal data, including: the agent can segment the current data engineering-related request into words and match each segment with data in an illegal database, wherein the illegal database stores a number of illegal data. If at least one segment in the current data engineering-related request successfully matches at least one illegal data in the illegal database, then it is determined that the current data engineering-related request includes illegal data; if all segments in the current data engineering-related request fail to match any illegal data in the illegal database, then it is determined that the current data engineering-related request does not include illegal data.
[0093] In some implementations, the agent detects whether the current data engineering-related request includes illegal data, including: the agent can segment the current data engineering-related request into words, select keywords, and match the keywords with data in an illegal database. If at least one keyword in the current data engineering-related request successfully matches at least one piece of illegal data in the illegal database, then the current data engineering-related request is determined to include illegal data; if all keywords in the current data engineering-related request fail to match any illegal data in the illegal database, then the current data engineering-related request is determined not to include illegal data.
[0094] It should be understood that intelligent agents can ensure data security by intercepting current data engineering-related requests that carry illegal data. In addition, intelligent agents can only perform illegality detection on keywords in current data engineering-related requests, while not performing illegality detection on other word segments, thereby improving detection efficiency and thus improving data processing efficiency while ensuring data security.
[0095] S320: Generate the target data engineering request based on the current data engineering related requests;
[0096] It should be understood that the target data engineering related request refers to the data engineering related request corresponding to the data engineering related response to be generated, which is ultimately determined by the agent. This target data engineering related request can be the current data engineering related request, or it can be a combination of the current data engineering related request and other data engineering related requests. The generation method of the target data engineering related request will be explained below:
[0097] It should be understood that, as mentioned above, the intelligent agent possesses a memory function. Based on this, the intelligent agent can determine historical data engineering-related requests associated with the current data engineering-related request, and generate a target data engineering-related request based on the current data engineering-related request and the historical data engineering-related requests associated with it. In other words, in this implementable method, the target data engineering-related request is a data engineering-related request formed by combining the current data engineering-related request and the historical data engineering-related requests associated with it.
[0098] For example, the current data engineering-related request is used to request the merging of the user activity table of video platform A this week and the user activity table of video platform A last week. The historical data engineering-related request related to the current data engineering-related request is used to request checking whether the user activity table of video platform A this week contains a user ID field. The agent combines the two and obtains the target data engineering-related request, which is used to request the merging of the user activity table of video platform A this week and the user activity table of video platform A last week according to user ID.
[0099] In some implementations, the agent determines historical data engineering-related requests related to the current data engineering-related request, including: the agent acquiring at least one historical data engineering-related request within a target time window; and based on the current data engineering-related request and at least one historical data engineering-related request, determining historical data engineering-related requests related to the current data engineering-related request.
[0100] The target time window is a time window that ends at the time the current data engineering-related request is received and begins at a preset time; alternatively, the target time window is a time window that ends at the time the current data engineering-related request is received and has a preset duration. This embodiment of the application does not impose any restrictions on the preset time or preset duration.
[0101] For example, assuming the current data engineering-related request is received at 15:00 on January 16, 2025, and the target time window is a preset duration of 10 minutes, then the agent can obtain historical data engineering-related requests from 14:50 on January 16, 2025 to 15:00 on January 16, 2025.
[0102] For example, assuming the current data engineering-related request is received at 15:00 on January 16, 2025, and the target time window starts at a preset time of 14:50 on January 16, 2025, then the agent can obtain historical data engineering-related requests from 14:50 on January 16, 2025 to 15:00 on January 16, 2025.
[0103] In some implementations, the agent determines historical data engineering related requests based on the current data engineering related request and at least one historical data engineering related request, including: the agent calculates the similarity between the current data engineering related request and at least one historical data engineering related request; and selects the historical data engineering related request with the highest similarity among the at least one historical data engineering related requests as the historical data engineering related request related to the current data engineering related request.
[0104] For example, suppose the current data engineering-related request is used to request the merging of the user activity table of video platform A this week and the user activity table of video platform A last week. Here, historical data engineering-related request 001 is used to request checking whether the user activity table of video platform A this week contains a user ID field, and historical data engineering-related request 002 is used to request checking whether the user activity table of video platform A last week contains a user ID field. Suppose the agent determines that the similarity between the current data engineering-related request and historical data engineering-related request 001 is 90%, and the similarity between the current data engineering-related request and historical data engineering-related request 002 is 80%. Then the agent can regard historical data engineering-related request 001 as a historical data engineering-related request related to the current data engineering-related request.
[0105] In some implementations, the agent calculates the similarity between the current data engineering-related request and at least one historical data engineering-related request, including: the agent determines the embedding corresponding to the current data engineering-related request and the embedding corresponding to each of the at least one historical data engineering-related request, and calculates the similarity between the embedding corresponding to the current data engineering-related request and the embedding corresponding to each historical data engineering-related request, as the similarity between the current data engineering-related request and the historical data engineering-related request.
[0106] In the embodiments of this application, for any two embeddings, any one of the following information can be used as the similarity between the two embeddings, but is not limited to: Pearson Correlation Coefficient, Cosine Similarity, Euclidean Distance, Manhattan Distance, and Dot Product Similarity.
[0107] It should be understood that the similarity calculation method for any two data engineering related requests will not be elaborated on below.
[0108] In some possible implementations, the agent determines historical data engineering related requests related to the current data engineering related request based on the current data engineering related request and at least one historical data engineering related request, including: the agent calculates the similarity between the current data engineering related request and at least one historical data engineering related request; and selects the historical data engineering related requests with the top K similarity scores among the at least one historical data engineering related requests as historical data engineering related requests related to the current data engineering related request, where K is an integer greater than 1.
[0109] For example, suppose the current data engineering-related request is used to merge the user activity table of video platform A for this week and the user activity table of video platform A for last week. Historical data engineering-related request 001 is used to check if the user activity table of video platform A for this week contains a user ID field, historical data engineering-related request 002 is used to check if the user activity table of video platform A for last week contains a user ID field, and historical data engineering-related request 003 is used to find the playback record table of video platform A for last week. Assuming K=2, the agent determines that the similarity between the current data engineering-related request and historical data engineering-related request 001 is 90%, the similarity between the current data engineering-related request and historical data engineering-related request 002 is 80%, and the similarity between the current data engineering-related request and historical data engineering-related request 003 is 20%. Then, the agent can consider historical data engineering-related requests 001 and 002 as related historical data engineering-related requests to the current data engineering-related request.
[0110] In some implementations, the agent determines historical data engineering-related requests related to the current data engineering-related request based on the current data engineering-related request and at least one historical data engineering-related request. This includes: the agent calculating the similarity between the current data engineering-related request and at least one historical data engineering-related request; and identifying historical data engineering-related requests with a similarity greater than a preset similarity among the at least one historical data engineering-related requests as historical data engineering-related requests related to the current data engineering request. In this embodiment, the preset similarity is not limited.
[0111] For example, suppose the current data engineering request requests to merge the user activity table of video platform A for this week and the user activity table of video platform A for last week. Historical data engineering request 001 requests to check if the user activity table of video platform A for this week contains a user ID field, historical data engineering request 002 requests to check if the user activity table of video platform A for last week contains a user ID field, and historical data engineering request 002 requests to find the playback record table of video platform A for last week. Assuming the preset similarity is 60%, the agent determines that the similarity between the current data engineering request and historical data engineering request 001 is... The similarity between the current data engineering-related request and the historical data engineering-related request 002 is 80%, and the similarity between the current data engineering-related request and the historical data engineering-related request 003 is 20%. Since the similarity between the current data engineering-related request and the historical data engineering-related request 001 is 90%, which is greater than the preset similarity of 60%, and the similarity between the current data engineering-related request and the historical data engineering-related request 002 is 80%, which is greater than the preset similarity of 60%, the agent can consider the historical data engineering-related requests 001 and 002 as historical data engineering-related requests related to the current data engineering-related request.
[0112] In some implementations, the agent determines historical data engineering related requests associated with the current data engineering related request based on the current data engineering related request and at least one historical data engineering related request, including: the agent searching in the knowledge graph for at least one historical data engineering related request that is associated with the current data engineering related request.
[0113] In some implementations, each time an agent receives a data engineering-related request, it can determine the relationship between that request and historical data engineering-related requests, and draw a knowledge graph based on these relationships. This knowledge graph includes the relationships between several data engineering-related requests.
[0114] In some implementations, the agent searches for historical data engineering related requests in the knowledge graph that are related to the current data engineering related request, including: the agent determining historical data engineering related requests that have a direct connection relationship with the current data engineering related request as historical data engineering related requests related to the current data engineering related request.
[0115] For example, Figure 4 This is a partial schematic diagram of the knowledge graph provided in the embodiments of this application, such as... Figure 4As shown, since historical data engineering related requests 001 and 002 are directly connected to the current data engineering related request, while historical data engineering related request 003 is indirectly connected to the current data engineering related request, the agent can determine historical data engineering related requests 001 and 002 as historical data engineering related requests related to the current data engineering related request.
[0116] In some possible implementations, the agent searches for historical data engineering related requests in the knowledge graph that are related to the current data engineering related request, including: the agent identifying historical data engineering related requests that have a direct connection relationship with the current data engineering related request and historical data engineering related requests that have an indirect connection relationship with the current data engineering related request as historical data engineering related requests related to the current data engineering related request.
[0117] For example, such as Figure 4 As shown, since historical data engineering related requests 001 and 002 are directly connected to the current data engineering related request, while historical data engineering related request 003 is indirectly connected to the current data engineering related request, the agent can identify historical data engineering related requests 001, 002, and 003 as historical data engineering related requests related to the current data engineering related request.
[0118] It should be understood that the embodiments of this application do not limit the method for determining the historical data engineering related requests related to the current data engineering related request based on the current data engineering related request and at least one historical data engineering related request.
[0119] In some implementations, the agent determines historical data engineering-related requests related to the current data engineering-related request, including: the agent acquiring at least one historical data engineering-related request from the same user within a target time window that is related to the current data engineering-related request; and based on the current data engineering-related request and at least one historical data engineering-related request, determining historical data engineering-related requests related to the current data engineering-related request.
[0120] It should be understood that the explanation of the target time window can be found above, and will not be repeated in the embodiments of this application.
[0121] For example, suppose the current data engineering-related request is received at 15:00 on January 16, 2025, and the target time window is a preset duration of 10 minutes. Within this time window, there are a total of 2,000 historical data engineering-related requests. However, historical data engineering-related requests belonging to the same user as the current data engineering-related request include: historical data engineering-related request 001 and historical data engineering-related request 002. Then, the agent can obtain these two historical data engineering-related requests.
[0122] For example, suppose the current data engineering-related request is received at 15:00 on January 16, 2025, and the target time window starts at a preset time of 14:50 on January 16, 2025. This time window includes a total of 2,000 historical data engineering-related requests. However, historical data engineering-related requests belonging to the same user as the current data engineering-related request include: historical data engineering-related request 001 and historical data engineering-related request 002. Then the agent can obtain these two historical data engineering-related requests.
[0123] It should be understood that, by setting a target time window, the embodiments of this application can obtain historical data engineering-related requests within a recent period, ensuring the relevance of these requests to the current data engineering-related requests. This results in more accurate target data engineering-related requests and improves the accuracy of data processing. Furthermore, by obtaining at least one historical data engineering-related request from the same user within the target time window that is related to the current data engineering-related request, the relevance of these requests to the current data engineering-related request can be further improved, leading to more accurate target data engineering-related requests and thus improving the accuracy of data processing.
[0124] It should be understood that, regarding how the embodiments of this application determine the historical data engineering related requests related to the current data engineering related requests based on the current data engineering related requests and at least one historical data engineering related request, please refer to the above, and the embodiments of this application will not repeat it here.
[0125] It should be understood that the embodiments of this application do not limit how to determine historical data engineering related requests associated with current data engineering related requests.
[0126] It should be understood that, since the intelligent agent performs data processing based on a large language model, the intelligent agent generates the target data engineering related request based on the current data engineering related request and the historical data engineering related requests related to the current data engineering related request. This includes: integrating the current data engineering related request and the historical data engineering related requests related to the current data engineering related request through a large language model to obtain the target data engineering related request.
[0127] In some possible implementations, the agent generates a target data engineering request based on the current data engineering request, including: the agent uses the current data engineering request as the target data engineering request.
[0128] For example, if the current data engineering related request is used to request the merging of the user activity table of video platform A for this week and the user activity table of video platform A for last week, the agent can send this current data engineering related request as the target data engineering related request.
[0129] In some implementations, before the agent generates the target data engineering request based on the current data engineering request, the process further includes: the agent rewriting the current data engineering request to conform to the input requirements of the agent's large language model.
[0130] It should be understood that, since the agent processes data based on a large language model, the agent rewrites the current data engineering-related requests to conform to the input requirements of the agent's large language model. This includes: the agent rewrites the current data engineering-related requests through the large language model to conform to the input requirements of the large language model.
[0131] It should be understood that the ability of a large language model to rewrite current data engineering-related requests primarily relies on its training process. This training process enables the large language model to possess the capability to rewrite current data engineering-related requests:
[0132] In some implementations, the large language model can be trained using training samples. Each training sample can include: the original data engineering-related request and the actual, rewritten data engineering-related request, where the actual, rewritten data engineering-related request can serve as the sample label. The training device can employ supervised training; for example, it can input the original data engineering-related request into the large language model and output the predicted rewritten data engineering-related request. Furthermore, the training device can calculate a loss based on all the actual and predicted rewritten data engineering-related requests included in the training samples, and adjust the parameters of the large language model based on this loss until the training iterations reach a preset number or the loss reaches its minimum value, at which point training stops.
[0133] In some implementations, the training device may use any of the following loss functions when training the large language model, but is not limited to: L1 loss function, mean squared error (MSE) loss function, cross-entropy loss function, etc.
[0134] In some possible implementations, the agent rewrites the current data engineering-related request using a large language model to conform to the input requirements of the large language model. This includes: the agent rewrites the structure of the current data engineering-related request using a large language model so that the rewritten structure conforms to the input requirements of the large language model.
[0135] For example, a current data engineering request that asks "How do I merge the user activity table for video platform A this week and the user activity table for video platform A last week?" could be rewritten as "What are the methods for merging the user activity table for video platform A this week and the user activity table for video platform A last week?"
[0136] In some implementations, the agent rewrites the current data engineering-related request using a large language model to conform to the input requirements of the large language model. This includes: the agent rewrites the current data engineering-related request using the large language model in conjunction with background information to conform to the input requirements of the large language model.
[0137] In some implementations, the background information includes at least one of the following, but is not limited to: the time of receipt of the current data engineering-related request and the user information corresponding to the request.
[0138] For example, the current data engineering request is "Merge the user activity table of video platform A this week and the user activity table of video platform A last week", which can be rewritten as "Please merge the user activity table of video platform A this week and the user activity table of video platform A last week up to the current time".
[0139] For example, the current data engineering request is "Merge the user activity table of video platform A this week and the user activity table of video platform A last week", which can be rewritten as "Merge the user activity table of video platform A this week and the user activity table of video platform A last week based on user ID".
[0140] In some possible implementations, the agent rewrites the current data engineering-related request using a large language model to conform to the input requirements of the large language model. This includes: the agent rewriting the structure of the current data engineering-related request using the large language model, and rewriting the current data engineering-related request using the large language model in conjunction with background information to conform to the input requirements of the large language model.
[0141] It should be understood that the explanation of the background information can be found above, and the embodiments of this application will not repeat it here.
[0142] For example, a current data engineering request that asks "How do I merge the user activity table for video platform A this week and the user activity table for video platform A last week?" could be rewritten as "Please merge the user activity table for video platform A this week and the user activity table for video platform A last week up to the current time".
[0143] It should be understood that the embodiments of this application rewrite the current data engineering related requests so that the rewritten current data engineering related requests are more in line with the input requirements of the large language model, thereby facilitating the large language model's understanding of the current data engineering related requests and improving the accuracy of data processing.
[0144] It should be understood that the embodiments of this application do not limit the rewriting methods of current data engineering-related requests.
[0145] S330: Search for the target data engineering-related request in an external knowledge base related to data engineering to obtain the associated knowledge of the target data engineering-related request, and generate prompt information based on the associated knowledge;
[0146] It should be understood that, in order to enhance the processing capabilities of the large language model, this application embodiment employs Retrieval-Augmented Generation (RAG) technology. RAG is an artificial intelligence technology that combines information retrieval techniques with a large language model. This technology retrieves relevant knowledge about the query statement (i.e., the target data engineering-related request here) from an external knowledge base and inputs it as a prompt message into the large language model to enhance its processing capabilities.
[0147] The RAG workflow typically includes three steps: retrieval, enhancement, and generation. First, the query (i.e., the target data engineering-related request) is retrieved from an external knowledge base to obtain relevant knowledge. Second, this relevant knowledge is used as hints for the large language model to enhance its understanding and ability to respond to the query. Finally, the large language model generates a response that meets the user's needs.
[0148] In some possible implementations, the external knowledge base includes at least one of the following, but is not limited to: database metadata, data lineage, data preview information, and definitions of business terms.
[0149] It should be understood that metadata is information about data structure, data description, and data management. It describes the attributes, sources, organization, relational constraints, and indexes of the data, enabling the data to be accurately understood, organized, and used.
[0150] It should be understood that metadata includes structured metadata, descriptive metadata, and administrative metadata. Structured metadata describes the structure and organization of data, such as table definitions, field types, and constraints. Descriptive metadata provides information about the data content, such as the data's source, creation time, and author. Administrative metadata involves information about data management and usage, such as data access permissions, backup and recovery strategies, etc.
[0151] It should be understood that data lineage refers to the natural, almost kinship-like relationship that forms between data throughout their entire lifecycle—from generation, processing, manipulation, fusion, and flow to their eventual demise. Simply put, it describes the upstream and downstream relationships between data, i.e., where the data comes from and where it goes. Data lineage involves not only the physical flow of data but also its logical relationships and transformation processes.
[0152] In some implementation methods, the data preview information includes at least one of the following: field information of the data table, data range, data type, data format, data quality, etc. Among these, the field information includes at least one of the following: field name, field type, field description, etc.
[0153] In some implementations, the data stored in the external database includes at least one of the following categories: structured data and unstructured data. Structured data refers to data with a defined row and column structure, which can be predefined by a data model or presented in the form of a two-dimensional table. This data can be stored in a database and is easily processed by computer programs. Structured data typically follows a fixed format or schema, such as tabular data in a relational database. Unstructured data, on the other hand, refers to information without a predefined data model or organized in a predefined manner. Unstructured data is usually text-based but may also include dates, numbers, and factual data.
[0154] In some possible implementations, the target data engineering-related request is retrieved from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request. This includes: the agent can calculate the similarity between the embedding of the target data engineering-related request and the embedding of each data in the external knowledge base, and take the data with the highest similarity to the target data engineering-related request as the associated knowledge of the target data engineering-related request.
[0155] For example, suppose the current data engineering request is to merge the user activity table of video platform A for this week and the user activity table of video platform A for last week. The external database stores the following data: user activity refers to the frequency and extent to which users use products or services within a certain period of time (such as daily, weekly, monthly, etc.). Assuming that this data has the highest similarity to the current data engineering request, then this data can be used as the associated knowledge of the current data engineering request.
[0156] In some possible implementations, the target data engineering-related request is retrieved from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request. This includes: the agent can calculate the similarity between the embedding of the target data engineering-related request and the embedding of each data in the external knowledge base, and use the top K data in the external database as associated knowledge of the target data engineering-related request.
[0157] For example, suppose the current data engineering request is to merge the user activity tables of video platform A for this week and last week. A piece of data 01 stored in the external database is as follows: User activity refers to the frequency and extent to which users use a product or service within a certain period (e.g., daily, weekly, monthly). Another piece of data 02 is as follows: The user activity table includes user ID, username, login time, login duration, number of logins, login channel, etc. Assuming that these two pieces of data have the highest and second-highest similarity to the current data engineering request, then these two pieces of data can be used as the association knowledge for the current data engineering request.
[0158] In some possible implementations, the target data engineering-related request is retrieved from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request. This includes: the agent can calculate the similarity between the embedding of the target data engineering-related request and the embedding of each data in the external knowledge base, and use the data in the external database with a similarity greater than a preset similarity as associated knowledge of the target data engineering-related request.
[0159] For example, suppose the current data engineering request is to merge the user activity table of video platform A for this week and the user activity table of video platform A for last week. The external database stores the following data: User activity refers to the frequency and extent to which users use a product or service within a certain period of time (such as daily, weekly, monthly, etc.). Assuming the preset similarity is 70%, and the similarity between this data and the current data engineering request is 80%, then this data can be used as the associated knowledge of the current data engineering request.
[0160] It should be understood that the embodiments of this application do not limit the way in which the associated knowledge of the target data engineering related requests is obtained.
[0161] In some implementations, the agent generates prompts based on associated knowledge, including: the agent uses the associated knowledge as the prompt.
[0162] In some possible implementations, the agent generates prompts based on associated knowledge, including: the agent combines its role information and associated knowledge to form the prompts; wherein the agent's role information is data engineering expert role information.
[0163] In some possible implementations, the agent combines its role information and associated knowledge into a prompt message, including: the agent combines its role information and associated knowledge into a prompt message according to a prompt message template.
[0164] For example, the prompt message template is as follows:
[0165] As a data engineering expert, please generate data engineering-related responses for data engineering-related requests based on the following related knowledge: User activity refers to the frequency and extent to which users use products or services within a certain period of time (such as daily, weekly, monthly, etc.).
[0166] It should be understood that the format of the prompt message template is not limited in the embodiments of this application.
[0167] S340: Decompose the target data engineering-related requests into multiple sub-tasks;
[0168] For example, a data engineering request to merge the user activity tables of video platform A for this week and last week can be understood as a task, which the agent can break down into the following sub-tasks:
[0169] Find the user activity table for video platform A this week;
[0170] Find the user activity table for video platform A for last week;
[0171] The user activity tables for video platform A this week and last week were merged by user ID.
[0172] S350: For each subtask and prompt message, call the corresponding data engineering tool to obtain the data engineering related response.
[0173] In some possible implementations, data engineering tools include at least one of the following, but are not limited to: data acquisition tools, data cleaning tools, data transformation tools, data integration tools, database lookup tools, and code generation tools.
[0174] It should be understood that the intelligent agent can analyze each subtask to determine the data engineering tools used to perform that subtask. For example, for the subtask of finding the user activity table of video platform A this week, a database table lookup tool can be called. For the subtask of finding the user activity table of video platform A last week, a database table lookup tool can also be called. For the subtask of merging the user activity table of video platform A this week and the user activity table of video platform A last week by user ID, a data integration tool can be called.
[0175] In some implementations, the multiple subtasks are executed sequentially, wherein when executing a later subtask, the agent can combine the execution results of the earlier subtasks, and the execution result of the last subtask is the data engineering-related response.
[0176] For example, a data engineering request to merge the user activity tables of video platform A for this week and last week can be understood as a task, which the agent can break down into the following sub-tasks:
[0177] Subtask 1: Find the user activity table for video platform A this week;
[0178] Subtask 2: Find the user activity table for video platform A last week;
[0179] Subtask 3: Merge the user activity table for this week and the user activity table for last week of video platform A by user ID.
[0180] The agent needs to execute subtask 1 and subtask 2 first, and then subtask 3. When the agent executes subtask 3, it can merge the user activity table of video platform A this week and the user activity table of video platform A last week according to user ID to obtain the merged user activity table.
[0181] In some implementations, before the agent calls the corresponding data engineering tool for each subtask and prompt to obtain the data engineering-related response, it also includes: selecting a graphics card that is compatible with the target data engineering-related request; accordingly, the agent can call the corresponding data engineering tool for each subtask and prompt to obtain the data engineering-related response through the graphics card that is compatible with the target data engineering-related request.
[0182] In some implementations, the agent can estimate the storage space required for the target data engineering-related request and select a graphics card that matches the storage space required for the target data engineering-related request.
[0183] In some implementations, the agent can estimate the running speed of the target data engineering-related requests and select a graphics card that matches the running speed of the target data engineering-related requests.
[0184] It should be understood that by running the target data engineering-related requests on the adapted video memory, this embodiment of the application can ensure the output speed of the large language model on the one hand, and ensure the reasonable allocation of video memory and the utilization rate of resources on the other hand.
[0185] In some implementations, after the agent calls the corresponding data engineering tool for each subtask and prompt, and obtains the data engineering-related response, it also includes: the agent controlling the agent's display interface to display the data engineering-related response.
[0186] In some implementations, if the data engineering-related response does not include illegal data, the agent controls the agent's display interface to show the data engineering-related response.
[0187] In some implementations, if the data engineering-related response includes illegal data, the agent's control agent's display interface does not show the data engineering-related response, or the control agent's display interface does not show illegal data.
[0188] In some implementations, the agent detects whether the current data engineering-related response includes illegal data, including: the agent can segment the current data engineering-related response into words and match each segment with data in an illegal database, wherein the illegal database stores a number of illegal data. If at least one segment in the current data engineering-related response successfully matches at least one illegal data in the illegal database, then it is determined that the current data engineering-related response includes illegal data; if all segments in the current data engineering-related response fail to match any illegal data in the illegal database, then it is determined that the current data engineering-related response does not include illegal data.
[0189] In some implementations, the agent detects whether the current data engineering-related response includes illegal data, including: the agent can segment the current data engineering-related response into words, select keywords, and match the keywords with data in an illegal database. If at least one keyword in the current data engineering-related response successfully matches at least one piece of illegal data in the illegal database, then the current data engineering-related response is determined to include illegal data; if all keywords in the current data engineering-related response fail to match any illegal data in the illegal database, then the current data engineering-related response is determined not to include illegal data.
[0190] It should be understood that intelligent agents can ensure data security by intercepting current data engineering-related responses that carry illegal data. In addition, intelligent agents can only perform illegality detection on keywords in current data engineering-related responses, while not performing illegality detection on other word segments, thereby improving detection efficiency and thus improving data processing efficiency while ensuring data security.
[0191] This application provides a data processing method, including: acquiring a current data engineering-related request; generating a target data engineering-related request based on the current data engineering-related request; retrieving the target data engineering-related request from an external knowledge base related to data engineering to obtain associated knowledge of the target data engineering-related request, and generating prompt information based on the associated knowledge; decomposing the target data engineering-related request into multiple sub-tasks; and calling the corresponding data engineering tool for each sub-task and prompt information to obtain a data engineering-related response. Because the intelligent agent possesses a high degree of autonomy, it can make decisions and take actions without relying on human intervention. Therefore, this autonomy enables the intelligent agent to quickly and accurately output data engineering-related responses, thereby improving data processing efficiency and accuracy.
[0192] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application.
[0193] It should also be understood that, in the various method embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0194] The methods provided in the embodiments of this application have been described above. The data processing apparatus and system provided in the embodiments of this application will be described below.
[0195] Figure 5 A schematic diagram of a data processing apparatus 500 provided in an embodiment of this application is shown below. Figure 5 As shown, the device 500 includes: an acquisition module 510, a generation module 520, a retrieval module 530, a disassembly module 540, and a calling module 550; wherein, the acquisition module 510 is used to acquire the current data engineering-related request; the generation module 520 is used to generate a target data engineering-related request based on the current data engineering-related request; the retrieval module 530 is used to retrieve the target data engineering-related request from an external knowledge base related to data engineering to obtain the associated knowledge of the target data engineering-related request; the generation module 520 is also used to generate prompt information based on the associated knowledge; the disassembly module 540 is used to disassemble the target data engineering-related request to obtain multiple sub-tasks; the calling module 550 is used to call the corresponding data engineering tool for each sub-task and prompt information to obtain the data engineering-related response.
[0196] In some implementations, the generation module 520 is specifically used to: determine historical data engineering related requests associated with the current data engineering related request; and generate a target data engineering related request based on the current data engineering related request and the historical data engineering related requests associated with the current data engineering related request.
[0197] In some implementations, the generation module 520 is specifically used to: acquire at least one historical data engineering-related request within a target time window; wherein the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset time as the start time, or the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset duration; and based on the current data engineering-related request and at least one historical data engineering-related request, determine the historical data engineering-related request related to the current data engineering-related request.
[0198] In some implementations, the generation module 520 is specifically used to: calculate the similarity between the current data engineering related request and at least one historical data engineering related request; and take the historical data engineering related request with the highest similarity among the at least one historical data engineering related requests as the historical data engineering related request related to the current data engineering related request.
[0199] In some implementations, the generation module 520 is specifically used to: search in the knowledge graph for at least one historical data engineering related request that is related to the current data engineering related request.
[0200] In some implementations, the device 500 further includes a rewriting module 560, which rewrites the current data engineering request to conform to the input requirements of the agent's large language model before the generation module 520 generates the target data engineering request based on the current data engineering request.
[0201] In some implementations, the generation module 520 is specifically used to: combine the agent's role information and associated knowledge into prompt information; wherein, the agent's role information is data engineering expert role information.
[0202] In some implementations, the generation module 520 is specifically used to: compose the agent's role information and associated knowledge into prompt information according to the prompt information template.
[0203] In some implementations, the acquisition module 510 is specifically used to: acquire current data engineering-related requests through the intelligent agent's display interface.
[0204] In some implementations, the device 500 further includes an interception module 570, configured to intercept the current data engineering request if the current data engineering request contains illegal data.
[0205] In some implementations, the device 500 further includes a control module 580, which, after the calling module 550 calls the corresponding data engineering tool for each subtask and prompt information and obtains the data engineering-related response, controls the display interface of the intelligent agent to display the data engineering-related response.
[0206] In some implementations, the control module 580 is specifically used to control the display interface of the intelligent agent to display the data engineering-related response if the data engineering-related response does not include illegal data.
[0207] In some implementations, the control module 580 is also configured to: if the data engineering-related response includes illegal data, then the control agent's display interface does not display the data engineering-related response, or the control agent's display interface does not display illegal data.
[0208] In some implementations, the device 500 further includes: a selection module 590, used to select a graphics card adapted to the target data engineering request before the calling module 550 calls the corresponding data engineering tool for each subtask and prompt message to obtain a data engineering-related response; correspondingly, the calling module 550 is specifically used to: call the corresponding data engineering tool for each subtask and prompt message through the graphics card adapted to the target data engineering request to obtain a data engineering-related response.
[0209] In some implementations, the external knowledge base includes at least one of the following: database metadata, data lineage, data preview information, and definitions of business terms.
[0210] In some implementations, data engineering tools include at least one of the following: data acquisition tools, data cleaning tools, data transformation tools, data integration tools, database lookup tools, and code generation tools.
[0211] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 5 The device 500 shown can perform Figure 3 The corresponding method embodiments, and the foregoing and other operations and / or functions of each module in device 500 are respectively implemented to achieve Figure 3 For the sake of brevity, the corresponding processes in each method are not described in detail here.
[0212] The apparatus 500 of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0213] Figure 6 This is a schematic diagram of a data processing system 600 provided in an embodiment of this application, as shown below. Figure 6As shown, the data processing system 600 includes: a user-agent interaction unit 610, a prompt word engineering unit 620, a language large model unit 630, an external knowledge base 640, and a data engineering tool 650. The user-agent interaction unit 610, the prompt word engineering unit 620, and the language large model unit 630 can be understood as units within the agent.
[0214] In some implementations, the user-agent interaction unit 610 is used to acquire the current data engineering-related request; the prompt word engineering unit 620 is used to generate a target data engineering-related request based on the current data engineering-related request; to search for the target data engineering-related request in an external knowledge base related to data engineering to obtain the associated knowledge of the target data engineering-related request, and to generate prompt information based on the associated knowledge; to decompose the target data engineering-related request into multiple sub-tasks; and the language large model unit 630 is used to call the corresponding data engineering tools for each sub-task and prompt information to obtain the data engineering-related response.
[0215] In some implementations, the prompt word engineering unit 620 is specifically used to: determine historical data engineering related requests associated with the current data engineering related request; and generate a target data engineering related request based on the current data engineering related request and the historical data engineering related requests associated with the current data engineering related request.
[0216] In some implementations, the prompt word engineering unit 620 is specifically used to: acquire at least one historical data engineering-related request within a target time window; wherein the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset time as the start time, or the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset duration; and based on the current data engineering-related request and at least one historical data engineering-related request, determine the historical data engineering-related request related to the current data engineering-related request.
[0217] In some implementations, the prompt word engineering unit 620 is specifically used to: calculate the similarity between the current data engineering related request and at least one historical data engineering related request; and take the historical data engineering related request with the highest similarity among the at least one historical data engineering related requests as the historical data engineering related request related to the current data engineering related request.
[0218] In some implementations, the prompt word engineering unit 620 is specifically used to: search in the knowledge graph for at least one historical data engineering related request that is related to the current data engineering related request.
[0219] In some implementations, the prompt word engineering unit 620 is also used to: rewrite the current data engineering-related request to conform to the input requirements of the agent's large language model.
[0220] In some implementations, the prompt word engineering unit 620 is specifically used to: compose prompt information from the agent's role information and associated knowledge; wherein, the agent's role information is data engineering expert role information.
[0221] In some implementations, the prompt word engineering unit 620 is specifically used to: compose prompt information from the agent's role information and associated knowledge according to the prompt information template.
[0222] In some implementations, the user-agent interaction unit 610 is specifically used to: obtain current data engineering-related requests through the agent's display interface.
[0223] In some implementations, the user-agent interaction unit 610 is also configured to: intercept the current data engineering-related request if the current data engineering-related request includes illegal data.
[0224] In some implementations, the user-agent interaction unit 610 is also used to: control the display interface of the agent to display data engineering-related responses.
[0225] In some implementations, the user-agent interaction unit 610 is specifically used to: control the agent's display interface to display the data engineering-related response if the data engineering-related response does not include illegal data.
[0226] In some implementations, the user-agent interaction unit 610 is further configured to: if the data engineering-related response includes illegal data, control the display interface of the agent to not display the data engineering-related response, or control the display interface of the agent to not display illegal data.
[0227] In some implementations, the language large model unit 630 is also used to select a graphics card that is compatible with the target data engineering request; correspondingly, the language large model unit 630 is specifically used to: through the graphics card that is compatible with the target data engineering request, for each subtask and prompt message, call the corresponding data engineering tool to obtain the data engineering related response.
[0228] In some implementations, the external knowledge base includes at least one of the following: database metadata, data lineage, data preview information, and definitions of business terms.
[0229] In some implementations, data engineering tools include at least one of the following: data acquisition tools, data cleaning tools, data transformation tools, data integration tools, database lookup tools, and code generation tools.
[0230] It should be understood that the system embodiments and method embodiments can correspond to each other, and similar descriptions can be found in the method embodiments. To avoid repetition, further details are omitted here. Specifically, Figure 6 The system 600 shown can execute Figure 3 The corresponding method embodiments, and the foregoing and other operations and / or functions of each module in system 600 are respectively for implementing Figure 3 For the sake of brevity, the corresponding processes in each method are not described in detail here.
[0231] Figure 7 This is a schematic block diagram of the electronic device 700 provided in an embodiment of this application. Figure 7 As shown, the electronic device 700 may include:
[0232] The system includes a memory 710 and a processor 720. The memory 710 stores a computer program 730 and transfers the computer program 730 to the processor 720. In other words, the processor 720 can retrieve and run the computer program 730 from the memory 710 to implement the methods described in the embodiments of this application.
[0233] For example, the processor 720 can be used to execute the steps in the above method according to the instructions in the computer program 730.
[0234] In some embodiments of this application, the processor 720 may include, but is not limited to:
[0235] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0236] In some embodiments of this application, the memory 710 includes, but is not limited to:
[0237] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0238] In some embodiments of this application, the computer program 730 may be divided into one or more modules, which are stored in the memory 710 and executed by the processor 720 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 730 in the electronic device.
[0239] like Figure 7 As shown, the electronic device 700 may further include:
[0240] Transceiver 740, which can be connected to processor 720 or memory 710.
[0241] The processor 720 can control the transceiver 740 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 740 may include a transmitter and a receiver. The transceiver 740 may further include antennas, and the number of antennas may be one or more.
[0242] It should be understood that the various components in the electronic device 700 are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0243] According to one aspect of this application, a computer storage medium is provided that stores a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0244] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in the above-described method embodiments.
[0245] In other words, when implemented using software, it can be implemented wholly or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0246] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0247] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0248] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0249] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method, characterized in that, The method is applied to an intelligent agent, and the method includes: Retrieve current data engineering related requests; Based on the current data engineering related requests, generate the target data engineering related requests; The target data engineering-related request is retrieved from an external knowledge base related to data engineering to obtain the associated knowledge of the target data engineering-related request, and prompt information is generated based on the associated knowledge. The target data engineering-related requests are broken down into multiple sub-tasks; For each subtask and the prompt message, the corresponding data engineering tool is invoked to obtain the data engineering-related response.
2. The method according to claim 1, characterized in that, The step of generating a target data engineering related request based on the current data engineering related request includes: Identify historical data engineering related requests that are associated with the current data engineering related request; The target data engineering request is generated based on the current data engineering request and the historical data engineering requests related to the current data engineering request.
3. The method according to claim 2, characterized in that, The determination of historical data engineering related requests associated with the current data engineering related request includes: Obtain at least one historical data engineering-related request within a target time window; wherein, the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset time as the start time, or, the target time window is a time window with the receiving time of the current data engineering-related request as the end time and a preset duration as the size. Based on the current data engineering related request and the at least one historical data engineering related request, determine the historical data engineering related request related to the current data engineering related request.
4. The method according to claim 3, characterized in that, The step of determining the historical data engineering related requests associated with the current data engineering related requests based on the current data engineering related requests and the at least one historical data engineering related requests includes: Calculate the similarity between the current data engineering-related request and the at least one historical data engineering-related request; The historical data engineering related request with the highest similarity among the at least one historical data engineering related requests is taken as the historical data engineering related request related to the current data engineering related request.
5. The method according to claim 3, characterized in that, The step of determining the historical data engineering related requests associated with the current data engineering related requests based on the current data engineering related requests and the at least one historical data engineering related requests includes: Search the knowledge graph for historical data engineering related requests that are related to the current data engineering related request among the at least one historical data engineering related requests.
6. The method according to any one of claims 2-5, wherein before generating the target data engineering related request based on the current data engineering related request, it further includes: The current data engineering-related requests are rewritten to meet the input requirements of the agent's large language model.
7. The method according to any one of claims 1-5, characterized in that, The generation of prompt information based on the associated knowledge includes: The prompt information is composed of the role information of the intelligent agent and the associated knowledge. The role information of the intelligent agent is data engineering expert role information.
8. The method according to claim 7, characterized in that, The step of assembling the prompt information from the agent's role information and the associated knowledge includes: The prompt message is composed of the agent's role information and the associated knowledge according to the prompt message template.
9. The method according to any one of claims 1-5, characterized in that, The request to obtain current data engineering related information includes: The current data engineering-related requests are obtained through the intelligent agent's display interface; The method further includes: If the current data engineering related request includes illegal data, then the current data engineering related request is blocked.
10. The method according to any one of claims 1-5, characterized in that, After invoking the corresponding data engineering tool for each subtask and the prompt information to obtain the data engineering-related response, the process further includes: If the data engineering-related response does not include illegal data, then the display interface of the intelligent agent is controlled to display the data engineering-related response; If the data engineering-related response includes illegal data, then the display interface of the intelligent agent is controlled not to display the data engineering-related response, or the display interface of the intelligent agent is controlled not to display the illegal data.
11. The method according to any one of claims 1-5, characterized in that, Before invoking the corresponding data engineering tool for each subtask and the prompt information to obtain the data engineering-related response, the process also includes: Select a graphics card that is compatible with the request related to the target data project; For each subtask and the prompt information, the corresponding data engineering tool is invoked to obtain a data engineering-related response, including: By using a graphics card adapted to the target data engineering request, for each subtask and the prompt information, the corresponding data engineering tool is invoked to obtain the data engineering related response.
12. A data processing apparatus, characterized in that, include: The module includes an acquisition module, a generation module, a retrieval module, a decomposition module, and a calling module. The acquisition module is used to acquire requests related to the current data engineering. The generation module is used to generate a target data engineering related request based on the current data engineering related request; The retrieval module is used to retrieve the target data engineering-related request from an external knowledge base related to data engineering, so as to obtain the associated knowledge of the target data engineering-related request; The generation module is also used to generate prompt information based on the associated knowledge; The decomposition module is used to decompose the target data engineering-related requests into multiple sub-tasks; The calling module is used to call the corresponding data engineering tool for each subtask and the prompt information to obtain data engineering related responses.
13. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 1 to 11.
15. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in any one of claims 1 to 11.