Data processing methods, apparatuses, electronic devices, storage media and software products
By using a data processing intelligent agent to analyze business scenarios and obtain matching service interfaces, the problem of automated governance of the data middle platform is solved, enabling rapid iteration and efficient data asset management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
The existing data asset management methods of data middleware require custom system development, have long links, many nodes, and complex manual interaction, making it difficult to achieve automated governance.
By analyzing business scenarios through a data processing intelligent agent, obtaining service interfaces that match the target problem, encapsulating metadata using the model context protocol, achieving atomic decomposition and automated processing, and generating data processing results.
It enables rapid iteration of data platform business scenarios and true automation of data governance, improving the efficiency of data asset governance.
Smart Images

Figure CN122132466A_ABST
Abstract
Description
Technical Field
[0001] In some cases, the field of computer technology is involved, specifically data processing methods, devices, electronic devices, storage media, and program products. Background Technology
[0002] As an enterprise-level data management and service platform, the data platform's data asset management methods are quite complex. Although data asset management methods can be customized to meet specific needs, these customized methods still suffer from problems such as long chains, numerous data asset nodes, and complex manual interactions. Data engineers need to spend a lot of time translating business problems into technical problems and proposing solutions, making it difficult to achieve true automation of data governance. Summary of the Invention
[0003] In view of this, a data processing method, apparatus, electronic device, storage medium, and program product are provided to solve the problem of difficulty in automating the governance of data assets.
[0004] Firstly, a data processing method is provided, including: acquiring the target problem on the data platform; using a data processing agent to parse the business scenario corresponding to the target problem and obtain a service interface matching the target problem, wherein the service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol; using the data processing agent to call the service interface to obtain target metadata information for the target problem; and generating the data processing result of the target problem based on the target metadata information.
[0005] Secondly, a data processing device is provided, comprising: an acquisition module for acquiring a target problem on a data platform; a service interface matching module for using a data processing agent to parse the business scenario corresponding to the target problem and obtain a service interface matching the target problem, wherein the service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on a model context protocol; an interface invocation module for using a data processing agent to invoke the service interface and obtain target metadata information for the target problem; and a data processing module for generating data processing results for the target problem based on the target metadata information.
[0006] Thirdly, an electronic device is provided, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the data processing method described in the first aspect or any corresponding embodiment thereof.
[0007] Fourthly, a computer-readable storage medium is provided, on which computer instructions are stored, for causing a computer to perform the data processing method described in the first aspect or any corresponding embodiment thereof.
[0008] Fifthly, a computer program product is provided, including computer instructions for causing a computer to execute the data processing method described in the first aspect or any corresponding embodiment thereof.
[0009] The aforementioned data processing methods, devices, electronic equipment, storage media, and program products acquire a target problem on the data platform, utilize a data processing agent to analyze the business scenario in which the target problem exists, and obtain a service interface matching the target problem within that business scenario. This service interface is an atomic interface encapsulated based on a model context protocol. The data processing agent calls the service interface to obtain target metadata information for the target problem, enabling the generation of corresponding data processing results based on this metadata. This atomically decomposes the metadata of the data platform, allowing the data processing agent to directly execute automated processing for the target problem. This facilitates rapid iteration of the data platform's business scenarios, promotes true data governance automation, and improves the efficiency of data asset governance within the data platform. Attached Figure Description
[0010] To more clearly illustrate the specific implementation methods or technical solutions in the prior art under certain circumstances, the accompanying drawings used in the description of the specific implementation methods or prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some implementation methods under certain circumstances. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 These are schematic diagrams illustrating application scenarios in various situations;
[0012] Figure 2 This is a flowchart illustrating the first type of data processing method in some situations; Figure 3 This is a schematic diagram of the second flow of data processing methods in some situations; Figure 4 This is a flowchart illustrating the third type of data processing method in some situations; Figure 5 These are structural block diagrams of data processing systems under certain circumstances; Figure 6 These are structural block diagrams of data processing devices under certain conditions; Figure 7These are schematic diagrams of the hardware structure of electronic devices in some scenarios. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages in some situations clearer, the technical solutions in some situations will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments in some situations, not all embodiments. Based on the embodiments shown in some situations, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection.
[0014] It is understood that before using the technical solutions disclosed in the various embodiments under certain circumstances, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0015] For example, upon receiving a user's proactive request, a prompt message can be sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the technical solution, based on the prompt message.
[0016] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0017] It is understood that the above notification and user authorization process are merely illustrative and do not limit the implementation methods in certain situations. Other methods that comply with relevant laws and regulations may also be applied to the implementation methods in certain situations.
[0018] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0019] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature marked "first" or "second" may explicitly or implicitly include one or more of that feature. In some descriptions, "multiple" means two or more, unless otherwise explicitly specified.
[0020] As the capabilities of general-purpose large models continue to upgrade, data platform applications for traditional complex scenarios can be abstracted using multi-agent agents and model context protocols (MCPs) to enable the construction of multi-agent agents for data asset management.
[0021] However, current data asset management methods require custom system development to meet specific needs. Furthermore, data platform systems have long chains, numerous data asset nodes, and complex manual interactions, making automation difficult. This requires data engineers to spend a lot of time transforming business problems into technical problems and proposing solutions.
[0022] Based on this, business problems are described using natural language through interactive processing, and metadata in the data platform system is atomically decomposed to construct data processing agents for relevant problem scenarios. These agents are then deployed to the data processing system to enable natural language interaction and serve engineers in resolving data platform-related issues, thereby improving the iteration and governance efficiency of the data platform.
[0023] As an optional application scenario in some situations, such as Figure 1 As shown, the electronic device 110 has an application 101 with a data processing system installed. The user 130 can interact with the application 101 through the electronic device 110 and / or the access device of the electronic device 110.
[0024] For example, application 101 is an application that provides services related to the data platform. Examples include data asset timeliness governance applications and data asset optimization assistant applications. Figure 1 In the application scenario shown, if application 101 is active, electronic device 110 can display the interface 102 of application 101. Interface 102 may include various pages that application 101 can provide, such as data governance pages, data asset usage pages, data optimization pages, etc.
[0025] In some scenarios, electronic device 110 communicates with server 120 to provide services to application 101. Electronic device 110 can be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, electronic device 110 can also support any type of interface, and server 120 can be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0026] It should be noted that, Figure 1 This is merely an example of an application scenario and is not limited to the scope of protection.
[0027] The following description, in conjunction with the accompanying drawings, will illustrate some scenarios. It should be understood that the actions described relative to electronic device 110 can be performed by application 101 on electronic device 110, or by application 101 in collaboration with its server (e.g., server 120).
[0028] According to certain circumstances, an embodiment of a data processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] In some cases, a data processing method is provided for use in a data processing system deployed on an electronic device. Figure 2 These are flowcharts of data processing methods in some situations, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain the target problem on the data platform.
[0030] The data platform is a data management and service platform that collects, governs, stores and processes data from various business lines in a unified manner, accumulating it into reusable data assets and efficiently supporting the data needs of various business departments in a service-oriented manner.
[0031] The target problem addresses business issues related to data asset management, such as timeliness governance of data asset output, redundant data, and data asset usage instructions. Specifically, the data processing system can access data generated by the data platform through a data service interface. The data platform has corresponding application pages where users can describe the current business problem to be addressed using natural language. Correspondingly, the data processing system can then obtain the target problem from the data platform.
[0032] Step S202: Use the data processing intelligent agent to analyze the business scenario corresponding to the target problem and obtain the service interface that matches the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol.
[0033] The data processing agent is an agent deployed in the data processing system. This data processing agent is an agent based on a large language model. It combines the powerful understanding and generation capabilities of the large language model with certain autonomous decision-making, task execution, and tool invocation capabilities.
[0034] Service interfaces are atomic interfaces obtained by encapsulating metadata using the Model Context Protocol (MCP). Different service interfaces correspond to different MCP services, which in turn support different business scenarios.
[0035] Specifically, the data processing agent understands the target problem, clarifies the business scenario that generates the target problem, and identifies one or more service interfaces that match the target problem in the business scenario.
[0036] Step S203: Use the data processing agent to call the service interface to obtain target metadata information for the target problem.
[0037] Establish a secure, two-way connection between the data storage and data processing agents in the data platform using service interfaces. Target metadata information comprises the metadata that constitutes the target problem, such as data assets, SLA information, task execution information, visualizations, data documents, business documents, and data lineage.
[0038] Specifically, after identifying the service interface that matches the target problem, the data processing agent can call the service interface to access the data storage unit of the data platform and obtain the corresponding target metadata information from the data storage unit.
[0039] Step S204: Based on the target metadata information, generate the data processing results for the target problem.
[0040] The data processing result is the output generated by the data processing agent through automated processing of the target problem. The agent understands the target metadata information to locate the corresponding problem point. Then, based on the relevant problem knowledge base, the agent decides on the processing method for that problem point and generates the corresponding data processing result.
[0041] The data processing method described above obtains the target problem on the data platform, uses a data processing agent to analyze the business scenario in which the target problem exists, and obtains a service interface matching the target problem within that business scenario. This service interface is an atomic interface encapsulated based on a model context protocol. The data processing agent calls the service interface to obtain target metadata information for the target problem, enabling the generation of corresponding data processing results based on this metadata. This atomically decomposes the metadata of the data platform, allowing the data processing agent to directly execute automated processing for the target problem. This facilitates rapid iteration of the data platform's business scenarios and promotes true automation of data governance.
[0042] In some cases, a data processing method is provided that can be used in the aforementioned electronic devices, such as computers and tablets. Figure 3 These are flowcharts of data processing methods in some situations, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain the target problem on the data platform. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be repeated here.
[0043] Step S302: Use the data processing intelligent agent to analyze the business scenario corresponding to the target problem and obtain the service interface that matches the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol.
[0044] Specifically, step S302 includes: Step S3021: Utilize the data processing intelligent agent to analyze the business scenario and obtain multiple first processing sub-processes corresponding to the target problem.
[0045] The first processing sub-process is the sub-process required to solve the target problem. Multiple first processing sub-processes constitute the complete processing flow for solving the target problem. By leveraging the understanding capabilities of the data intelligence agent to analyze the business scenario corresponding to the target problem, and combining the autonomous decision-making capabilities of the data intelligence agent, the processing flow of the target problem is broken down into multiple first processing sub-processes.
[0046] Step S3022: Based on the matching relationship between the first processing sub-process and the service interface, obtain the service interface that matches each first processing sub-process.
[0047] Each first processing sub-process has corresponding processing metadata. Therefore, each first processing sub-process needs to call the corresponding service interface to obtain the corresponding processing metadata. Specifically, all available service interfaces have corresponding data service functions, each first processing sub-process has corresponding data processing requirements, and there is a one-to-one match between data processing requirements and data service functions, that is, there is a matching relationship between the first processing sub-process and the service interface.
[0048] Therefore, by analyzing the data processing requirements of each first processing sub-process and combining the matching relationship between the first processing sub-process and the service interface, the service interface matching each first processing sub-process can be obtained.
[0049] In a specific example, if the processing flow of the target problem is divided into a first processing sub-flow A and a first processing sub-flow B, where the first processing sub-flow A represents the need to query a specified data asset, and the first processing sub-flow B represents the need to obtain upstream and downstream data assets of the specified data asset.
[0050] Among all available service interfaces, there are Service Interface 1, Service Interface 2, and Service Interface 3. Service Interface 1 is used to precisely query information of a specified data asset based on a clear resource type; Service Interface 2 is used to retrieve upstream and downstream data assets of a specified data asset; and Service Interface 3 is used for fuzzy matching of key information of assets to recall data assets.
[0051] Therefore, by matching the data processing requirements with the data service functions, we can find that service interface 1 matches the first processing sub-process A, and service interface 2 matches the second processing sub-process B.
[0052] In some optional scenarios, the metadata of the data platform is encapsulated based on the model context protocol to obtain service interfaces, including: Step a1: Based on the data asset type represented by metadata, encapsulate the data asset query service interface using the model context protocol.
[0053] Data asset type indicates the type of data asset, such as data table, visualization image, document data, etc.; data asset query service interface is a service interface formed by independently encapsulating the metadata corresponding to the data asset type.
[0054] Specifically, based on the metadata representing the data asset type, corresponding API layers, business logic layers, and data access layers are created. The API layer implements API endpoints, the business logic layer implements interface calls, and the data access layer implements data storage access logic, such as database operations. The API layer, business logic layer, and data access layer are encapsulated using a model context protocol to obtain a corresponding data asset query service interface. This interface receives data asset query requests and accurately retrieves data asset information of the specified data asset type from the corresponding database.
[0055] Step a2: Based on the key information of data assets represented by metadata, encapsulate the data asset recall service interface using the model context protocol.
[0056] Key information about data assets includes crucial information characterizing the data asset, such as data source, data type, data timeliness, and data attributes. The data asset retrieval service interface is a service interface formed by independently encapsulating the metadata corresponding to the key information of the data asset.
[0057] Specifically, based on the metadata representing key information of data assets, corresponding API layers, business logic layers, and data access layers are created. The API layer, business logic layer, and data access layer are encapsulated using the model context protocol to obtain the corresponding data asset retrieval service interface. Through this data asset retrieval service interface, data asset query requests are received, and key information of the data assets is fuzzily matched from the corresponding database to retrieve the data assets.
[0058] Step a3: Based on the structured query language of metadata representation, use the model context protocol to encapsulate fields to obtain the service interface.
[0059] The field retrieval service interface is a service interface formed by independently encapsulating the metadata corresponding to the field processing logic.
[0060] Specifically, based on the metadata representing the structured query language, corresponding API, business logic, and data access layers are created. These layers are then encapsulated using a model context protocol to obtain corresponding field retrieval service interfaces. These interfaces receive the trimmed structured query language and, based on field-level processing logic, retrieve the corresponding field definitions from the trimmed structured query language.
[0061] Step a4: Based on the metadata representation of upstream and downstream assets, encapsulate the service interfaces for obtaining upstream and downstream assets using the model context protocol.
[0062] The upstream and downstream asset acquisition service interface is a service interface formed by independently encapsulating the metadata corresponding to the upstream and downstream assets.
[0063] Specifically, based on the metadata representing upstream and downstream assets, corresponding API layers, business logic layers, and data access layers are created. These layers are then encapsulated using a model context protocol to obtain the corresponding upstream and downstream asset retrieval service interfaces. When a specific data asset is specified, the upstream and downstream asset retrieval service interface receives the query request for that data asset and retrieves the corresponding upstream and downstream assets from the relevant database.
[0064] Step a5: Based on the dataset represented by metadata, encapsulate the dataset query service interface using the model context protocol.
[0065] The dataset query service interface is a service interface formed by independently encapsulating the metadata of the dataset dashboard.
[0066] Specifically, based on the metadata representing the dataset, corresponding API, business logic, and data access layers are created. These layers are then combined to encapsulate a dataset query service interface. This interface receives query requests for a specified dataset, retrieves high-popularity dashboards downstream of the specified dataset from the relevant database, and obtains the corresponding dashboard data.
[0067] In the above implementation, independent service interfaces are encapsulated for different metadata to provide atomic service interfaces for data processing agents to call, thereby realizing the atomic decomposition of metadata for the data middle platform. This allows for the acquisition of different combinations of atomic service interfaces according to the business scenario of the target problem, which can then be applied to different business scenarios of the data middle platform.
[0068] In some optional cases, the above method also includes updating the service interface set in response to an extension operation on the service interface.
[0069] Extension operations are operations that expand upon existing service interfaces. Specifically, service interfaces support extensions. When metadata information for a new data type exists, an independent MCP service can be encapsulated for that new data type's metadata information to provide a service interface for data processing agents according to that MCP service. Subsequently, the newly added service interface is updated to the service interface set to extend the existing service interfaces in the service interface set.
[0070] Specifically, when the business scenarios of the data platform are updated, a new MCP service can be encapsulated based on the data required by the updated business scenarios to provide service interfaces adapted to the new business scenarios and geared towards data processing intelligent agents. This newly added service interface is then updated to the service interface set, thus extending the existing service interfaces.
[0071] Of course, the service interface can also be expanded according to other situations, without specific limitations here.
[0072] By supporting the expansion of service interfaces, it is easier to adapt to the iteration of the data platform and meet the business needs of the data platform for rapid iteration.
[0073] Step S303: The data processing agent invokes the service interface to obtain target metadata information related to the target problem. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above; they will not be repeated here.
[0074] Step S304: Based on the target metadata information, generate the data processing result for the target problem. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above, which will not be repeated here.
[0075] The data processing method described above utilizes a data processing agent to analyze the business scenario in which the target problem exists, obtaining multiple first processing sub-processes for solving the target problem and service interfaces matching each first processing sub-process. This achieves automatic decomposition of the processing sub-processes of the target problem, facilitating subsequent retrieval of the corresponding target metadata using atomic service interfaces, and greatly improving the accuracy of the data processing agent in generating data processing results.
[0076] In some cases, a data processing method is provided that can be used in the aforementioned electronic devices, such as computers and tablets. Figure 4 These are flowcharts of data processing methods in some situations, such as... Figure 4 As shown, the process includes the following steps: Step S401: Obtain the target problem on the data platform. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be repeated here.
[0077] Step S402: The data processing agent analyzes the business scenario corresponding to the target problem to obtain a service interface matching the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be repeated here.
[0078] Step S403: Use the data processing agent to call the service interface to obtain target metadata information for the target problem.
[0079] Specifically, step S403 includes: Step S4031: Obtain the metadata type corresponding to the target problem, and the database that matches the metadata type.
[0080] Metadata type refers to the data type of metadata. Different types of metadata use different data storage methods and have corresponding storage databases.
[0081] Specifically, metadata types include asset metadata, visual image metadata, document metadata, and data lineage metadata. For asset metadata, it can be stored in a columnar storage database in the form of data tables, including data assets, SLA information, and task execution information. For visual image metadata, it is stored in the cloud as files using Cloud Storage Objects (TOS). For document metadata, it is vectorized and stored in a vector database. For data lineage metadata, each data asset is used as a node, and the edges between data assets are used as edges, stored in a graph database.
[0082] Step S4032: Use the data processing agent to call the service interface and extract the target metadata from the database.
[0083] As described above, the service interface establishes a secure, two-way connection between data storage and data processing agents. The data processing agent can access a database that matches the metadata type corresponding to the target problem by calling the service interface, and extract the target metadata associated with the target problem from that database.
[0084] Step S404: Based on the target metadata information, generate the data processing results for the target problem.
[0085] Specifically, step S404 includes: Step S4041: In response to the target problem being a data timeliness problem, the target metadata corresponding to the data timeliness problem is understood to obtain the timeliness anomaly node.
[0086] Data timeliness issues indicate discrepancies in the timeliness of data output; timeliness anomaly nodes are data asset nodes with delayed data output.
[0087] Specifically, when the target problem is data timeliness, the data processing agent in the data processing system understands the target metadata such as data asset nodes, output baseline requirements, data lineage, and data output timeliness of the data platform, and locates the timeliness anomaly nodes of data output delay.
[0088] Step S4042: Based on the data link corresponding to the timeliness anomaly node, generate data timeliness governance results using a data processing intelligent agent.
[0089] Based on data lineage, the upstream data asset nodes are located in the data chain of the time-sensitive node. Combining the time-sensitive node and the upstream data asset nodes, a data processing agent is used to diagnose the root cause of the time-sensitive node's anomaly, such as abnormal fluctuations in data timeliness or data task failures.
[0090] Subsequently, a data processing intelligent agent is used to generate a summary report on the data timeliness issue, obtain the corresponding data timeliness governance results, and push the data timeliness governance results to the relevant responsible personnel for optimization of the data timeliness issue.
[0091] Therefore, compared to the manual process of sorting out upstream nodes of the data link, locating problems, summarizing historical performance, and proposing solutions in related technologies, this can be achieved directly through a data processing system for automated processing.
[0092] Specifically, step S404 may further include: Step S4043: In response to the target question being a data usage question, the data processing agent is used to understand the target metadata corresponding to the data usage question and generate a response result that matches the data usage question.
[0093] Data usage issues refer to instructions for using data assets. These issues can include data metric definitions, data asset lookup, and metric data querying.
[0094] Specifically, when the target problem is a data usage problem, the data processing agent in the data processing system understands the target metadata information such as basic data, technical documents, and business documents of the data platform, and uses the interactive page provided by the data processing agent to interact with the data asset usage problem and generate corresponding response results.
[0095] Therefore, compared to manual Q&A in related technologies, this method requires additional effort and may suffer from issues such as delayed responses or the questioner asking the wrong question, resulting in a poor user experience. By automatically generating responses that match the data usage questions through a data processing system, most basic problems can be resolved through pre-emptive interception, improving interaction efficiency.
[0096] Specifically, step S404 may further include: Step S4044: In response to the target problem being a data duplication problem, the data processing agent compares the target metadata information to obtain duplicate data.
[0097] Data duplication issues indicate that data assets contain redundant data; duplicate data indicates that duplicate data exists within data assets.
[0098] Specifically, when the target problem is data duplication, the data processing system understands the target metadata information such as basic data, metadata, technical documents, and code of the data platform to perform comparative analysis of duplicate content on the developed data assets (including table nodes, data indicator definitions in reports, and development logic) and obtain the corresponding duplicate data.
[0099] Step S4045: Optimize the data assets of the data platform based on duplicate data.
[0100] Consolidate duplicate data in data assets, assist relevant data personnel in iteratively optimizing redundant duplicate data, reduce data processing and storage consumption, and avoid redundant development of data assets.
[0101] Therefore, compared to related technologies that require different data personnel to collect data assets within the data platform, manually analyze and label them, and then uniformly integrate them—which necessitates not only individual work but also inter-team communication and collaboration—a data processing system can help reduce the collection, analysis, and communication processes, directly identify duplicate content, and generate reports on redundant content, thus achieving automated pre-emptive diagnosis of data assets within the data platform.
[0102] In some optional cases, the generation of the above-mentioned data processing agent includes: Step b1: Obtain the pending issues from the data platform and the corresponding second processing sub-processes for each pending issue.
[0103] Step b2: Use the pre-trained data processing model to learn each of the second processing sub-processes and generate a data processing agent based on the pre-trained data processing model.
[0104] Among them, the data processing intelligent agent is matched with the business scenario corresponding to the problem to be processed, and is used to execute the second processing sub-process for the problem to be processed, so as to obtain the data processing result corresponding to the problem to be processed.
[0105] The issues to be addressed refer to problems encountered by the data platform during the data governance process. Specifically, data engineers can analyze the business scenario in which the data platform operates to identify the issues to be addressed during data governance. Simultaneously, the processing flow for these issues can be broken down to obtain multiple secondary processing sub-flows for resolving them.
[0106] Accordingly, data engineers can provide the problem to be processed and the second processing sub-process as problem samples to the pre-trained data processing model, so that the pre-trained data processing model can learn and iterate on the processing content of each second processing sub-process, thereby abstracting the problem to be processed into a corresponding data processing intelligent agent. The pre-trained data processing model is an artificial intelligence model based on deep learning using large-scale data, enabling it to possess understanding and task processing capabilities.
[0107] By combining different problems to be addressed, different data processing agents are abstracted and deposited into the service layer of the data processing system to provide different data processing capabilities and obtain corresponding data processing results, thereby improving the iteration efficiency and data governance efficiency of the data platform.
[0108] The data processing method described above utilizes different metadata types with corresponding storage databases. It calls service interfaces that match the business scenario of the target problem, retrieving target metadata from the appropriate database based on the metadata type corresponding to the target problem, thus improving the accuracy of target metadata extraction. Therefore, based on the target metadata corresponding to different target problems, corresponding data processing flows are executed to obtain the corresponding data processing results, achieving automated processing of the target problem and automated generation of data processing results.
[0109] As a specific application example in certain scenarios, an iterative data platform data processing system is provided, such as... Figure 5 As shown, the data processing system includes a data storage module 501, a data service module 502, and an intelligent agent application module 503.
[0110] The data storage module 501 consists of data storage components, including a columnar storage database, a file storage component, a vector database, and a graph database. This module is used to store the metadata that constitutes the data platform, including data assets, SLA information, task execution information, visualizations, data documents, business documents, and data lineage.
[0111] The data service module 502 consists of an API application interface call layer and an MCP service management layer. By querying the metadata information in the data storage module 501, it encapsulates independent MCP services and provides atomic service interfaces for intelligent agents to call.
[0112] The intelligent agent application module 503 consists of several data processing agents, which interact with the MCP service management layer in the data service module 502 and provide corresponding data platform system construction solutions. Specifically, the services provided by the multiple data processing agents include: data asset output timeliness governance solutions; data asset usage instructions; and data asset optimization and slimming assistants.
[0113] Specifically, data engineers interact with the data processing agent in the intelligent agent application module 503 using natural language to clarify the business scenario of the data platform that needs to be addressed. They then invoke the MCP service interface in the data service module 502 to extract metadata information matching the business scenario from the data storage module 501 and return this metadata information to the corresponding data processing agent. The data processing agent can then understand the metadata information, generate corresponding data processing results, and feed these results back to the data engineer, enabling the engineer to optimize relevant issues on the data platform.
[0114] In some cases, a data processing apparatus is also provided for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0115] In some cases, a data processing apparatus is provided, such as Figure 6 As shown, it includes: The acquisition module 601 is used to acquire the target problem on the data platform.
[0116] The service interface matching module 602 is used to analyze the business scenario corresponding to the target problem using the data processing intelligent agent, and obtain the service interface that matches the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol.
[0117] The interface call module 603 is used to call the service interface using the data processing intelligent agent to obtain target metadata information for the target problem.
[0118] The data processing module 604 is used to generate data processing results for the target problem based on the target metadata information.
[0119] In some optional cases, the service interface matching module 602 includes: The parsing unit is used to analyze the business scenario using a data processing intelligent agent to obtain multiple first processing sub-processes corresponding to the target problem.
[0120] The matching unit is used to obtain the service interface that matches each of the first processing sub-processes based on the matching relationship between the first processing sub-processes and the service interface.
[0121] In some optional cases, the above-mentioned device further includes: The encapsulation module is used to encapsulate the metadata retrieval service interface of the data platform based on the model context protocol.
[0122] Specifically, the encapsulation module includes: The type encapsulation unit is used to encapsulate the data asset query service interface based on the data asset type represented by metadata, using the model context protocol.
[0123] The key information encapsulation unit is used to encapsulate key information of data assets based on metadata representation and to encapsulate the data asset retrieval service interface using the model context protocol.
[0124] The field encapsulation unit is used for structured query language based on metadata representation, and uses the model context protocol to encapsulate fields to obtain service interfaces.
[0125] Upstream and downstream encapsulation units are used to encapsulate upstream and downstream assets based on metadata representation and to encapsulate service interfaces for upstream and downstream assets using model context protocols.
[0126] The dataset encapsulation unit is used for datasets based on metadata representations and encapsulates the dataset query service interface using the model context protocol.
[0127] In some optional cases, the above-mentioned device further includes: The interface extension module is used to update the service interface set in response to extension operations on the service interface.
[0128] In some alternative implementations, the interface invocation module 603 includes: The database matching unit is used to obtain the metadata type corresponding to the target problem, as well as the database that matches the metadata type.
[0129] The data extraction unit is used to extract target metadata from the database by calling the service interface using the data processing agent.
[0130] In some optional cases, the data processing module 604 includes: The abnormal node determination unit is used to understand the target metadata corresponding to the data timeliness problem in response to the target problem being a data timeliness problem, and to obtain the timeliness abnormal node.
[0131] The data governance unit is used to generate data timeliness governance results based on the data links corresponding to the timeliness anomaly nodes using data processing intelligent agents.
[0132] In some optional cases, the data processing module 604 may also include: The data usage processing unit is used to respond to the target question as a data usage question, and utilizes a data processing agent to understand the target metadata corresponding to the data usage question, and generate a response result that matches the data usage question.
[0133] In some optional cases, the data processing module 604 may also include: The duplicate data identification unit is used to identify duplicate data by comparing the target metadata information with the target metadata information in response to the target problem being a data duplication problem.
[0134] The optimization unit is used to optimize the data assets of the data platform based on duplicate data.
[0135] In some optional cases, the above-mentioned device further includes: The agent generation module is used to generate data processing agents.
[0136] Specifically, the agent generation module includes: The pending issues acquisition unit is used to acquire pending issues from the data platform and the corresponding second processing sub-processes for each pending issue.
[0137] The learning unit is used to learn each second processing sub-process using a pre-trained data processing model to generate a data processing agent based on the pre-trained data processing model. Among them, the data processing intelligent agent is matched with the business scenario corresponding to the problem to be processed, and is used to execute the second processing sub-process for the problem to be processed, so as to obtain the data processing result corresponding to the problem to be processed.
[0138] This data processing device can execute the data processing methods provided in the above-described scenarios, possessing the corresponding functional modules and beneficial effects. By acquiring the target problem on the data platform, the data processing agent analyzes the business scenario in which the target problem exists, obtaining a service interface matching the target problem within that business scenario. This service interface is an atomic interface obtained through encapsulation based on the model context protocol. The data processing agent calls the service interface to obtain target metadata information for the target problem, enabling the generation of corresponding data processing results based on this metadata. Thus, by atomically decomposing the metadata of the data platform, the data processing agent can directly execute automated processing for the target problem, facilitating rapid iteration of the data platform's business scenarios and promoting true automation of data governance.
[0139] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0140] Figure 7 This is a schematic diagram of the structure of an electronic device provided in certain situations.
[0141] The following is a detailed reference. Figure 7 This diagram illustrates the structure of an electronic device suitable for implementing certain scenarios. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 701, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 702 or a program loaded from memory 708 into random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0142] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0143] In particular, under certain circumstances, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, some cases may include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from memory 708, or installed from ROM 702. When the computer program is executed by processor 701, it performs the functions defined in the data processing methods under certain circumstances.
[0144] Figure 7 The electronic devices shown are merely examples and should not impose any limitations on their functionality or scope of use.
[0145] In some cases, a computer-readable storage medium is also provided, in which the methods described above can be implemented in hardware or firmware, or implemented as recordable on the storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded via a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the data processing methods shown in the above embodiments.
[0146] In some cases, certain components can be applied as computer program products, such as computer program instructions. When executed by a computer, these instructions, through the operation of the computer, can invoke or provide methods and / or technical solutions according to the aforementioned situations. Those skilled in the art will understand that the forms in which computer program instructions exist in computer-readable media include, but are not limited to, source files, executable files, and installation package files. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0147] Although some embodiments have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the description, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A data processing method, comprising: The target problem on the data acquisition platform; The data processing agent is used to analyze the business scenario corresponding to the target problem and obtain a service interface that matches the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol. The data processing agent invokes the service interface to obtain target metadata information for the target problem; Based on the target metadata information, the data processing result of the target problem is generated.
2. The method according to claim 1, wherein the step of using a data processing agent to parse the business scenario corresponding to the target problem and obtain a service interface matching the target problem includes: The data processing agent is used to analyze the business scenario to obtain multiple first processing sub-processes corresponding to the target problem; Based on the matching relationship between the first processing sub-process and the service interface, the service interface that matches each of the first processing sub-processes is obtained.
3. The method according to claim 1 or 2, wherein the service interface is obtained by encapsulating the metadata of the data platform based on the model context protocol, comprising: Based on the data asset type represented by the metadata, the data asset query service interface is encapsulated using the model context protocol. And / or, Based on the key information of the data assets represented by the metadata, the data asset recall service interface is encapsulated using the model context protocol. And / or, Based on the structured query language represented by the metadata, the service interface is obtained by encapsulating fields using the model context protocol; and / or, Based on the upstream and downstream assets represented by the metadata, the upstream and downstream asset acquisition service interface is encapsulated using the model context protocol; And / or, Based on the dataset represented by the metadata, the dataset query service interface is encapsulated using the model context protocol.
4. The method according to claim 3, further comprising: In response to an extension operation for the service interface, update the service interface set.
5. The method according to claim 1, wherein the step of using the data processing agent to call the service interface to obtain target metadata information for the target problem includes: Obtain the metadata type corresponding to the target problem, and the database that matches the metadata type; The data processing agent invokes the service interface to extract the target metadata from the database.
6. The method according to claim 1, wherein generating the data processing result of the target problem based on the target metadata information includes: In response to the target problem being a data timeliness issue, the target metadata corresponding to the data timeliness issue is understood to obtain timeliness anomaly nodes; Based on the data link corresponding to the timeliness anomaly node, the data processing agent generates data timeliness governance results.
7. The method according to claim 1, wherein generating the data processing result of the target problem based on the target metadata information includes: In response to the target question being a data usage question, the data processing agent is used to understand the target metadata corresponding to the data usage question and generate a response that matches the data usage question.
8. The method according to claim 1, wherein generating the data processing result of the target problem based on the target metadata information includes: In response to the target problem being a data duplication problem, the data processing agent compares the target metadata information to obtain duplicate data; Based on the duplicate data, the data assets of the data platform are optimized.
9. The method according to any one of claims 1 to 8, wherein the generation of the data processing agent comprises: Obtain the pending issues of the data platform and the corresponding second processing sub-processes for each pending issue; The pre-trained data processing model is used to learn each of the second processing sub-processes to generate the data processing agent based on the pre-trained data processing model; The data processing agent is matched with the business scenario corresponding to the problem to be processed, and is used to execute the second processing sub-process for the problem to be processed, so as to obtain the data processing result corresponding to the problem to be processed.
10. A data processing apparatus, comprising: The acquisition module is used to acquire the target problem on the data platform; The service interface matching module is used to analyze the business scenario corresponding to the target problem using a data processing intelligent agent, and obtain a service interface that matches the target problem. The service interface is an atomic interface obtained by encapsulating the metadata of the data platform based on the model context protocol. The interface invocation module is used to invoke the service interface using the data processing intelligent agent to obtain target metadata information for the target problem. The data processing module is used to generate data processing results for the target problem based on the target metadata information.
11. An electronic device, comprising: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the data processing method of any one of claims 1 to 9.
12. A computer-readable storage medium storing computer instructions for causing a computer to perform the data processing method of any one of claims 1 to 9.
13. A computer program product comprising computer instructions for causing a computer to perform the data processing method of any one of claims 1 to 9.