Task processing method, data supplement method, and task processing system
By using a unified cross-task solution framework based on large language models, the problem of low data processing efficiency in data lakes is solved, achieving efficient task processing and cost reduction, adapting to various data tasks in data lakes, and improving model processing performance.
Patent Information
- Application Number
- CN202310495569.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2043-04-28
AI Technical Summary
When dealing with large datasets of unknown structure, content, and quality in a data lake, user processing efficiency is extremely low, the effectiveness and performance of machine learning models are limited, and solutions based on machine learning models usually require designing corresponding models for specific scenarios and tasks, resulting in high inference costs.
We adopt a unified cross-task solution framework based on a large language model. By leveraging the language understanding and cross-task capabilities of the task processing model, we can adapt to various data tasks in the data lake, reduce repetitive process development, fully consider the data organization structure of the data lake, and perform data retrieval to extract relevant information and filter out invalid information.
It improves task processing efficiency, reduces computational load and usage costs, enhances model processing performance, and reduces computational consumption of invalid information.
Smart Images

Figure CN116680245B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of computer technology, and particularly relate to a task processing method. One or more embodiments of the present specification also relate to a data supplement method, a task processing system, a task processing device, a data supplement device, a computing device, a computer-readable storage medium, and a computer program. BACKGROUND
[0002] With the development of computer technology, the data generated by enterprises and individual users is growing explosively. A data lake is a storage system that stores data in raw format, which can store data as is without prior structured processing of the data. How to efficiently process a large amount of data in the data lake has gradually become a research focus.
[0003] Currently, a user can search and browse data sets in the data lake through a console. However, in the face of a large number of data sets with unknown structure, content, and data quality in the data lake, the user's self-searching and browsing data is extremely inefficient. Therefore, there is an urgent need for an efficient data processing solution. SUMMARY
[0004] In view of this, embodiments of the present specification provide a task processing method. One or more embodiments of the present specification also relate to a data supplement method, a task processing system, a task processing device, a data supplement device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.
[0005] According to a first aspect of embodiments of the present specification, a task processing method is provided, comprising:
[0006] In response to a task processing request for a data lake, task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of a target task;
[0007] The task attribute information and the task processing data are input into a retrieval unit in a task processing model to obtain reference data of the target task;
[0008] The task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request;
[0009] The target processing data is input into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0010] According to a second aspect of embodiments of the present specification, a task processing method is provided, applied to a cloud-side device, comprising:
[0011] In response to a task processing request for the data lake sent by the end-side device, task processing data corresponding to the task processing request is acquired, wherein the task processing request carries task attribute information of a target task;
[0012] The task attribute information and the task processing data are input into a retrieval unit in the task processing model to obtain reference data of the target task;
[0013] The task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request;
[0014] The target processing data is input into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request;
[0015] The task processing result corresponding to the task processing request is sent to the end-side device.
[0016] According to a third aspect of an embodiment of the present specification, a data supplement method is provided, comprising:
[0017] In response to a data supplement request for the data lake, task processing data corresponding to the data supplement request is acquired, wherein the data supplement request carries task attribute information of a target data supplement task;
[0018] The task attribute information and the task processing data are input into a retrieval unit in the task processing model to obtain reference data of the target data supplement task;
[0019] The task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the data supplement request;
[0020] The target processing data is input into a processing unit in the task processing model to obtain a data supplement result corresponding to the data supplement request.
[0021] According to a fourth aspect of an embodiment of the present specification, a task processing system is provided, comprising a data processing component and a task processing component;
[0022] The data processing component is configured to, in response to a task processing request for the data lake, acquire task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task; input the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target task; and send the reference data of the target task to the task processing component;
[0023] The task processing component is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; and input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0024] According to a fifth aspect of the embodiments of the present specification, a task processing apparatus is provided, comprising:
[0025] The first obtaining module is configured to, in response to a task processing request for a data lake, obtain task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task.
[0026] The first input module is configured to input the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target task.
[0027] The second input module is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request.
[0028] The third input module is configured to input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0029] According to a sixth aspect of the embodiments of the present specification, a task processing apparatus is provided, applied to a cloud-side device, comprising:
[0030] The second obtaining module is configured to, in response to a task processing request for a data lake sent by an edge-side device, obtain task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task.
[0031] The fourth input module is configured to input the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target task.
[0032] The fifth input module is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request.
[0033] The sixth input module is configured to input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0034] The sending module is configured to send the task processing result corresponding to the task processing request to the edge-side device.
[0035] According to a seventh aspect of the embodiments of the present specification, a data supplementing apparatus is provided, comprising:
[0036] a third obtaining module configured to obtain task processing data corresponding to the data supplementing request in response to the data supplementing request for the data lake, wherein the data supplementing request carries task attribute information of a target data supplementing task;
[0037] a seventh input module configured to input the task attribute information and the task processing data into a retrieving unit in the task processing model to obtain reference data of the target data supplementing task;
[0038] an eighth input module configured to input the task attribute information and the reference data into a converting unit in the task processing model to obtain target processing data corresponding to the data supplementing request;
[0039] a ninth input module configured to input the target processing data into a processing unit in the task processing model to obtain a data supplementing result corresponding to the data supplementing request.
[0040] According to an eighth aspect of the embodiments of the present specification, a computing device is provided, comprising:
[0041] a memory and a processor;
[0042] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the method provided in the first aspect or the second aspect or the third aspect.
[0043] According to a ninth aspect of the embodiments of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions, when executed by a processor, implement the steps of the method provided in the first aspect or the second aspect or the third aspect.
[0044] According to a tenth aspect of the embodiments of the present specification, a computer program is provided, and when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect or the second aspect or the third aspect.
[0045] The task processing method provided by one embodiment of the present specification comprises the following steps: in response to a task processing request for a data lake, obtaining task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task; inputting the task attribute information and the task processing data into a retrieval unit in a task processing model to obtain reference data of the target task; inputting the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; and inputting the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request. By utilizing the language understanding capability and cross-task capability of the task processing model, different tasks in the data lake can be better adapted, repeated process development can be reduced, and in the task processing process, various data organization structures under the data lake are fully considered, data retrieval is performed by utilizing the task processing model, on the one hand, information related to the target task is extracted, the model processing performance is improved, the task processing efficiency is further improved, on the other hand, invalid information is filtered, the consumption of the calculation amount is reduced, and the use cost of the task processing model is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 is an architecture diagram of a task processing system provided by one embodiment of the present specification;
[0047] Figure 2 is an architecture diagram of another task processing system provided by one embodiment of the present specification;
[0048] Figure 3 is a flowchart of a task processing method provided by one embodiment of the present specification;
[0049] Figure 4 is a flowchart of another task processing method provided by one embodiment of the present specification;
[0050] Figure 5 is a flowchart of a data supplement method provided by one embodiment of the present specification;
[0051] Figure 6 is a framework diagram of a task processing system provided by one embodiment of the present specification;
[0052] Figure 7 is a processing process flowchart of a task processing method provided by one embodiment of the present specification;
[0053] Figure 8 is a processing process flowchart of a data supplement method provided by one embodiment of the present specification;
[0054] Figure 9 is an interface schematic diagram of a task processing interface provided by one embodiment of the present specification;
[0055] Figure 10 is a structural schematic diagram of a task processing apparatus provided by an embodiment of the present specification;
[0056] Figure 11 is a structural schematic diagram of another task processing apparatus provided by an embodiment of the present specification;
[0057] Figure 12 is a structural schematic diagram of a data supplement apparatus provided by an embodiment of the present specification;
[0058] Figure 13 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0059] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples, set forth in this description. Those skilled in the art, in light of the description, can implement the present specification without limiting to the specific details disclosed in this description.
[0060] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0061] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second, and similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."
[0062] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0063] Firstly, the terms involved in one or more embodiments of the present specification are explained.
[0064] Data lake: A data lake is a centralized storage area for storing, processing, and protecting large amounts of structured, semi-structured, and unstructured data. Data lakes can ingest data very quickly and then dynamically prepare data for user access.
[0065] Large language model: A large language model (LLM) is an artificial intelligence system that is trained on large amounts of text data to learn language structures. A large-scale language model is pre-trained in the field of natural language processing, and excellent performance can be achieved through parameter fine-tuning or prompt learning in downstream tasks. Downstream tasks can be various natural language tasks such as text classification, question answering, and dialogue, and are an important way to artificial intelligence.
[0066] Prompt learning: Prompt-based learning is a strategy that can be used to use large language models. The prompt learning method does not require parameter updates to the model, so the same model can be used for different tasks without the need for retraining.
[0067] A data lake is a storage system that stores data in its original format. It can store data as is without the need for prior structured processing. A data lake can store structured data, semi-structured data, unstructured data, and binary data, etc. As the data generated by enterprises and individual users grows explosively, data lakes face the problem of how to support data discovery, extraction, cleaning and integration on large data sets with unknown structure, content and data quality.
[0068] Currently, data processing can generally be performed in the following ways: first, by deploying a console that can manage a specific subset of data, users can access the console to search and browse available data sets that meet their project needs. Second, deploy an analysis and machine learning-based optimization workload that enables users to view data lineage across clouds and transient clusters, with a single management platform across hybrid and multi-cloud environments. Third, provide serverless hosted enterprise data warehouses that enable organizations to analyze data by creating logical data warehouses on hosted, columnar storage, and data from object storage and spreadsheets. Fourth, provide a solution designed for cloud environments that enables users to store data of any size, shape, and speed, and also perform data processing and analysis across platforms and languages.
[0069] However, the above solutions have low efficiency in processing data in the data lake, and the effect and performance of the machine learning model are limited. Moreover, the machine learning model-based solution usually needs to design a corresponding model for a specific scene and a specific task, so that the cost of inference through the machine learning model is also high.
[0070] To solve the above problems, the present solution proposes a unified solution framework based on a large language model for data lake scenarios, which enhances the performance of the large language model through data retrieval, and is suitable for various data tasks in the data lake. Specifically, in response to a task processing request for the data lake, task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of a target task; the task attribute information and the task processing data are input into a retrieval unit in a task processing model to obtain reference data of the target task; the task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; and the target processing data is input into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request. By utilizing the language understanding ability and cross-task ability of the task processing model, the present solution better adapts to different tasks in the data lake, reduces repeated process development, and fully considers various data organization structures under the data lake during task processing. Through data retrieval using the task processing model, on the one hand, information related to the target task is extracted, the model processing performance is improved, and the task processing efficiency is further improved; on the other hand, invalid information is filtered, the consumption of calculation amount is reduced, and the use cost of the task processing model is reduced.
[0071] In the specification, a task processing method is provided, and the specification also relates to a data supplement method, a task processing system, a task processing device, a data supplement device, a computing device, a computer-readable storage medium and a computer program, which are described in the following embodiments one by one.
[0072] Referring to Figure 1 , Figure 1 An architecture diagram of a task processing system provided by an embodiment of the specification is shown, and the task processing system includes a client 100 and a server 200.
[0073] The client 100 is configured to send a task processing request for a data lake to the server 200, wherein the task processing request carries task attribute information of a target task.
[0074] The server 200 is configured to, in response to the task processing request for the data lake, acquire task processing data corresponding to the task processing request; input the task attribute information and the task processing data into a retrieval unit in a task processing model to obtain reference data of the target task; input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request; and send the task processing result to the client 100.
[0075] The client 100 is further configured to receive the task processing result sent by the server 200.
[0076] By using the language understanding capability and the cross-task capability of the task processing model, the scheme of the embodiment of the specification can better adapt to different tasks in the data lake, reduce repeated process development, and fully consider various data organization structures under the data lake in the task processing process. By using the task processing model for data retrieval, on the one hand, information related to the target task is extracted, the model processing performance is improved, the task processing efficiency is further improved, on the other hand, invalid information is filtered, the consumption of the calculation amount is reduced, and the use cost of the task processing model is reduced.
[0077] Referring to Figure 2 , Figure 2An architecture diagram of another task processing system provided by an embodiment of the present specification is shown, which can include a plurality of clients 100 and a server 200, wherein the client 100 can be referred to as an end-side device, and the server 200 can be referred to as a cloud-side device. The plurality of clients 100 can establish a communication connection through the server 200, and in a task processing scenario, the server 200 is used to provide a task processing service between the plurality of clients 100, and the plurality of clients 100 can respectively act as a sending end or a receiving end to realize communication through the server 200.
[0078] A user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, and the like. In a task processing scenario, the user can publish a data stream to the server 200 through the client 100, the server 200 generates a task processing result according to the data stream, and pushes the task processing result to other clients that establish a communication connection.
[0079] Among them, the client 100 and the server 200 establish a connection through a network. The network provides a medium for a communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or optical fiber cables, and the like. The data transmitted by the client 100 can need to be processed through encoding, transcoding, compression, and the like before being published to the server 200.
[0080] The client 100 can be a browser, an APP (Application, application program), or a web application such as an H5 (HyperText Markup Language 5, HyperText Markup Language version 5) application, or a light application (also known as a small program, a lightweight application program), or a cloud application, and the like. The client 100 can be developed based on a software development kit (SDK, Software Development Kit) provided by the server 200 for the corresponding service, such as based on a real-time communication (RTC, Real Time Communication) SDK, and the like. The client 100 can be deployed in an electronic device and needs to rely on the device or some APP in the device to run, and the like. The electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue type applications, model training type applications, text processing type applications, web browser applications, shopping type applications, search type applications, instant communication tools, mailbox clients, social platform software, and the like.
[0081] The server 200 can include a server providing various services, for example, a server providing a communication service for a plurality of clients, for example, a server for background training supporting a model used on a client, for example, a server processing data sent by a client, and the like. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of a cloud service, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, and the like. Basic cloud computing services of artificial intelligence technology, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0082] It should be noted that the task processing method provided in the embodiments of the present specification is generally executed by the server, but in other embodiments of the present specification, the client can also have similar functions as the server, so as to execute the task processing method provided in the embodiments of the present specification. In other embodiments, the task processing method provided in the embodiments of the present specification can also be executed by the client and the server together.
[0083] Referring to Figure 3 , Figure 3 A flowchart of a task processing method provided by one embodiment of the present specification is shown, which specifically includes the following steps:
[0084] Step 302: In response to a task processing request for a data lake, task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of a target task.
[0085] In one or more embodiments of the present specification, a user can send a task processing request for a data lake, so as to obtain task processing data corresponding to the task processing request in response to the data processing request, and perform task processing based on the task processing data and the task attribute information of the target task carried by the task processing request.
[0086] Specifically, the data lake can store structured data, semi-structured data, unstructured data, and binary data, etc., the structured data is like a table in a relational database. The semi-structured data is like comma-separated values (CSV), logs, extensible markup language (XML), and JavaScript object notation (JSON). The unstructured data is like emails, documents, and portable document format (PDF). The binary data is like graphics, audio, and video.
[0087] The request processing object of the task processing request is a target task for the data lake. The target task includes at least one of a data discovery task, a data cleaning task, a data supplement task, and a data integration task, wherein the data discovery task is like a relationship discovery task, and the data supplement task is like a missing value filling task. The target task can also include a data format conversion task, etc. In addition, the target task can be a task in different task scenarios, and the task scenarios include but are not limited to a translation scenario, a commodity identification scenario, and an article main idea extraction scenario. The target task is specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation thereto.
[0088] The task processing data refers to data required by the target task processing process, which is used to assist and guide the task processing process. The task attribute information refers to the information of the target task itself, and the data source of the task attribute information can also be the data lake. The task attribute information includes but is not limited to the task name, the task type, the target task content, and the task description information of the target task, which are specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation thereto.
[0089] It should be noted that when the task processing request does not carry a specific target task, but carries task description information of the target task, the task processing model can be used to understand the task description information and adapt the target task. For example, the task description information is "filling missing values in table data", and the task processing model can understand the task and automatically adapt the target task to "data filling".
[0090] In actual application, after receiving the task processing request for the data lake, the task processing data corresponding to the task processing request can be obtained in various ways in response to the task processing request for the data lake, which is specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation thereto. In one possible implementation manner of the present specification, the task processing data input by the user can be directly received.
[0091] In another possible implementation of the present specification, the task processing data corresponding to the task processing request can be obtained from the data lake, that is, the above-mentioned obtaining the task processing data corresponding to the task processing request in response to the task processing request for the data lake can include the following steps:
[0092] In response to the task processing request for the data lake, the task processing data corresponding to the task processing request is obtained from the data lake according to the task attribute information of the target task.
[0093] It should be noted that the task processing data obtained from the data lake can be structured data, unstructured data, semi-structured data, or data of multiple different structures. Moreover, the data format of the task processing data and the data format of the task attribute information can be the same or different, and the task processing data is selected according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.
[0094] For example, assuming that the task attribute information is "the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; zip code: xxxxx", the task processing data obtained from the data lake according to the task attribute information can be shown in Table 1 below, wherein "?" is the filling object corresponding to the data filling task.
[0095] Table 1: Task processing data example table
[0096]
[0097]
[0098] By applying the scheme of the embodiments of the present specification, in response to the task processing request for the data lake, the task processing data corresponding to the task processing request is obtained from the data lake according to the task attribute information of the target task, so that the task processing data can be searched subsequently, and reference is provided for the processing process of the task processing model, so that the task processing model can distinguish the data organization structure in the data lake, and unified processing across tasks is realized.
[0099] Step 304: inputting the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target task.
[0100] In one or more embodiments of the present specification, after obtaining the task processing data corresponding to the task processing request in response to the task processing request for the data lake, the task attribute information and the task processing data can be further input into a retrieval unit in the task processing model to obtain reference data of the target task.
[0101] Specifically, the task processing model is a large language model, which can be trained on a large amount of text data, so as to perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. The task processing model includes but is not limited to a generative pre-training language model (GPT, Generative Pre-trained Transformer), a bidirectional encoding language model (BERT, Bidirectional Encoder Representations from Transformers), a text-to-text conversion model (T5, Transfer Text-to-Text Transformer), and the task processing model is specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation on this. The retrieval unit in the task processing model is used to retrieve and extract information from the data input into the task processing model, so as to filter invalid information and obtain reference data of the target task. The reference data is used to prompt the processing process of the task processing model, so as to realize prompt learning of the task processing model.
[0102] In actual application, the retrieval unit retrieves the task attribute information and the task processing data to obtain the reference data of the target task in multiple ways, and the embodiments of the present specification do not make any limitation on this.
[0103] In a possible implementation of the present specification, the reference data of the target task can be retrieved from the task attribute information and the task processing data by using a rule matching method, wherein the matching rule is set according to actual conditions, such as keyword matching.
[0104] In another possible implementation of the present specification, the retrieval unit can use a preset retrieval template to extract information, that is, the above-mentioned inputting of the task attribute information and the task processing data into the retrieval unit in the task processing model to obtain the reference data of the target task can include the following steps:
[0105] extracting information from the task attribute information and the task processing data according to the preset retrieval template to obtain task filling data and candidate filling data;
[0106] filling the task filling data and the candidate filling data into the preset retrieval template to obtain target retrieval data;
[0107] inputting the target retrieval data into the retrieval unit in the task processing model to obtain the reference data of the target task.
[0108] Specifically, the preset retrieval template refers to a template for retrieving reference data, including but not limited to a meta-information retrieval template and an instance retrieval template. The task filling data is a retrieval target of the retrieval unit, and the candidate filling data is a retrieval object of the retrieval unit.
[0109] For example, it is assumed that the task attribute information is "the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; and zip code: xxxxx", and the task processing information is shown in Table 1. According to the preset retrieval template, the task attribute information and the task processing data are information extracted, the task filling data is "data filling; M city; time zone", and the candidate filling data is "city; country; population; and zip code".
[0110] It should be noted that when the task attribute information and the task processing data are information extracted according to the preset retrieval template, the filling position in the preset retrieval template can be obtained, and the corresponding filling data can be extracted from the task attribute information and the task processing data according to the filling prompt information before the filling position. Further, when the task filling data and the candidate filling data are filled into the preset retrieval template, the corresponding filling position can be filled, so as to ensure the accuracy and logical coherence of the target retrieval data.
[0111] According to the scheme of the embodiments of the present specification, the task attribute information and the task processing data are information extracted according to the preset retrieval template, the task filling data and the candidate filling data are obtained, the task filling data and the candidate filling data are filled into the preset retrieval template, the target retrieval data is obtained, and the target retrieval data is input into the retrieval unit in the task processing model, and the reference data of the target task is obtained. The task processing model can determine the processing target through the task filling data, and the processing process of the task processing model is prompted through the candidate filling data, and the task processing efficiency is improved.
[0112] In actual application, there are various ways to input the target retrieval data into the retrieval unit in the task processing model to obtain the reference data of the target task, which is specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.
[0113] In a possible implementation manner of the present specification, the retrieval unit can use its semantic understanding ability to take the target retrieval data as a question and answer task, and generate the reference data of the target task according to the target retrieval data.
[0114] In another possible implementation manner of the present specification, the retrieval unit can determine an association index of the task reference data and the candidate reference data, and further determine the reference data according to the association index, that is, the above-mentioned inputting the target retrieval data into the retrieval unit in the task processing model to obtain the reference data of the target task can include the following steps:
[0115] The target search data is input into a search unit in the task processing model, in which a correlation index between the task filling data and the candidate filling data is calculated;
[0116] According to the correlation index, reference data of the target task is filtered from the candidate filling data.
[0117] Specifically, the correlation index is an index obtained by performing a correlation score on the task filling data and the candidate filling data. There are various ways to calculate the correlation index between the task filling data and the candidate filling data in the search unit, including but not limited to cosine similarity, Euclidean distance, etc.
[0118] It should be noted that after the correlation index between the task filling data and the candidate filling data is calculated, the correlation index can be sorted, so that a preset number of candidate filling data is filtered from the candidate filling data as reference data of the target task according to the sorted correlation index.
[0119] By applying the scheme of the embodiments of the present specification, the target search data is input into a search unit in the task processing model, in which a correlation index between the task filling data and the candidate filling data is calculated; according to the correlation index, reference data of the target task is filtered from the candidate filling data, which improves the accuracy of the reference data.
[0120] In an optional embodiment of the present specification, the preset search template includes a meta information search template; the above information extraction on the task attribute information and the task processing data according to the preset search template to obtain the task filling data and the candidate filling data can include the following steps:
[0121] The meta information search template is used to extract information from the task attribute information to obtain task filling meta information;
[0122] The meta information search template is used to extract information from the task processing data to obtain candidate filling meta information.
[0123] Specifically, meta information refers to descriptive information, which includes but is not limited to the summary of a data lake and the name of each column of a table. The purpose of information extraction on the task attribute information and the task processing data according to the meta information search template is to obtain task filling meta information in the task attribute information and candidate filling meta information in the task processing data. After obtaining the task filling meta information and the candidate filling meta information, meta information retrieval can be performed in the search unit, the meta information related to the task filling meta information in the candidate filling meta information is measured, and the related meta information is extracted.
[0124] Exemplarily, assuming that the task attribute information is "the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; and zip code: xxxxx", the task processing information is shown in Table 1. The meta information retrieval template is "the features about [s1] are [s2, s3,..., sn], the task is [T], the target task content is [Y], and which features are related to the task and the target task content?", where the [] in the meta information retrieval template is used to identify the filling position, T is the target task, sn is the feature of the meta information, and sk is the missing feature that needs to be filled. According to the meta information retrieval template, the information extraction is performed on the task attribute information, and the task filling meta information is obtained as "data filling; time zone"; and the candidate filling meta information is obtained as "city; country; population; and zip code".
[0125] Further, the task filling meta information and the candidate filling meta information are filled into the meta information retrieval template, and the target retrieval data is obtained as "the features about [city] are [country, population, and zip code], the task is [data filling], the target task content is [time zone], and which features are related to the task and the target task content?". The target retrieval data "the features about [city] are [country, population, and zip code], the task is [data filling], the target task content is [time zone], and which features are related to the task and the target task content?" is input into the retrieval unit in the task processing model, and the reference meta information of the target task is obtained as "country".
[0126] According to the scheme of the embodiment of the present specification, the information extraction is performed on the task attribute information according to the meta information retrieval template, and the task filling meta information is obtained; the information extraction is performed on the task processing data according to the meta information retrieval template, and the candidate filling meta information is obtained, thereby improving the accuracy of the candidate filling meta information and the task filling meta information.
[0127] In another optional embodiment of the present specification, the preset retrieval template includes an instance retrieval template; and the information extraction is performed on the task attribute information and the task processing data according to the preset retrieval template to obtain the task filling data and the candidate filling data, which can include the following steps:
[0128] The information extraction is performed on the task attribute information according to the instance retrieval template to obtain a task filling instance;
[0129] The information extraction is performed on the task processing data according to the instance retrieval template to obtain a candidate filling instance.
[0130] Specifically, the instance refers to specific data, such as a piece of data in a data lake. The purpose of information extraction of the task attribute information and the task processing data according to the instance retrieval template is to obtain a task filling instance in the task attribute information and a candidate filling instance in the task processing data, respectively. After obtaining the task filling instance and the candidate filling instance, instance retrieval can be performed in the retrieval unit, the relevance between the candidate filling instance and the task filling instance is measured, and the relevant instance is extracted.
[0131] Exemplarily, it is assumed that the task attribute information is “the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; and zip code: xxxxx”, and the task processing information is shown in Table 1. The instance retrieval template is “the task is [T], the target task content is [Y], and the relevance of the following instances is scored (1 to 5 points): [r1, r3, …, rn]”, wherein the [] in the instance retrieval template is used to identify a filling position, T is a target task, and the target task content is a piece of data instance with features. According to the instance retrieval template, the information extraction is performed on the task attribute information, and the task filling instance is “data filling; M city”, and the candidate filling instance is “A city; B city; C city; D city; and E city”.
[0132] Further, the task filling instance and the candidate filling instance are filled into the instance retrieval template, and the target retrieval data is obtained as “the task is [data filling], the target task content is [M city], and the relevance of the following instances is scored (1 to 5 points): [A city; B city; C city; D city; and E city]”. The target retrieval data “the task is [data filling], the target task content is [M city], and the relevance of the following instances is scored (1 to 5 points): [A city; B city; C city; D city; and E city]” is input into the retrieval unit in the task processing model, the retrieval unit outputs the relevance indexes of each candidate filling instance, and the two candidate instances with the highest relevance indexes with the target task content are selected as reference instances according to the sorting of the relevance indexes, and the reference instances are “A city; B city; and C city”.
[0133] By applying the scheme of the embodiments of the present specification, the task filling instance is obtained by performing information extraction on the task attribute information according to the instance retrieval template, and the candidate filling instance is obtained by performing information extraction on the task processing data according to the instance retrieval template, thereby improving the accuracy of the candidate filling instance and the task filling instance.
[0134] Step 306: Input the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request.
[0135] In one or more embodiments of the present specification, in response to a task processing request for a data lake, task processing data corresponding to the task processing request is obtained, task attribute information and the task processing data are input into a retrieval unit in a task processing model, reference data of a target task is obtained, and then the task attribute information and the reference data can be input into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request.
[0136] Specifically, the data format of the target processing data is a format that is more conducive to processing by the task processing model, such as a natural description language format. The conversion unit is used to convert the data format of the task attribute information and the reference data, so that the data format of the task attribute information and the reference data is unified to a format that is more conducive to processing by the task processing model, such as a natural description language format. It should be noted that, taking the data before conversion as table data as an example, the conversion unit can use the language understanding capability of the task processing model to extract discrete effective data from the table data, and organize the effective data into a piece of natural description language.
[0137] In actual applications, there are various ways to input the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request, which are selected according to actual conditions, and the present specification does not make any limitation in this regard.
[0138] In a possible implementation manner of the present specification, the data format of the task attribute information and the reference data can be the same, so the task attribute information and the reference data can be directly spliced to obtain spliced processing data, and the spliced processing data can be input into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request.
[0139] In another possible implementation manner of the present specification, the data format of the task attribute information and the reference data can be different, and therefore, the data format of the reference data needs to be unified with the data format of the task attribute information. Taking conversion of the data format of the reference data according to the data format of the task attribute information as an example, after obtaining the converted reference data, the task attribute information and the converted reference data can be input into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request. When the data format of the reference data is converted according to the data format of the task attribute information, the conversion unit of the task processing model can be used for format conversion, or other data format conversion manners can be used, which are selected according to actual conditions, and the present specification does not make any limitation in this regard.
[0140] In an optional embodiment of the present specification, the above-mentioned inputting of the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request can include the following steps:
[0141] data decomposition is performed on the reference data, and the reference data after data decomposition is serialized to obtain reference data sequence;
[0142] The reference data sequence is input into a conversion unit in the task processing model to obtain converted reference data;
[0143] The task attribute information and the converted reference data are input into the conversion unit in the task processing model to generate target processing data corresponding to the task processing request.
[0144] It should be noted that data decomposition refers to converting reference data in an original data format into reference data in another data format, and only the data format is changed, and the data content is not changed. The conversion unit included in the task processing model can be one or more. For example, the task processing model includes a first conversion unit and a second conversion unit. After the reference data is data-decomposed and the reference data after data decomposition is serialized to obtain the reference data sequence, the reference data sequence can be input into the first conversion unit in the task processing model to obtain the converted reference data; the task attribute information and the converted reference data are input into the second conversion unit in the task processing model to generate target processing data corresponding to the task processing request.
[0145] Exemplarily, assuming that the task attribute information is “the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; postcode: xxxxx”, the reference information is “country”, and the reference instance is “A city; B city; C city”, the determined reference data is shown in Table 2 as follows:
[0146] Table 2 Reference data example table
[0147] City Country Time Zone City A a XX Central Time Zone City B b XX Central Time Zone City C c XX Eastern Time Zone
[0148] The reference data is data-decomposed to obtain data-decomposed reference data “city, country, time zone, A city, a, XX central time zone, B city, b, XX central time zone, C city, c, XX eastern time zone”, and the data-decomposed reference data is serialized to obtain reference data sequence “city: A city, country: a, time zone: XX central time zone; city: B city, country: b, time zone: XX central time zone; city: C city, country: c, time zone: XX eastern time zone”. The reference data sequence is input into the conversion unit in the task processing model to obtain converted reference data “A city is a city in a country, and its time zone is XX central time zone; B city is a city in b country, and its time zone is XX central time zone; C city is a city in c country, and its time zone is XX eastern time zone”.
[0149] In actual application, there are various ways to input the reference data sequence into the conversion unit in the task processing model to obtain the converted reference data, which are selected according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.
[0150] In a possible implementation of the present specification, the reference data sequence can be directly input into the conversion unit in the task processing model to obtain the converted reference data.
[0151] In another possible implementation of the present specification, the conversion prompt information can be obtained, the reference data sequence and the conversion prompt information are input into the task processing model, and the converted reference data is obtained by using the conversion prompt information to assist the task processing model.
[0152] For example, assuming that the conversion prompt information is “convert the information into a text format, and include all relevant information in a logical order”, the conversion prompt information and the reference data sequence “city: A city, country: a, time zone: XX Central Time Zone; city: B city, country: b, time zone: XX Central Time Zone; city: C city, country: c, time zone: XX Eastern Time Zone” are input into the conversion unit of the task processing model, and the converted reference data is “A city is a city in a country, and its time zone is XX Central Time Zone; B city is a city in a country, and its time zone is XX Central Time Zone; C city is a city in a country, and its time zone is XX Eastern Time Zone”.
[0153] By applying the scheme of the embodiments of the present specification, the reference data is subjected to data decomposition, and the reference data after data decomposition is subjected to serialization processing to obtain a reference data sequence; the reference data sequence is input into the conversion unit in the task processing model to obtain converted reference data; and the task attribute information and the converted reference data are input into the conversion unit in the task processing model to generate target processing data corresponding to a task processing request. Through the conversion unit in the task processing model, the data formats of the reference data and the task attribute information are converted into formats more suitable for processing by the task processing model, which enables the task processing model to be better adapted to different tasks in the data lake, and reduces repeated process development.
[0154] In actual application, there are various ways to input the task attribute information and the converted reference data into the conversion unit in the task processing model to generate target processing data corresponding to a task processing request, which are selected according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.
[0155] In a possible implementation of the present specification, the task attribute information and the converted reference data can be directly merged, and the merged data is input into the conversion unit in the task processing model to generate target processing data corresponding to a task processing request.
[0156] In another possible implementation of the present specification, before the task attribute information and the converted reference data are input into the conversion unit in the task processing model, the task attribute information and the converted reference data can also be processed by using a preset processing template for guiding the conversion unit in the task processing model to perform data format conversion, that is, the above-mentioned inputting the task attribute information and the converted reference data into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request can include the following steps:
[0157] obtaining a preset processing template;
[0158] filling the task attribute information and the converted reference data into the preset processing template to obtain to-be-processed data;
[0159] inputting the to-be-processed data into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request.
[0160] Exemplarily, it is assumed that the task attribute information is "the target task is data filling, the target task content is city: M city; time zone:?; country: m; population: 875437; and zip code: xxxxx", and the preset processing template is "context content is [d], task is [T], and target task content is [Y]", where the context content is the converted reference data. The converted reference data is "A city is a city in a country, and its time zone is XX central time zone; B city is a city in a country, and its time zone is XX central time zone; and C city is a city in a country, and its time zone is XX eastern time zone".
[0161] filling the task attribute information and the converted reference data into the preset processing template to obtain the to-be-processed data is "context content is [A city is a city in a country, and its time zone is XX central time zone; B city is a city in a country, and its time zone is XX central time zone; and C city is a city in a country, and its time zone is XX eastern time zone], task is [data filling], and target task content is [M city; time zone]"; and inputting the to-be-processed data into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request is "the task is to fill in the missing values in the data. The context content is A city is a city in a country, and its time zone is XX central time zone; B city is a city in a country, and its time zone is XX central time zone; and C city is a city in a country, and its time zone is XX eastern time zone. M city is a city in a country, and it belongs to which time zone?".
[0162] In an optional embodiment of the present specification, in order to improve the accuracy of the target processing data generated by the conversion unit for the task processing request, a preset conversion strategy can be pre-set in the conversion unit, or the preset conversion strategy, the task attribute information and the converted reference data are directly input into the conversion unit in the task processing model, and the conversion unit can learn the conversion rule from the preset conversion strategy, so as to generate the target processing data corresponding to the task processing request.
[0163] For example, the preset conversion strategy can be "convert the declaration description into processing data. Declaration: the context content is "city is the area where human beings live... smart city...", and the task is "data discovery". The target content is "smart city". Processing data: the task is to find data from the context content... city is the area where human beings live... what is a smart city?", therefore, the conversion unit can determine the to-be-processed data as a declaration, and generate the target processing data "the task is to fill in the missing values in the data. The context content is that A city is a city in a country, whose time zone is XX central time zone; B city is a city in a country, whose time zone is XX central time zone; C city is a city in a country, whose time zone is XX eastern time zone. M city is a city in a country, which belongs to which time zone?" corresponding to the task processing request based on the preset conversion strategy.
[0164] By applying the scheme of the embodiments of the present specification, the preset processing template is obtained; the task attribute information and the converted reference data are filled into the preset processing template to obtain the to-be-processed data; and the to-be-processed data is input into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request. The accuracy of the target processing data is improved by the preset processing template, and the data in the location structure in the data lake is converted into the target processing data which is more conducive to the processing of the task processing model. The task processing model is applied to the data lake scene, so that the task processing model can better adapt to different tasks in the data lake and reduce repeated process development.
[0165] Step 308: inputting the target processing data into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request.
[0166] In one or more embodiments of the present specification, in response to the task processing request for the data lake, the task processing data corresponding to the task processing request is obtained, the task attribute information and the task processing data are input into the retrieval unit in the task processing model to obtain the reference data of the target task, the task attribute information and the reference data are input into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request, and further, the target processing data can be input into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request.
[0167] Exemplarily, the target processing data "task is to fill in the missing values in the data. The context content is that A city is a city in a country, and its time zone is XX Central Time Zone; B city is a city in a country, and its time zone is XX Central Time Zone; C city is a city in a country, and its time zone is XX Eastern Time Zone. Which time zone does M city belong to?" The processing unit in the task processing model is input, and the task processing result corresponding to the task processing request is "XX Central Time Zone".
[0168] By using the language understanding capability and cross-task capability of the task processing model, the scheme of the embodiment of the present specification better adapts to different tasks in the data lake, reduces repeated process development, and fully considers various data organization structures under the data lake in the task processing process. By using the task processing model for data retrieval, on the one hand, information related to the target task is extracted, the model processing performance is improved, and the task processing efficiency is further improved, and on the other hand, invalid information is filtered, the consumption of the calculation amount is reduced, and the use cost of the task processing model is reduced.
[0169] Referring to Figure 4 , Figure 4 A flowchart of another task processing method provided by an embodiment of the present specification is shown, which is applied to a cloud-side device, and specifically includes the following steps:
[0170] Step 402: In response to the task processing request for the data lake sent by the end-side device, obtain the task processing data corresponding to the task processing request, wherein the task processing request carries the task attribute information of the target task.
[0171] Step 404: Input the task attribute information and the task processing data into the retrieval unit in the task processing model to obtain the reference data of the target task.
[0172] Step 406: Input the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request.
[0173] Step 408: Input the target processing data into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request.
[0174] Step 410: Send the task processing result corresponding to the task processing request to the end-side device.
[0175] It should be noted that the implementation modes of steps 402 to 408 are the same as those of steps 302 to 308 described above, and the embodiments of the present specification will not be repeated.
[0176] By applying the scheme of the embodiment of the present specification, the end-side device sends a task processing request for the data lake to the cloud-side device, the cloud-side device better adapts to different tasks in the data lake by using the language understanding capability and cross-task capability of the task processing model, reduces repeated process development of the cloud-side device, and fully considers various data organization structures under the data lake in the task processing process. By using the task processing model for data retrieval, on the one hand, information related to the target task is extracted, the model processing performance is improved, and the task processing efficiency is further improved. On the other hand, invalid information is filtered, the consumption of the calculation amount of the cloud-side device is reduced, and the use cost of the task processing model is reduced. Moreover, the cloud-side device generates a task processing result corresponding to the task processing request without consuming the calculation resources of the end-side device, and the resource consumption of the end-side device is reduced.
[0177] Referring to Figure 5 , Figure 5 A flowchart of a data supplement method provided by one embodiment of the present specification is shown, which specifically includes the following steps:
[0178] Step 502: In response to a data supplement request for the data lake, task processing data corresponding to the data supplement request is obtained, wherein the data supplement request carries task attribute information of a target data supplement task.
[0179] Step 504: The task attribute information and the task processing data are input into a retrieval unit in the task processing model to obtain reference data of the target data supplement task.
[0180] Step 506: The task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the data supplement request.
[0181] Step 508: The target processing data is input into a processing unit in the task processing model to obtain a data supplement result corresponding to the data supplement request.
[0182] It should be noted that the implementation manners of steps 502 to 508 are the same as those of steps 302 to 308 described above, and the embodiment of the present specification will not be described again.
[0183] By applying the scheme of the embodiment of the present specification, the language understanding capability and cross-task capability of the task processing model are used to better adapt to the data supplement task in the data lake, reduce repeated process development, and use the task processing model for data retrieval in the target data supplement task processing process. On the one hand, information related to the target data supplement task is extracted, the model processing performance is improved, and the task processing efficiency is further improved. On the other hand, invalid information is filtered, the consumption of the calculation amount is reduced, and the use cost of the task processing model is reduced.
[0184] Referring to Figure 6 , Figure 6 A framework diagram of a task processing system is shown, the task processing system comprising a data processing component and a task processing component;
[0185] The data processing component is configured to, in response to a task processing request for a data lake, obtain task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task; input the task attribute information and the task processing data into a retrieval unit in a task processing model to obtain reference data of the target task; and send the reference data of the target task to the task processing component;
[0186] The task processing component is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; and input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0187] It is worth noting that, as Figure 6 shown, in the embodiments of the present specification, a task processing system for a data lake scenario based on a task processing model can be built based on the semantic understanding ability, general knowledge reserve, cross-domain ability and cross-task ability of the task processing model. The data processing component in the task processing system can cope with a large amount of data in the data lake that is not uniform in format (such as databases, tables, and texts), retrieve the data, and further perform downstream tasks based on the data that is not uniform in format. The task processing component can cope with tool problems in the data lake and can use the task processing model to solve different tasks in the data lake (such as data discovery tasks, data cleaning tasks, and data integration tasks).
[0188] Referring to Figure 7 , Figure 7 A processing process flow diagram of a task processing method is shown, the method comprising:
[0189] Data retrieval: After obtaining data from a data source, the task processing model can extract data required by a target task according to the target task. Data retrieval can be divided into meta-information retrieval and instance retrieval, and the data source can be a data lake including unstructured data, semi-structured data and structured data, and the target task can be data management, data discovery, data extraction, data cleaning and data integration.
[0190] Data decomposition: The data obtained by retrieval and extraction is decomposed to obtain decomposed data. The decomposed data is converted into a format more suitable for processing by a large model using the task processing model.
[0191] Automatic question integration: automatically integrate the target task attribute information and the decomposed data after format conversion, convert the integrated data into the question format of the task processing model, that is, the target processing data.
[0192] Output: input the data obtained by automatic question integration into the task processing model to obtain the final task processing result.
[0193] It should be noted that the efficiency of the task processing model can be improved by using the prompting method engineering in the automatic question integration process. For different data sources and data tasks, the task processing method provided by the embodiments of the present specification can adaptively adjust the model using the capabilities of the task processing model, automatically perform data retrieval based on the task processing model, and automatically associate related data according to specific tasks. While improving the method effect, it can reduce the cost of using the model.
[0194] The following describes the data supplement method provided by the present specification in detail with reference to the accompanying drawings. Figure 8 The data supplement method provided by the present specification is further described by taking the application of the data supplement method in the data lake scene as an example. Among them, Figure 8 FIG. 1 shows a process flow diagram of a data supplement method provided by an embodiment of the present specification, which specifically includes the following steps:
[0195] Data input: in response to a task processing request for a data lake, obtaining task processing data corresponding to the task processing request, wherein the task processing request carries task attribute information of a target task.
[0196] Automatic retrieval: performing information extraction on the task attribute information and the task processing data according to a preset retrieval template to obtain task filling data and candidate filling data; filling the task filling data and the candidate filling data into the preset retrieval template to obtain target retrieval data; inputting the target retrieval data into a retrieval unit in the task processing model, in which the retrieval unit, calculating an association index between the task filling data and the candidate filling data; according to the association index, screening reference data of the target task from the candidate filling data.
[0197] Data decomposition: decomposing the reference data and serializing the decomposed reference data to obtain a reference data sequence; inputting the reference data sequence into a conversion unit in the task processing model to obtain converted reference data.
[0198] Target processing data integration: obtaining a preset processing template; filling the task attribute information and the converted reference data into the preset processing template to obtain to-be-processed data; inputting the to-be-processed data into the conversion unit in the task processing model to generate target processing data corresponding to the task processing request.
[0199] Output: input the target processing data into the processing unit in the task processing model, obtain the task processing result corresponding to the task processing request.
[0200] By applying the scheme of the embodiments of the present specification, the language understanding capability and cross-task capability of the task processing model are fully utilized, the use process of the task processing model is designed for the multi-process scenario under the data lake, so that the task processing framework can better adapt to different tasks and reduce repeated process development. In addition, for the data scenario under the data lake, the data organization structure under the data lake is fully considered, the data organization structure is discriminated by the task processing model, and retrieval is performed from two aspects of meta information and instances. Through data retrieval, on the one hand, relevant information is extracted, the performance of the model is improved, and on the other hand, invalid information is filtered, and the consumption of calculation amount is reduced.
[0201] Referring to Figure 9 , Figure 9 An interface schematic diagram of a task processing interface provided by an embodiment of the present specification is shown. The task processing interface includes a task processing request input interface and a task processing result display interface. The task processing request input interface includes a task processing request input box, a "confirm" control and a "cancel" control. The task processing result display interface includes a task processing result display box.
[0202] The user inputs the task processing request through the task processing request input box displayed by the client, selects the "confirm" control, and the client sends the task processing request to the server; the server obtains the task processing data corresponding to the task processing request in response to the task processing request for the data lake, wherein the task processing request carries the task attribute information of the target task; inputs the task attribute information and the task processing data into the retrieval unit in the task processing model, obtains the reference data of the target task; inputs the task attribute information and the reference data into the conversion unit in the task processing model, obtains the target processing data corresponding to the task processing request; inputs the target processing data into the processing unit in the task processing model, obtains the task processing result corresponding to the task processing request, and sends the task processing result to the client. The client displays the task processing result in the task processing result display box.
[0203] In actual application, the operation mode of the user on the control includes any one of clicking, double-clicking, touch control, mouse hovering, sliding, long pressing, voice control or shaking, and the specific selection is made according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.
[0204] Corresponding to the method embodiments described above, the present specification also provides task processing device embodiments, Figure 10 A structural schematic diagram of a task processing device provided by an embodiment of the present specification is shown. As Figure 10 shown, the device includes:
[0205] The first obtaining module 1002 is configured to obtain task processing data corresponding to a task processing request in response to the task processing request for a data lake, wherein the task processing request carries task attribute information of a target task;
[0206] The first input module 1004 is configured to input the task attribute information and the task processing data into a retrieval unit in a task processing model to obtain reference data of the target task.
[0207] The second input module 1006 is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request.
[0208] The third input module 1008 is configured to input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request.
[0209] Optionally, the first input module 1004 is further configured to perform information extraction on the task attribute information and the task processing data according to a preset retrieval template to obtain task filling data and candidate filling data; fill the task filling data and the candidate filling data into the preset retrieval template to obtain target retrieval data; and input the target retrieval data into the retrieval unit in the task processing model to obtain the reference data of the target task.
[0210] Optionally, the first input module 1004 is further configured to input the target retrieval data into the retrieval unit in the task processing model, calculate an association index between the task filling data and the candidate filling data in the retrieval unit, and select the reference data of the target task from the candidate filling data according to the association index.
[0211] Optionally, the preset retrieval template includes a meta information retrieval template; and the first input module 1004 is further configured to perform information extraction on the task attribute information according to the meta information retrieval template to obtain task filling meta information, and perform information extraction on the task processing data according to the meta information retrieval template to obtain candidate filling meta information.
[0212] Optionally, the preset retrieval template includes an instance retrieval template; and the first input module 1004 is further configured to perform information extraction on the task attribute information according to the instance retrieval template to obtain task filling instances, and perform information extraction on the task processing data according to the instance retrieval template to obtain candidate filling instances.
[0213] Optionally, the second input module 1006 is further configured to perform data decomposition on the reference data, and perform serialization processing on the data-decomposed reference data to obtain reference data sequences; input the reference data sequences into a conversion unit in the task processing model to obtain converted reference data; and input the task attribute information and the converted reference data into the conversion unit in the task processing model to generate target processing data corresponding to the task processing request.
[0214] Optionally, the second input module 1006 is further configured to obtain a preset processing template; fill the task attribute information and the converted reference data into the preset processing template to obtain to-be-processed data; and input the to-be-processed data into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request.
[0215] Optionally, the first obtaining module 1002 is further configured to, in response to the task processing request for the data lake, obtain task processing data corresponding to the task processing request from the data lake according to the task attribute information of the target task.
[0216] Optionally, the target task includes at least one of a data discovery task, a data cleaning task, a data supplement task, and a data integration task.
[0217] By applying the scheme of the embodiments of the present specification, in response to a task processing request for a data lake, task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of a target task; the task attribute information and the task processing data are input into a retrieval unit in a task processing model to obtain reference data of the target task; the task attribute information and the reference data are input into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request; and the target processing data is input into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request. By utilizing the language understanding capability and the cross-task capability of the task processing model, different tasks in the data lake are better adapted, repeated process development is reduced, and in the task processing process, various data organization structures under the data lake are fully considered, data retrieval is performed by utilizing the task processing model, on the one hand, information related to the target task is extracted, the model processing performance is improved, the task processing efficiency is further improved, on the other hand, invalid information is filtered, the consumption of the calculation amount is reduced, and the use cost of the task processing model is reduced.
[0218] The above is a schematic scheme of a task processing device of the present embodiment. It should be noted that the technical scheme of the task processing device belongs to the same concept as the technical scheme of the task processing method described above, and the details of the technical scheme of the task processing device that are not described in detail can be referred to the description of the technical scheme of the task processing method.
[0219] Corresponding to the method embodiments described above, the specification also provides task processing device embodiments applied to a cloud-side device, Figure 11 A structural schematic diagram of another task processing device provided by an embodiment of the specification is shown, which is applied to a cloud-side device, as shown in Figure 11 The device comprises:
[0220] The second acquisition module 1102 is configured to acquire task processing data corresponding to the task processing request in response to the task processing request for the data lake sent by the end-side device, wherein the task processing request carries task attribute information of a target task;
[0221] The fourth input module 1104 is configured to input the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target task;
[0222] The fifth input module 1106 is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the task processing request;
[0223] The sixth input module 1108 is configured to input the target processing data into a processing unit in the task processing model to obtain a task processing result corresponding to the task processing request;
[0224] The sending module 1110 is configured to send the task processing result corresponding to the task processing request to the end-side device.
[0225] By applying the scheme of the embodiments of the specification, the end-side device sends a task processing request for the data lake to the cloud-side device, the cloud-side device better adapts to different tasks in the data lake by utilizing the language understanding capability and cross-task capability of the task processing model, reduces repeated process development of the cloud-side device, and fully considers various data organization structures under the data lake in the task processing process, performs data retrieval by utilizing the task processing model, extracts information related to the target task on the one hand, improves the model processing performance, further improves the task processing efficiency, filters invalid information on the other hand, reduces the consumption of the cloud-side device computing amount, and reduces the use cost of the task processing model. Moreover, the cloud-side device generates a task processing result corresponding to the task processing request without consuming the computing resources of the end-side device, thereby reducing the resource consumption of the end-side device.
[0226] The above is a schematic scheme of the task processing device applied to the cloud-side equipment according to an embodiment of the present specification. It should be noted that the technical scheme of the task processing device belongs to the same concept as the technical scheme of the task processing method applied to the cloud-side equipment described above. The technical scheme of the task processing device that is not described in detail can be seen from the description of the technical scheme of the task processing method applied to the cloud-side equipment described above.
[0227] Corresponding to the method embodiments described above, the present specification also provides data supplement device embodiments, Figure 12 A structural schematic diagram of a data supplement device provided by an embodiment of the present specification is shown. As shown in the figure, Figure 12 The device includes:
[0228] The third acquisition module 1202 is configured to acquire task processing data corresponding to the data supplement request in response to the data supplement request for the data lake, wherein the data supplement request carries task attribute information of a target data supplement task;
[0229] The seventh input module 1204 is configured to input the task attribute information and the task processing data into a retrieval unit in the task processing model to obtain reference data of the target data supplement task;
[0230] The eighth input module 1206 is configured to input the task attribute information and the reference data into a conversion unit in the task processing model to obtain target processing data corresponding to the data supplement request;
[0231] The ninth input module 1208 is configured to input the target processing data into a processing unit in the task processing model to obtain a data supplement result corresponding to the data supplement request.
[0232] By using the language understanding ability and cross-task ability of the task processing model, the scheme of the embodiment of the present specification better adapts to the data supplement task in the data lake, reduces repeated process development, and uses the task processing model to retrieve data in the target data supplement task processing process, on the one hand, extracts information related to the target data supplement task, improves the model processing performance, further improves the task processing efficiency, and on the other hand, filters invalid information, reduces the consumption of calculation amount, and reduces the use cost of the task processing model.
[0233] The above is a schematic scheme of the data supplement device according to an embodiment of the present specification. It should be noted that the technical scheme of the data supplement device belongs to the same concept as the technical scheme of the data supplement method described above. The technical scheme of the data supplement device that is not described in detail can be seen from the description of the technical scheme of the data supplement method.
[0234] Figure 13A structural block diagram of a computing device is shown. The components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 through a bus 1330, and a database 1350 is used to store data.
[0235] The computing device 1300 also includes an access device 1340 that enables the computing device 1300 to communicate via one or more networks 1360. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 1340 can include one or more of any type of network interface (for example, a network interface card (NIC)), wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0236] In one embodiment of the present specification, the above-mentioned components of the computing device 1300 and other components not shown in the Figure 13 may be connected to each other, for example, through a bus. It should be understood that Figure 13 The structural block diagram of the computing device shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art.
[0237] The computing device 1300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1300 can also be a mobile or stationary server.
[0238] The processor 1320 is configured to execute computer-executable instructions to perform the steps of the task processing method or the data supplement method.
[0239] The above is a schematic solution of the computing device of the embodiment. It should be noted that the technical solution of the computing device belongs to the same concept as the technical solutions of the task processing method and the data supplement method, and the details of the technical solution of the computing device that are not described in detail can be referred to the description of the technical solutions of the task processing method or the data supplement method.
[0240] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the task processing method or the data supplement method.
[0241] The above is a schematic solution of the computer-readable storage medium of the embodiment. It should be noted that the technical solution of the storage medium belongs to the same concept as the technical solutions of the task processing method and the data supplement method, and the details of the technical solution of the storage medium that are not described in detail can be referred to the description of the technical solutions of the task processing method or the data supplement method.
[0242] An embodiment of the present specification further provides a computer program, which, when executed in a computer, causes the computer to perform the steps of the task processing method or the data supplement method.
[0243] The above is a schematic solution of the computer program of the embodiment. It should be noted that the technical solution of the computer program belongs to the same concept as the technical solutions of the task processing method and the data supplement method, and the details of the technical solution of the computer program that are not described in detail can be referred to the description of the technical solutions of the task processing method or the data supplement method.
[0244] The above-described embodiments of the application have several aspects, no single one of which is solely responsible for the application's desirable attributes. Without limiting the scope of the application as expressed by the claims which follow, some further embodiments make these aspects even more useful. Other embodiments can result in less desirable attributes.
[0245] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc.
[0246] It should be noted that, for the foregoing method embodiments, the acts described can be performed in serial, parallel, or in some other order. In other words, the order of the acts of the embodiments can be modified without changing the underlying nature of the embodiments. Also, the embodiments described herein can be performed in a different order than the order described. Furthermore, some acts can be performed in a different order than the order described, or even at the same time. In addition, it should be noted that described acts can be carried out in one embodiment, but can be optional in other embodiments.
[0247] In the above-described embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0248] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The alternative embodiments do not describe all the details, nor limit the application to the specific embodiments described. Obviously, according to the content of the embodiments of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical application of the embodiments of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited only by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: In response to a task processing request for the data lake, task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of the target task, and the task processing data is data used to assist and guide the task processing process corresponding to the task processing request. The task attribute information and the task processing data are input into the retrieval unit in the task processing model to obtain reference data for the target task. The retrieval unit is used to retrieve and extract information from the task attribute information and the task processing data, and the reference data is used to prompt the processing process of the task processing model. The task attribute information and the reference data are input into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request. The target processing data is data in a data format that is easy for the task processing model to process. The target processing data is input into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request.
2. The method according to claim 1, wherein inputting the task attribute information and the task processing data into the retrieval unit in the task processing model to obtain reference data for the target task includes: Based on a preset retrieval template, information is extracted from the task attribute information and the task processing data to obtain task filling data and candidate filling data. The task data and candidate data are filled into the preset search template to obtain the target search data. The target retrieval data is input into the retrieval unit in the task processing model to obtain reference data for the target task.
3. The method according to claim 2, wherein inputting the target retrieval data into the retrieval unit in the task processing model to obtain reference data for the target task includes: The target retrieval data is input into the retrieval unit in the task processing model, and the correlation index between the task filling data and the candidate filling data is calculated in the retrieval unit. Reference data for the target task is selected from the candidate filling data based on the correlation indicators.
4. The method according to claim 2, wherein the preset search template includes a metadata search template; The step of extracting information from the task attribute information and the task processing data according to a preset retrieval template to obtain task filling data and candidate filling data includes: Based on the meta-information retrieval template, information is extracted from the task attribute information to obtain task filling meta-information; Information is extracted from the task processing data according to the meta-information retrieval template to obtain candidate filling meta-information.
5. The method according to claim 2, wherein the preset search template includes an instance search template; The step of extracting information from the task attribute information and the task processing data according to a preset retrieval template to obtain task filling data and candidate filling data includes: Based on the instance retrieval template, information is extracted from the task attribute information to obtain a task filling instance; Based on the instance retrieval template, information is extracted from the task processing data to obtain candidate filling instances.
6. The method according to claim 1, wherein inputting the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request includes: The reference data is decomposed, and the decomposed reference data is serialized to obtain a reference data sequence. The reference data sequence is input into the conversion unit in the task processing model to obtain the converted reference data. The task attribute information and the converted reference data are input into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request.
7. The method according to claim 6, wherein inputting the task attribute information and the converted reference data into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request includes: Get the preset processing template; The task attribute information and the converted reference data are filled into a preset processing template to obtain the data to be processed. The data to be processed is input into the conversion unit in the task processing model to generate the target processing data corresponding to the task processing request.
8. The method according to claim 1, wherein obtaining task processing data corresponding to a task processing request for a data lake comprises: In response to a task processing request for the data lake, task processing data corresponding to the task processing request is obtained from the data lake based on the task attribute information of the target task.
9. The method according to claim 1, wherein the target task includes at least one of a data discovery task, a data cleaning task, a data supplementation task, and a data integration task.
10. A task processing method applied to a cloud-side device, the method comprising: In response to a task processing request for the data lake sent by the edge device, the task processing data corresponding to the task processing request is obtained, wherein the task processing request carries task attribute information of the target task, and the task processing data is data used to assist and guide the task processing process corresponding to the task processing request. The task attribute information and the task processing data are input into the retrieval unit in the task processing model to obtain reference data for the target task. The retrieval unit is used to retrieve and extract information from the task attribute information and the task processing data, and the reference data is used to prompt the processing process of the task processing model. The task attribute information and the reference data are input into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request. The target processing data is data in a data format that is easy for the task processing model to process. The target processing data is input into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request. Send the task processing result corresponding to the task processing request to the terminal device.
11. A data supplementation method, comprising: In response to a data supplementation request for a data lake, task processing data corresponding to the data supplementation request is obtained, wherein the data supplementation request carries task attribute information of the target data supplementation task, and the task processing data is data used to assist and guide the task processing process corresponding to the data supplementation request; The task attribute information and the task processing data are input into the retrieval unit in the task processing model to obtain the reference data for the target data supplementing the task. The retrieval unit is used to retrieve and extract information from the task attribute information and the task processing data, and the reference data is used to prompt the processing process of the task processing model. The task attribute information and the reference data are input into the conversion unit in the task processing model to obtain the target processing data corresponding to the data supplementation request. The target processing data is data in a data format that is easy for the task processing model to process. The target processing data is input into the processing unit in the task processing model to obtain the data supplementation result corresponding to the data supplementation request.
12. A task processing system, comprising a data processing component and a task processing component; The data processing component is configured to, in response to a task processing request for the data lake, obtain task processing data corresponding to the task processing request, wherein... The task processing request carries the task attribute information of the target task, and the task processing data is data used to assist and guide the task processing process corresponding to the task processing request; the task attribute information and the task processing data are input into the retrieval unit in the task processing model to obtain reference data of the target task, wherein the retrieval unit is used to retrieve and extract information from the task attribute information and the task processing data, and the reference data is used to prompt the processing process of the task processing model; the reference data of the target task is sent to the task processing component; The task processing component is configured to input the task attribute information and the reference data into the conversion unit in the task processing model to obtain the target processing data corresponding to the task processing request; and to input the target processing data into the processing unit in the task processing model to obtain the task processing result corresponding to the task processing request, wherein the target processing data is data in a format that is conducive to the processing of the task processing model.
13. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9, or claim 10 or claim 11.
14. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method of any one of claims 1 to 9, or claim 10, or claim 11.
Citation Information
Patent Citations
Question answering method, device, electronic device and storage medium
CN109284363A