Data processing method, computing device, storage medium and computer program product
By using problem analysis model and problem processing model in data processing, we can efficiently obtain target reference data and quickly determine the problem processing results, and solve the problem of low data processing efficiency in the prior art, and achieve more efficient data processing.
Patent Information
- Application Number
- CN202411732000.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, data processing efficiency is low, mainly due to excessive time and labor cost consumption caused by excessive data volume.
By using the problem analysis model to determine the target problem and data storage unit of the pending problem, and using the problem processing model to determine the data query statement, thereby efficiently obtaining the target reference data and quickly determining the problem processing results based on the data.
This improves the efficiency of problem processing, reduces the time and labor cost consumption caused by excessive data volume, and avoids the problem of low problem processing efficiency.
Smart Images

Figure CN119961280A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a data processing method. One or more embodiments of this specification also relate to another data processing method, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the continuous development of computer technology, a large amount of data will be generated in the process of using computers to perform various data processing operations; and these generated data contain important information that can be used to solve the problems faced by users.
[0003] Currently, in the process of using data to deal with questions raised by users, manual sorting is often used to find data related to the problem from the data, and the found data is used to deal with the problem; but due to the large amount of data, a lot of time and manpower costs will be consumed for data processing, resulting in low problem handling efficiency; therefore, how to improve problem handling efficiency has become an issue that needs to be urgently addressed. Summary of the invention
[0004] In view of this, an embodiment of this specification provides a data processing method. One or more embodiments of this specification also relate to another data processing method, a data processing device, another data processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a data processing method is provided, including: Determine a problem to be processed, and use a problem analysis model to analyze the problem to be processed, determine a target problem and a data storage unit corresponding to the problem to be processed, wherein the data storage unit stores reference data obtained from multiple types of data sources; Utilizing the problem processing model, a data query statement is determined based on the data storage unit and the target problem. By executing the data query statement, target reference data corresponding to the target problem is obtained from the data storage unit, and based on the target reference data, a problem processing result corresponding to the problem to be processed is determined.
[0006] According to a second aspect of an embodiment of this specification, there is provided a data processing device, including: A problem determination module is configured to determine a problem to be processed, and use a problem analysis model to perform problem analysis on the problem to be processed, and determine a target problem corresponding to the problem to be processed and a data storage unit, wherein the data storage unit stores reference data obtained from multiple types of data sources; The result determination module is configured to utilize the problem processing model to determine a data query statement based on the data storage unit and the target problem, obtain target reference data corresponding to the target problem from the data storage unit by executing the data query statement, and determine the problem processing result corresponding to the problem to be processed based on the target reference data.
[0007] According to a third aspect of an embodiment of this specification, a data processing method is provided, including: Determine a logistics problem to be processed, and use a problem analysis model to analyze the logistics problem to be processed, determine a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, wherein the logistics data storage unit stores logistics reference data obtained from multiple types of data sources; Utilizing the problem analysis model, a data query statement is determined based on the logistics data storage unit and the target logistics problem. By executing the data query statement, target logistics reference data corresponding to the target logistics problem is obtained from the logistics data storage unit. Based on the target logistics reference data, a logistics problem processing result corresponding to the logistics problem to be processed is determined.
[0008] According to a fourth aspect of an embodiment of this specification, there is provided a data processing device, including: A problem determination module is configured to determine a logistics problem to be processed, and use a problem analysis model to perform problem analysis on the logistics problem to be processed, and determine a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, wherein the logistics data storage unit stores logistics reference data obtained from multiple types of data sources; The result determination module is configured to utilize the problem analysis model to determine a data query statement based on the logistics data storage unit and the target logistics problem, obtain target logistics reference data corresponding to the target logistics problem from the logistics data storage unit by executing the data query statement, and determine the logistics problem processing result corresponding to the logistics problem to be processed based on the target logistics reference data.
[0009] According to a fifth aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of any one of the above methods are implemented.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the computer program / instruction implements the steps of any one of the above methods when executed by a processor.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of any one of the above methods when executed by a processor.
[0012] One or more embodiments of the present specification provide a data processing method. In the process of processing a problem to be processed, a problem analysis model can be used to determine a target problem and a data storage unit corresponding to the problem to be processed, and a data query statement can be determined using the problem processing model; then, by executing the data query statement, target reference data corresponding to the target problem can be efficiently obtained from the data storage unit, thereby overcoming the defects of manual sorting methods and avoiding the problem of consuming a lot of time and manpower costs for data processing due to excessive data volume; and based on the target reference data, the problem processing result corresponding to the problem to be processed can be quickly determined; the problem processing efficiency is improved; and the problem of low problem processing efficiency is avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram of a data processing solution provided by an embodiment of this specification; Figure 2 It is an application schematic diagram of a data processing method provided by an embodiment of this specification; Figure 3 is a flow chart of a data processing method provided by an embodiment of this specification; Figure 4 is a schematic diagram of a data processing interface in a data processing method provided by an embodiment of this specification; Figure 5 is a processing flow chart of a data processing method provided by an embodiment of this specification; Figure 6 is a flow chart of another data processing method provided by an embodiment of this specification; Figure 7 is a structural schematic diagram of a data processing device provided by an embodiment of this specification; Figure 8 is a schematic diagram of the structure of another data processing device provided by an embodiment of this specification; Fig. 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0014] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0015] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0017] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0018] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. The large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0019] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0020] First, the terms involved in one or more embodiments of this specification are explained.
[0021] NL2SQL: Text-to-SQL refers to the task of converting natural language queries into SQL queries that can be executed on a relational database. The goal is to generate SQL that accurately reflects the user's intentions and ensure appropriate results after execution.
[0022] NL2Code: refers to generating code based on natural language description.
[0023] Metadata: It is information that describes the characteristics of data, including the data's attributes, sources, formats, etc., and is used to help users find, understand, and manage data.
[0024] Index system: The index system refers to an organic whole composed of a number of relatively independent and interrelated statistical indicators that reflect the overall quantitative characteristics of economic phenomena. The index system is composed of multiple interrelated statistical indicators and is organized according to certain logical relationships and internal connections.
[0025] SOP (Standard Operating Procedure): refers to standard operating procedure.
[0026] DataGPT: is an artificial intelligence-based data analysis product. DataGPT uses natural language dialogue to analyze data without the need for SQL. Users only need to enter a simple question to understand the data path, what the data is, why, and what to do, to assist in high-quality decision-making.
[0027] SQL (Structured Query Language): refers to structured query language.
[0028] NoSQL: A JSON-like query language for operating its database, which allows complex queries to filter, sort, and limit results.
[0029] GraphQL: A data query and manipulation language for APIs.
[0030] TableGPT: is a large model that can understand and operate tables using external function commands to process table data.
[0031] Graph data processing model: It is a large model that improves the understanding and generalization ability of the large language model for graph structured data through graph instruction adjustment.
[0032] Geographic data processing model: is a large model that is specifically used to process and generate geospatial data.
[0033] Time series data processing model: is a large model that is specifically used to process and generate time series data.
[0034] With the continuous development of computer technology, a large amount of data will be generated in the process of using computers to perform various data processing operations; and these generated data contain important information that can be used to solve the problems faced by users. At present, in the process of using data to deal with the problems raised by users, manual sorting is often used to find data related to the problem from the data, and the data found is used to deal with the problem; but due to the large amount of data, a lot of time and manpower costs will be consumed for data processing, resulting in low problem handling efficiency.
[0035] For example, Figure 1 is a schematic diagram of a data processing scheme provided by an embodiment of this specification, based on Figure 1It can be seen that the processing of data can be divided into six stages; the first stage is: find the table and ask the data warehouse whether there is an available table. If not, execute the second stage: the data warehouse processes the table; if so, execute the third stage: manually design the analysis ideas and analyze the analysis path. Specifically, in the period when large model intelligence was not widely used, if the user wanted to find data for data analysis, he needed to first find the data warehouse R&D developer to ask which table the original data was, and whether there was a ready-made table that had been cleaned by the data warehouse or the R&D intermediate process; if not, it would need to be processed by the data warehouse. If so, it would be necessary to analyze the ideas and paths for data acquisition. The fourth stage: manually writing SQL queries to obtain data, the fifth stage: writing complex SQL based on various complex functions for data analysis, and the sixth stage: manually summarizing and organizing the analysis results for display, which means that the data users write SQL by themselves or ask developers to write SQL to obtain data, so as to find the data; if the data needs to be analyzed further, the data users with code writing experience will continue to perform data analysis through SQL functions and display the analysis results; in this process, they must be able to write data analysis code and have a threshold for data query and analysis; this leads to relatively low efficiency in data acquisition, query and analysis.
[0036] In response to the above problems, this specification provides an intelligent question-and-answer assistant for data processing, which processes data around: finding / asking / obtaining data, intelligent insight attribution, etc.; however, the intelligent question-and-answer assistant extracts data and analyzes the defined data set, and the entire process is a standard link of NL2SQL+SOP analysis, so there are some standard SOP rules. For example, some methods in the industry mainly require experienced people to manually decompose the problem, input sub-questions step by step to get answers, and manually gain insight into the conclusions, and then the user draws the conclusion. This is a human-assisted exploratory analysis process, rather than a truly productized automated analysis, especially for some fuzzy analysis problems that have not been defined by SOP, which cannot be understood, answered, or given the correct answer.
[0037] Based on this, in this specification, a data processing method is provided. One or more embodiments of this specification also relate to another data processing method, a data processing device, another data processing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0038] See also Figure 2 , Figure 2 A schematic diagram of an application of a data processing method provided according to an embodiment of the present specification is shown. Figure 2 It can be seen that, considering the huge amount of model parameters of the large model and the limited computing resources of the mobile terminal, the data processing method provided in the embodiment of the present application can be applied to Figure 2The application scenarios shown are not limited to these. Figure 2 In the application scenario shown, the large model is deployed in the server 10, and the server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. The client device 20 here may include but is not limited to: smart phones, tablet computers, laptops, PDAs, personal computers, smart home devices, vehicle-mounted devices, etc. The client device 20 can interact with the user through a graphical user interface to implement the call of the large model, thereby implementing the method provided in the embodiments of this specification.
[0039] In an embodiment of the present specification, the system composed of the client device 20 and the server 10 can execute the following steps: the client device 20 executes sending the logistics problem to be processed provided by the user to the server 10; the server 10 executes steps such as problem analysis, data query and data processing; wherein, data analysis refers to using the problem analysis model to analyze the logistics problem to be processed, thereby breaking down the logistics problem to be processed into corresponding target logistics problems, and determining the data storage unit corresponding to the target logistics problem; data query refers to using the problem processing model to determine the logistics data query statement based on the data storage unit and the target logistics problem, and by executing the logistics data query statement, obtaining the target logistics reference data corresponding to the target logistics problem from the data storage unit; data processing refers to using the data processing model to process the target logistics reference data based on the target logistics problem, and determine the problem processing result corresponding to the logistics problem to be processed.
[0040] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device.
[0041] See also Figure 3 , Figure 3 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0042] Step 302: Determine the problem to be processed, and use the problem analysis model to analyze the problem to be processed, determine the target problem and data storage unit corresponding to the problem to be processed, wherein the data storage unit stores reference data obtained from multiple types of data sources.
[0043] The problem to be processed can be understood as a problem that needs to be processed using a problem analysis model or a problem processing model. The problem to be processed can be a data processing problem, a logistics problem to be processed, a medical problem to be processed, etc., which are not specifically limited here. The problem to be processed can be a text type problem, an image type problem, a voice type problem, or a video type problem.
[0044] The target problem can be understood as a target problem obtained after problem analysis on the problem to be processed. The target problem can be a sub-problem obtained by decomposing the problem to be processed. The target problem can be one or at least two.
[0045] The problem analysis model can be understood as a model that can analyze and process the problem to be processed. The problem analysis model can be a large model (LLM), a neural network model, a deep learning model, etc., and is not specifically limited here; for example, the problem analysis model can be DataGPT.
[0046] The data storage unit can be understood as a unit for storing reference data, and the data storage unit can be a server, a database, a local disk, a cloud server, etc. For example, the data storage unit can be a unified indicator library.
[0047] The data source can be understood as a data source that can obtain the reference data. For example, the data source can be a server, client, database and other devices, or the data source can be a document, application, web page and other software modules. No specific restrictions are made here. Different types of data sources refer to different types of data sources. For example, data sources such as servers, databases, and applications can be understood as different types of data sources.
[0048] Reference data can be understood as data related to the problem to be processed and used to process the problem to be processed. For example, when the problem to be processed is a logistics problem to be processed, the reference data can be logistics reference data; the logistics reference data can be used to answer the logistics problem to be processed, for example, the logistics reference data can be the parcel collection volume data of Province A throughout the year, the parcel collection volume data of Province A each month this year, etc. When the problem to be processed is a medical problem to be processed, the reference data can be medical reference data; the medical reference data can be used to answer the medical problem to be processed, for example, the medical reference data can be the number of patients in Hospital A each month this year, etc.
[0049] Specifically, the method can determine the problem to be processed, and after determining the problem to be processed, input the problem to be processed into a problem analysis model; use the problem analysis model to decompose the problem to be processed, and split the problem to be processed into at least two target problems.
[0050] In one or more embodiments provided in this specification, in the process of analyzing the problem to be processed to obtain the target problem, in order to accurately decompose the problem to be processed into at least two target problems, it is necessary to use the problem analysis model to determine the problem processing steps of the problem to be processed, and accurately decompose the problem to be processed into at least two target problems based on the problem processing steps. The specific implementation method is as follows: The problem analysis model is used to analyze the problem to be processed and determine the target problem corresponding to the problem to be processed, including: The problem to be processed is input into the problem analysis model for problem analysis, problem processing steps corresponding to the problem to be processed are determined, and the problem to be processed is decomposed according to the problem processing steps to obtain at least two types of target problems corresponding to the problem to be processed.
[0051] Specifically, the method can input the problem to be processed into a problem analysis model, and perform problem analysis on the problem to be processed in the problem analysis model, so as to determine multiple problem processing steps to be performed to process the problem to be processed; then, the problem analysis model converts the problem to be processed into at least two types of sub-problems to be processed based on the problem processing steps. The sub-problems to be processed can be understood as a part of the problem to be processed, and the processing of the problem to be processed can be achieved by processing multiple types of sub-problems to be processed; finally, the at least two types of sub-problems to be processed are determined as at least two types of target problems, and the processing of the problem to be processed is achieved by executing at least two types of target problems.
[0052] Taking the application of the data processing method provided in this specification in a logistics problem scenario as an example, the data processing method is explained, wherein the problem to be processed may be a logistics problem to be processed, the problem to be processed may be a question asked by a user through natural language, the problem to be processed may be a fuzzy problem or a problem with unclear semantics; the problem analysis model may be a large model.
[0053] Based on this, users ask questions in natural language to obtain the logistics problems to be processed; then, the logistics problems to be processed are input into the big model, and the big model is used to understand and analyze them, sort out the analysis ideas, and form an analysis path; the question-answering task is divided into three stages (i.e., problem processing steps): planning ideas-NL2SQL (natural language to SQL)-NL2Code (natural language to code) for processing; through these three stages, the logistics problems to be processed can be converted into at least two types of sub-problems (i.e., at least two types of target problems), and then by executing at least two types of sub-problems in sequence, self-service product content generation can be achieved.
[0054] In one or more embodiments provided in this specification, in the process of data processing, it is necessary to process multiple types of data to obtain reference data with a unified format and accuracy, so as to facilitate the subsequent use of the reference data to accurately process the problem to be processed; the specific implementation method is as follows.
[0055] Before analyzing the problem to be processed by using the problem analysis model and determining the target problem and the data storage unit corresponding to the problem to be processed, the method further includes: Determine the multiple types of data sources, and obtain to-be-processed reference data of multiple data types from the various types of data sources; According to the data storage format of the data storage unit, the data formats of the reference data to be processed of various data types are adjusted respectively to obtain reference data of multiple data types, and the reference data of the multiple data types are stored.
[0056] The data storage format may be understood as a unified format for data configuration of the data storage unit.
[0057] Using the above example, the data storage unit can be a unified index library, and the reference data to be processed is the logistics reference data to be processed. Based on this, this method adopts a unified index library, so that data indicators from different data sources (i.e., logistics reference data to be processed) are centrally stored in a unified database (unified index library). In addition, these data indicators are managed and maintained, including data cleaning, standardization, index optimization, etc.
[0058] In the process of storing the logistics reference data to be processed into the unified indicator library, the unified indicator library defines a standardized indicator system, including indicator name, calculation logic, data source, etc. (i.e. data storage format), and extracts, converts and loads the logistics reference data to be processed into the unified indicator library.
[0059] It should be noted that the metadata of the logistics reference data, such as name, description, data type, unit, etc., is stored in the unified indicator library, so as to facilitate the management and retrieval of the logistics reference data based on the metadata.
[0060] In one or more embodiments provided in this specification, the data processing method provided in this specification can be applied to the server, and receive the problem to be processed sent by the client, and meet the user's needs for processing the problem to be processed. The specific implementation method is as follows: The determining of the issues to be addressed includes: Receiving a pending issue sent by a client, wherein the pending issue is sent by the client when a user performs a data processing operation based on a data processing interface; The data processing interface can be understood as a human-computer interaction interface for users to input problems to be processed; the data processing interface can be an application program interface or a web page. The data processing operation can be understood as an operation in which a user inputs a problem to be processed to a client, and through the data processing operation, the client and the server can be instructed to perform data processing operations to obtain the problem processing results corresponding to the problem to be processed.
[0061] Specifically, the client in the method can provide a data processing interface for the user, and the user sends the pending issue to the client through the data processing interface; after receiving the pending issue, the client forwards the pending issue to the server.
[0062] After receiving the problem to be processed, the server executes the operation steps recorded in the data processing method for the problem to be processed, thereby obtaining the problem processing result corresponding to the problem to be processed.
[0063] Using the above example, Figure 4 is a schematic diagram of a data processing interface in a data processing method provided in an embodiment of this specification, based on Figure 4 It can be seen that users can input the logistics problem to be processed to the client through the problem input control (input box, button for controlling voice input, etc.) in the data processing interface; the client forwards the logistics problem to be processed to the server for processing, thereby obtaining the answer to the problem corresponding to the logistics problem to be processed (problem processing result); the answer to the problem is the content displayed in the dotted box in the data processing interface.
[0064] Step 304: Utilize the problem processing model to determine a data query statement based on the data storage unit and the target problem, obtain target reference data corresponding to the target problem from the data storage unit by executing the data query statement, and determine a problem processing result corresponding to the problem to be processed based on the target reference data.
[0065] Among them, the problem processing model can be understood as a model that determines and executes the data query language. The problem processing model can be a large model (LLM), a neural network model, a deep learning model, etc., which is not specifically limited here; for example, the problem processing model can be DataGPT.
[0066] It should be noted that the problem processing model and the problem analysis model can be the same model or two different models.
[0067] The data query statement can be understood as a language used to query target parameter data from the data storage unit. The data query statement can be an SQL statement, a NoSQL statement, a GraphQL statement, etc., and no specific limitation is made here.
[0068] The target reference data can be understood as data related to the problem to be processed and used to process the problem to be processed. The target reference data can be used to answer the problem to be processed, or to generate the problem processing result corresponding to the problem to be processed. For example, when the problem to be processed is a logistics problem to be processed, the target reference data can be reference data used to answer the logistics problem to be processed. For example, when the logistics problem to be processed is "the three months with the highest parcel collection volume in Province A this year", the corresponding logistics reference data can be the parcel collection volume data of Province A each month this year, etc. When the problem to be processed is a medical problem to be processed, the target reference data can be the number of patients visiting Hospital A each month this year, etc.
[0069] The problem processing result can be understood as the result obtained after the problem is processed for the problem to be processed. The problem processing result can be the answer to the question, or the data processing result corresponding to the problem to be processed. For example, when the problem to be processed is "What is today's date", the corresponding answer to the question can be "Today is XX / XX / XX"; for example, when the problem to be processed is "The three months with the highest parcel collection volume in Province A this year", the corresponding data processing result can be "The three months of month a, month b, and month c in Province A", and the data processing result is obtained after data processing of the target reference data "The parcel collection volume of Province A each month this year".
[0070] In one or more embodiments provided in this specification, the at least two types of target problems include data acquisition problems and data processing problems; The problem processing model is used to determine a data query statement based on the data storage unit and the target problem, obtain target reference data corresponding to the target problem from the data storage unit by executing the data query statement, and determine a problem processing result corresponding to the problem to be processed based on the target reference data, including: Determine a data query statement based on the data storage unit and the data acquisition problem by using a problem processing model, and acquire target reference data corresponding to the data acquisition problem from the data storage unit by executing the data query statement; The data processing model is used to process the target reference data based on the data processing problem to obtain a problem processing result corresponding to the problem to be processed.
[0071] Among them, the data acquisition problem can be understood as the problem used to obtain target reference data; the data processing problem can be understood as the problem used to process the obtained target reference data, and the data processing includes but is not limited to processing operations such as data sorting, data filtering, and / or data feature extraction.
[0072] The data processing model may be a model for processing target reference data. The data processing model may be a large model (LLM), a neural network model, a deep learning model, etc., which is not specifically limited here; for example, the problem processing model may be DataGPT.
[0073] It should be noted that the problem processing model, problem analysis model and data processing model can be the same model or different models.
[0074] Specifically, the method may input the data acquisition problem and the unit information of the data storage unit into the problem processing model; In the problem processing model, a data query statement corresponding to the data acquisition problem is generated according to the unit information, and by executing the data query statement, target reference data corresponding to the data acquisition problem is acquired from the data storage unit.
[0075] Determining a target data type of the target reference data, and based on the target data type, determining a target data processing model from a plurality of data processing models, wherein the target data processing model is used to process the target reference data of the target data type; The data processing problem and the target reference data are input into the target data processing model. In the data processing model, data processing is performed on the target reference data based on the data processing problem to obtain a problem processing result corresponding to the problem to be processed.
[0076] Using the above example, the question-answering task is divided into three stages through the big model: planning ideas-NL2SQL (natural language to SQL)-NL2Code (natural language to code). Through these three stages, the logistics problem to be processed can be converted into at least two types of target problems (i.e., at least two types of target problems). At least two types of target problems can be classified into data acquisition problems and data processing problems, among which the data acquisition problem corresponds to the planning ideas stage and the NL2SQL stage. The data acquisition problem is used to generate SQL statements for obtaining target reference data. The data processing problem corresponds to the NL2Code stage and is used to process the obtained target reference data to obtain the problem processing result.
[0077] Based on this, after obtaining the data acquisition problem and data processing problem, this method inputs the identification of the unified indicator library and the data acquisition problem (Prompt) into the problem processing model. The problem processing model generates SQL statements corresponding to the unified indicator library based on the identification and data acquisition problem of the unified indicator library; the big model directly connects to the customer's source database (unified indicator library) and executes the generated SQL query. The database executes the SQL query statement and returns the query result (target reference data) to the big model. After the big model processes the result, it obtains the processed result (i.e., the problem processing result).
[0078] In one or more embodiments provided in this specification, the using the problem processing model to determine a data query statement based on the data storage unit and the target problem, and acquiring target reference data corresponding to the target problem from the data storage unit by executing the data query statement, includes: Inputting the target problem and the unit information of the data storage unit into the problem processing model; In the problem processing model, a data query statement corresponding to the target problem is generated according to the unit information, and the target reference data corresponding to the target problem is acquired from the data storage unit by executing the data query statement.
[0079] The unit information may be understood as a unit identifier or a unit type of a data storage unit.
[0080] Continuing with the above example, after obtaining the data acquisition problem and the data processing problem, this method inputs the identifier of the unified indicator library and the data acquisition problem (Prompt) into the problem processing model. The problem processing model generates an SQL statement corresponding to the unified indicator library based on the identifier of the unified indicator library and the data acquisition problem; this facilitates the subsequent smooth query of the target reference data based on the SQL statement corresponding to the unified indicator library, thereby avoiding data query failure caused by query statement mismatch.
[0081] In one or more embodiments provided in this specification, determining the problem processing result corresponding to the problem to be processed based on the target reference data includes: Determining a target data type of the target reference data, and based on the target data type, determining a target data processing model from a plurality of data processing models, wherein the target data processing model is used to process the target reference data of the target data type; The target problem and the target reference data are input into the target data processing model. In the data processing model, data processing is performed on the target reference data based on the target problem to obtain a problem processing result corresponding to the problem to be processed.
[0082] Among them, multiple data processing models can be understood as data processing models for processing different types of target reference data; in the embodiments of this specification, in order to improve data processing efficiency and avoid problems caused by type mismatch between the data processed by the model and the target reference data, this method provides multiple data processing models.
[0083] The data processing model includes but is not limited to: a table data processing model (such as TableGPT), an image data processing model, a map data processing model, a time series data processing model, etc.
[0084] Using the above example, since the target reference data can be any type of data such as table data, map data, terrain data, etc., the process of processing the target reference data can determine a matching data processing model for the target parameter data. Based on this, in the process of processing the target logistics reference data (table type), this method can determine a table data processing model for it, and input the data processing problem and the target logistics reference data into the table data processing model for data sorting, data analysis, etc., to obtain the code corresponding to the problem to be processed (i.e., the problem processing result); wherein, the code refers to the code.
[0085] In one or more embodiments provided in this specification, after determining the problem processing result corresponding to the problem to be processed based on the target reference data, the method further includes: The problem processing result is sent to the client.
[0086] Continuing with the above example, after obtaining the code corresponding to the problem to be processed, the code can be sent to the client and displayed directly in the client.
[0087] One or more embodiments of the present specification provide a data processing method. In the process of processing a problem to be processed, a problem analysis model can be used to determine a target problem and a data storage unit corresponding to the problem to be processed, and a data query statement can be determined using the problem processing model; then, by executing the data query statement, target reference data corresponding to the target problem can be efficiently obtained from the data storage unit, thereby overcoming the defects of manual sorting methods and avoiding the problem of consuming a lot of time and manpower costs for data processing due to excessive data volume; and based on the target reference data, the problem processing result corresponding to the problem to be processed can be quickly determined; the problem processing efficiency is improved; and the problem of low problem processing efficiency is avoided.
[0088] The following combination Figure 5 , taking the application of the data processing method provided in this specification in a logistics scenario as an example, the data processing method is further explained. Figure 5A processing flow chart of a data processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0089] Step 502: The user asks a question in natural language.
[0090] Specifically, the user can input logistics problems to the client in the form of natural language questions through the human-computer interaction interface; the client forwards the logistics problems to the server for processing.
[0091] It should be noted that this method supports the input of fuzzy questions.
[0092] Step 504: Convert the user question into sub-questions to be analyzed, and complete the analysis of the problem path decomposition.
[0093] Specifically, the logistics problem is input into the big model, and through the big model, the analysis ideas are sorted out and the analysis path is formed; thus, the question-answering task is divided into three stages: planning ideas-NL2SQL (natural language to SQL)-NL2Code (natural language to code); Among them, the understanding and analysis of the large model can be divided into two steps, namely understanding the organization's analysis ideas and forming a complete analysis path.
[0094] Through these three stages, the logistics problem to be processed can be converted into at least two types of sub-problems, and then at least two types of sub-problems are executed in sequence to achieve self-service product content generation.
[0095] Step 506: Analyze each sub-problem to obtain analysis results.
[0096] Specifically, after obtaining multiple sub-problems, this method analyzes and processes each sub-problem respectively and obtains corresponding processing results; the specific method is shown in the following manner: 1. Sub-problem analysis - define the metrics.
[0097] This step is used to determine the scope of data to be queried when querying data from the unified indicator library, so as to facilitate the large model to generate accurate SQL statements.
[0098] Specifically, the sub-questions for defining the metrics and the sub-questions for generating SQL statements (prompt) are input into the big model, and the big model is used to determine the data query scope, such as data types, data keywords, etc.
[0099] 2. Sub-problem analysis - data acquisition logic and execution of data acquisition.
[0100] This step is used to generate and execute SQL statements to query corresponding logistics data from the unified indicator library.
[0101] For example, when the logistics problem is "the three months with the highest parcel collection volume in Province A this year", the corresponding retrieved data may be "the parcel collection volume data of Province A for each month this year" and so on.
[0102] 3. Sub-problem analysis-analysis insights.
[0103] This step is used to perform operations such as data filtering and data sorting on the retrieved data, so as to obtain answers corresponding to user questions.
[0104] It should be noted that since the retrieved data can be any type of data such as table data, map data, terrain data, etc., a matching data processing model can be determined during data processing. The data processing model includes but is not limited to: TableGPT, or other data processing models, for example, the other data processing model can be a geographic data processing model, a time series data processing model, a graphic data processing model, etc.; no specific limitation is made here.
[0105] For example, when the retrieved data is tabular data, TableGPT can be used to perform data sorting, data filtering, data analysis, and other processing to obtain answers to user questions.
[0106] Step 508: Summarize and analyze the results.
[0107] Specifically, summary analysis and conclusion refers to analyzing and processing the results output by the large model and determining the content of the results that need to be output.
[0108] The result content can be in two forms: TableQA (table question and answer) and Table2Code (table generation code).
[0109] Among them, table question and answer means: using the answers obtained by analyzing the table data as the output content.
[0110] The table generation code refers to: converting the answer obtained by analyzing the table data into programming code (Code), and using the programming code as output content. The programming code is obtained by converting the answer into code through the large model.
[0111] It should be noted that the programming code can be directly executed on the client side. Therefore, when the programming code is used as output content, after the programming code is sent to the terminal, it can be converted into the display content of the application.
[0112] In addition, the recommended question means: using the big model to generate other similar or related questions based on the user's question, and then synchronously sending the other similar or related questions to the terminal for display to the user.
[0113] Step 510: Output the result content.
[0114] Specifically, after the output content is determined, the output content is sent to the terminal to be displayed to the user.
[0115] Based on the above steps, the data processing method in this specification provides a solution for a data intelligent question-and-answer assistant that combines a large model with metadata in a logistics scenario. This solution implements an intelligent question-and-answer assistant that helps users quickly obtain data through natural language questions and improves data search efficiency; by diagnosing data anomalies and obtaining insight results, the threshold for data analysis is greatly reduced and the efficiency of data analysis is improved. In the logistics scenario, the threshold for use by logistics operators is low, so problems can be solved quickly and conveniently.
[0116] In addition, by providing a unified intelligent data query and data use portal, including data query, anomaly analysis, query, and analysis results, users do not need to define the SOP for each scenario one by one and input manual definition. They only need to access the data source and indicator library, and this solution design uses the big model to analyze and get the results; users can ask questions in natural language, and the big model will perform three stages: question and answer analysis -> query and obtain data NL2SQL and perform analysis and insight -> summarize and analyze the conclusions of each sub-analysis task NL2Code, realizing results such as checking details, asking comparisons, asking trends, looking at rankings, attribution analysis, asking distribution, and asking extreme values.
[0117] Based on the above steps, users ask questions in natural language, and then use the big model to understand and analyze, sort out analysis ideas, form analysis paths, and divide the question-answering tasks into three stages: planning ideas - NL2SQL (natural language to SQL) - NL2Code (natural language to code) for self-service product content generation. This has higher knowledge limitation and scalability, and there is no need to manually define the knowledge question-answering conclusions.
[0118] Through productization, data query and analysis results such as checking details, asking for comparison, asking for trends, viewing rankings, gaining insight into attribution analysis, asking for distribution, and asking for extreme values are realized, meeting the needs of actual applications. This method does not require developers to manually input and output the standards of various scenarios before users use the data, and define them one by one as standard SOPs. It only requires access to the unified indicator library of the database, and users can input questions and ask questions. In the end, there is no big gap in the user experience.
[0119] Moreover, after analyzing the problem with the big model, the analysis data range based on the metadata indicator system is further calculated through the multi-problem big model to obtain the results, thereby greatly reducing the threshold for using data groups to find data, check data, and analyze data.
[0120] See also Figure 6 , Figure 6A flowchart of another data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0121] Step 602: determining a logistics problem to be processed, and using a problem analysis model to analyze the logistics problem to be processed, determining a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, wherein the logistics data storage unit stores logistics reference data obtained from multiple types of data sources; Step 604: Utilize the problem analysis model to determine a data query statement based on the logistics data storage unit and the target logistics problem; obtain target logistics reference data corresponding to the target logistics problem from the logistics data storage unit by executing the data query statement; and determine a logistics problem processing result corresponding to the logistics problem to be processed based on the target logistics reference data.
[0122] One or more embodiments of the present specification provide another data processing method. In the process of processing a logistics problem to be processed, a problem analysis model can be used to determine a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, and a data query statement can be determined using the problem processing model; then, by executing the data query statement, target logistics reference data corresponding to the target logistics problem can be efficiently obtained from the logistics data storage unit, thereby overcoming the defects of the manual sorting method and avoiding the problem of consuming a lot of time and manpower costs for logistics data sorting due to excessive logistics data volume; and, based on the target logistics reference data, the logistics problem processing result corresponding to the logistics problem to be processed can be quickly determined; the efficiency of logistics problem processing is improved; and the problem of low efficiency in logistics problem processing is avoided.
[0123] The above is a schematic scheme of another data processing method of this embodiment. It should be noted that the technical scheme of the other data processing method and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the other data processing method that are not described in detail can all be referred to the description of the technical scheme of the above data processing method.
[0124] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 7 FIG. 1 is a schematic diagram showing the structure of a data processing device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises: The problem determination module 702 is configured to determine the problem to be processed, and use the problem analysis model to perform problem analysis on the problem to be processed, and determine the target problem and the data storage unit corresponding to the problem to be processed, wherein the data storage unit stores reference data obtained from multiple types of data sources; The result determination module 704 is configured to utilize the problem processing model to determine a data query statement based on the data storage unit and the target problem, obtain target reference data corresponding to the target problem from the data storage unit by executing the data query statement, and determine the problem processing result corresponding to the problem to be processed based on the target reference data.
[0125] Optionally, the problem determination module 702 is further configured to: The problem to be processed is input into the problem analysis model for problem analysis, problem processing steps corresponding to the problem to be processed are determined, and the problem to be processed is decomposed according to the problem processing steps to obtain at least two types of target problems corresponding to the problem to be processed.
[0126] Optionally, the at least two types of target problems include data acquisition problems and data processing problems; The result determination module 704 is further configured to: Determine a data query statement based on the data storage unit and the data acquisition problem by using a problem processing model, and acquire target reference data corresponding to the data acquisition problem from the data storage unit by executing the data query statement; The data processing model is used to process the target reference data based on the data processing problem to obtain a problem processing result corresponding to the problem to be processed.
[0127] Optionally, the result determination module 704 is further configured to: Inputting the target problem and the unit information of the data storage unit into the problem processing model; In the problem processing model, a data query statement corresponding to the target problem is generated according to the unit information, and the target reference data corresponding to the target problem is acquired from the data storage unit by executing the data query statement.
[0128] Optionally, the result determination module 704 is further configured to: Determining a target data type of the target reference data, and based on the target data type, determining a target data processing model from a plurality of data processing models, wherein the target data processing model is used to process the target reference data of the target data type; The target problem and the target reference data are input into the target data processing model. In the data processing model, data processing is performed on the target reference data based on the target problem to obtain a problem processing result corresponding to the problem to be processed.
[0129] Optionally, the data processing device further includes a reference data processing module configured to: Determine the multiple types of data sources, and obtain to-be-processed reference data of multiple data types from the various types of data sources; According to the data storage format of the data storage unit, the data formats of the reference data to be processed of various data types are adjusted respectively to obtain reference data of multiple data types, and the reference data of the multiple data types are stored.
[0130] Optionally, the problem determination module 702 is further configured to: Receiving a pending issue sent by a client, wherein the pending issue is sent by the client when a user performs a data processing operation based on a data processing interface; The data processing device further includes a result sending module, which is configured to: The problem processing result is sent to the client.
[0131] One or more embodiments of the present specification provide a data processing device. In the process of processing a problem to be processed, a problem analysis model can be used to determine a target problem and a data storage unit corresponding to the problem to be processed, and a data query statement can be determined using the problem processing model; then, by executing the data query statement, target reference data corresponding to the target problem can be efficiently obtained from the data storage unit, thereby overcoming the defects of the manual sorting method and avoiding the problem of consuming a lot of time and labor costs for data processing due to excessive data volume; and based on the target reference data, the problem processing result corresponding to the problem to be processed can be quickly determined; the problem processing efficiency is improved; and the problem of low problem processing efficiency is avoided.
[0132] The above is a schematic scheme of a data processing device of this embodiment. It should be noted that the technical scheme of the data processing device and the technical scheme of the above data processing method belong to the same concept, and the details of the technical scheme of the data processing device that are not described in detail can be referred to the description of the technical scheme of the above data processing method.
[0133] Corresponding to the above method embodiment, this specification also provides another data processing device embodiment, Figure 8 FIG. 2 shows a schematic diagram of the structure of another data processing device provided by an embodiment of the present specification. Figure 8 As shown, the device comprises: The problem determination module 802 is configured to determine the logistics problem to be processed, and use the problem analysis model to analyze the logistics problem to be processed, determine the target logistics problem and the logistics data storage unit corresponding to the logistics problem to be processed, wherein the logistics data storage unit stores logistics reference data obtained from multiple types of data sources; The result determination module 804 is configured to utilize the problem analysis model to determine a data query statement based on the logistics data storage unit and the target logistics problem, obtain target logistics reference data corresponding to the target logistics problem from the logistics data storage unit by executing the data query statement, and determine the logistics problem processing result corresponding to the logistics problem to be processed based on the target logistics reference data.
[0134] One or more embodiments of the present specification provide another data processing device. In the process of processing a logistics problem to be processed, a problem analysis model can be used to determine a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, and a data query statement can be determined using the problem processing model; then, by executing the data query statement, target logistics reference data corresponding to the target logistics problem can be efficiently obtained from the logistics data storage unit, thereby overcoming the defects of the manual sorting method and avoiding the problem of consuming a lot of time and manpower costs for logistics data sorting due to excessive logistics data volume; and, based on the target logistics reference data, the logistics problem processing result corresponding to the logistics problem to be processed can be quickly determined; the efficiency of logistics problem processing is improved; and the problem of low efficiency in logistics problem processing is avoided.
[0135] The above is a schematic scheme of another data processing device of this embodiment. It should be noted that the technical scheme of the another data processing device and the technical scheme of the another data processing method mentioned above belong to the same concept, and the details of the technical scheme of the another data processing device that are not described in detail can all be referred to the description of the technical scheme of the another data processing method mentioned above.
[0136] Fig. 9 The structure block diagram of a computing device 900 provided according to an embodiment of the present specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0137] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0138] In one embodiment of the present specification, the above components of the computing device 900 and Fig. 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig. 9 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0139] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.
[0140] The processor 920 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of any one of the above-mentioned data processing methods.
[0141] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to any of the above-mentioned data processing method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above-mentioned data processing method embodiments.
[0142] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of any one of the above-mentioned data processing methods when executed by a processor.
[0143] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to any of the above-mentioned data processing method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of any of the above-mentioned data processing method embodiments.
[0144] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of any one of the above-mentioned data processing methods when executed by a processor.
[0145] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of any of the above data processing methods belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of any of the above data processing methods.
[0146] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0147] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0148] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0149] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0150] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Determine a problem to be processed, and use a problem analysis model to analyze the problem to be processed, determine a target problem and a data storage unit corresponding to the problem to be processed, wherein the data storage unit stores reference data obtained from multiple types of data sources; Utilizing the problem processing model, a data query statement is determined based on the data storage unit and the target problem. By executing the data query statement, target reference data corresponding to the target problem is obtained from the data storage unit, and based on the target reference data, a problem processing result corresponding to the problem to be processed is determined.
2. According to the data processing method of claim 1, the step of using the problem analysis model to analyze the problem to be processed and determining the target problem corresponding to the problem to be processed comprises: The problem to be processed is input into the problem analysis model for problem analysis, problem processing steps corresponding to the problem to be processed are determined, and the problem to be processed is decomposed according to the problem processing steps to obtain at least two types of target problems corresponding to the problem to be processed.
3. The data processing method according to claim 2, wherein the at least two types of target problems include data acquisition problems and data processing problems; The problem processing model is used to determine a data query statement based on the data storage unit and the target problem, and by executing the data query statement, target reference data corresponding to the target problem is obtained from the data storage unit, and based on the target reference data, a problem processing result corresponding to the problem to be processed is determined, including: Determine a data query statement based on the data storage unit and the data acquisition problem by using a problem processing model, and acquire target reference data corresponding to the data acquisition problem from the data storage unit by executing the data query statement; The target reference data is processed by using a data processing model based on the data processing problem to obtain a problem processing result corresponding to the problem to be processed.
4. The data processing method according to claim 3, wherein the problem processing model is used to determine a data query statement based on the data storage unit and the target problem, and the target reference data corresponding to the target problem is obtained from the data storage unit by executing the data query statement, comprising: Inputting the target problem and the unit information of the data storage unit into the problem processing model; In the problem processing model, a data query statement corresponding to the target problem is generated according to the unit information, and the target reference data corresponding to the target problem is acquired from the data storage unit by executing the data query statement.
5. The data processing method according to claim 1, wherein determining the problem processing result corresponding to the problem to be processed based on the target reference data comprises: Determining a target data type of the target reference data, and based on the target data type, determining a target data processing model from a plurality of data processing models, wherein the target data processing model is used to process the target reference data of the target data type; The target problem and the target reference data are input into the target data processing model. In the data processing model, data processing is performed on the target reference data based on the target problem to obtain a problem processing result corresponding to the problem to be processed.
6. The data processing method according to claim 1, wherein determining the problem to be processed comprises: Receiving a pending issue sent by a client, wherein the pending issue is sent by the client when a user performs a data processing operation based on a data processing interface; After determining the problem processing result corresponding to the problem to be processed based on the target reference data, the method further includes: The problem processing result is sent to the client.
7. A data processing method, comprising: Determine a logistics problem to be processed, and use a problem analysis model to analyze the logistics problem to be processed, determine a target logistics problem and a logistics data storage unit corresponding to the logistics problem to be processed, wherein the logistics data storage unit stores logistics reference data obtained from multiple types of data sources; Utilizing the problem analysis model, a data query statement is determined based on the logistics data storage unit and the target logistics problem. By executing the data query statement, target logistics reference data corresponding to the target logistics problem is obtained from the logistics data storage unit. Based on the target logistics reference data, a logistics problem processing result corresponding to the logistics problem to be processed is determined.
8. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Query statement generation method and device based on large model, medium and equipment
CN120541191A