GPT model-based data query method and device, storage medium and electronic device
Through the data query method based on the GPT model, users' needs are automatically analyzed, database metadata is matched and structured query statements are generated, which solves the problems of inefficient and low accuracy of manual writing of SQL statements, and improves the efficiency and accuracy of data query.
Patent Information
- Application Number
- CN202311574685.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
Manually writing user requirements into SQL statements has problems of inefficiency and low accuracy, resulting in inaccurate query information retrieved and analyzed in the database.
The data query method based on the GPT model is adopted to analyze the data demand information by calling the preset first GPT model interface to determine the data query attributes; then, based on these attributes, match the metadata information in the database, call the preset second GPT model for query statement analysis, generate a structured query statement in a preset format, and execute it in the database to obtain the target data.
It realizes the automatic writing and execution of SQL statements corresponding to user data requirements, which improves the processing efficiency and accuracy of generating query information.
Smart Images

Figure CN120030031A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data query method, device, storage medium and electronic device based on a GPT model. Background Art
[0002] It is known from relevant technologies that in the field of data query and data analysis, data often exists in a database. Users usually need to manually write their needs into structured query statements (also known as SQL) before they can effectively search and analyze the database to obtain query data or query information corresponding to the user needs.
[0003] However, manually writing user requirements into SQL statements has the problems of low efficiency and low precision, which leads to inaccurate query information obtained through retrieval and analysis in the database. Summary of the invention
[0004] The present application provides a data query method, device, storage medium and electronic device based on the GPT model, which realizes the automatic writing and execution of SQL statements corresponding to the user's data demand information, thereby improving the processing efficiency and accuracy of generating query information.
[0005] The present application provides a data query method based on a GPT model, the method comprising: calling a preset first GPT model interface to perform data query attribute analysis on acquired data demand information, and determining the data query attributes contained in the data demand information; according to the data query attributes, determining metadata information of a data table to be queried that meets preset matching conditions in a preset database; calling a preset second GPT model to perform query statement analysis on the data query attributes and the metadata information, and obtaining a target structured data query statement in a preset format; executing the target structured query statement in the preset database, and obtaining target data corresponding to the data demand information.
[0006] According to a data query method based on a GPT model provided by the present application, the calling of a preset second GPT model to perform query statement analysis on the data query attributes and the metadata information specifically includes: determining text prompt information based on the data query attributes and the metadata information, wherein the text prompt information is used to guide the preset second GPT model to generate prompt information of a structured query statement in a preset format; and inputting the text prompt into the preset second GPT model to implement query statement analysis on the data query attributes and the metadata information.
[0007] According to a data query method based on a GPT model provided by the present application, metadata information of a data table to be queried that meets preset matching conditions is determined in a preset database based on the data query attributes, specifically including: determining metadata information of a data table to be queried that meets preset matching conditions in a preset database through similarity matching based on the data query attributes.
[0008] According to a data query method based on a GPT model provided by the present application, the preset database includes multiple groups of mapping relationships, and the mapping relationship includes the table name of the data table to be queried and the data vector corresponding to the data table to be queried; the metadata information of the data table to be queried that meets the preset matching conditions is determined in the preset database through similarity matching based on the data query attribute, specifically including: vector conversion of the data query attribute to obtain a query attribute vector corresponding to the data query attribute; based on the query attribute vector and the data vector in the mapping relationship, through similarity matching, determine in the preset database a target mapping relationship that matches the query attribute vector; based on the table name of the data table to be queried in the target mapping relationship, obtain the data table to be queried, and based on the data table to be queried, obtain the metadata information of the data table to be queried.
[0009] According to a data query method based on a GPT model provided by the present application, before determining the text prompt information based on the data query attributes and the metadata information, the method also includes: obtaining format requirements for generating a structured query statement in a preset format, and obtaining the data requirement information; determining the text prompt information based on the data query attributes and the metadata information specifically includes: splicing the data requirement information, the data query attributes, the metadata information, and the format requirements to obtain the text prompt information.
[0010] According to a data query method based on a GPT model provided by the present application, in the event that the execution of the target structured query statement fails, the method further includes: obtaining the target structured query statement that failed to be executed; based on the data query attributes, the metadata information, and the target structured query statement that failed to be executed, re-obtaining text prompts, and inputting the re-obtained text prompts into the preset second GPT model to obtain the target data corresponding to the query requirement information.
[0011] The present application also provides a data query device based on the GPT model, the device comprising: a calling module, used to call a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determine the data query attributes contained in the data demand information; a determination module, used to determine the metadata information of the data table to be queried that meets the preset matching conditions in the preset database according to the data query attributes; an analysis module, used to call a preset second GPT model to perform query statement analysis on the data query attributes and the metadata information, and obtain a target structured data query statement in a preset format; a generation module, used to execute the target structured query statement in the preset database to obtain the target data corresponding to the data demand information.
[0012] The present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute and implement any of the above-mentioned GPT model-based data query methods through the computer program.
[0013] The present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, it executes and implements any of the data query methods based on the GPT model as described above.
[0014] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned data query methods based on the GPT model.
[0015] The data query method, device, storage medium and electronic device based on the GPT model provided by the present application perform data query attribute analysis on the acquired data demand information by calling a preset first GPT model interface to determine the data query attributes contained in the data demand information; according to the data query attributes, metadata information of the data table to be queried that meets the preset matching conditions is determined in the preset database; and then a preset second GPT model is called to perform query statement analysis on the data query attributes and metadata information, so as to obtain a target structured data query statement in a preset format, thereby avoiding manual writing of SQL statements, thereby laying the foundation for the automated processing of the user's data demand information, and then executing the target structured query statement in the preset database to obtain the target data corresponding to the data demand information, thereby improving the processing efficiency and accuracy of generating query information. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 It is a schematic diagram of the hardware environment of a data query method based on a GPT model according to an embodiment of the present application;
[0019] Figure 2 It is a flowchart of the data query method based on the GPT model provided by this application;
[0020] Figure 3 It is a flow chart of calling the preset second GPT model to analyze the query statement of the data query attribute and metadata information provided by the present application;
[0021] Figure 4 It is a flow chart of determining metadata information of a data table to be queried that meets preset matching conditions in a preset database based on data query attributes through similarity matching provided by the present application;
[0022] Figure 5 It is a flowchart of building a preset database provided by this application;
[0023] Figure 6 This is a schematic diagram of an application scenario of a data query method based on a GPT model provided in this application;
[0024] Figure 7 It is a structural schematic diagram of a data query device based on a GPT model provided by the present application;
[0025] Figure 8 It is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] According to one aspect of the embodiments of the present application, a data query method based on a GPT model is provided. The data query method based on the GPT model is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned data query method based on the GPT model can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0029] The network may include but is not limited to at least one of the following: wired network, wireless network. The wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0030] In another embodiment, the data query method based on the GPT model provided in this application can be applied to smart home appliances. Among them, smart home appliances refer to home appliances formed by introducing microprocessors, sensor technology, and network communication technology into home appliances, which have the function of automatically sensing the state of residential space, the state of the home appliances themselves, and the state of home appliance services, and can automatically control and receive control instructions from residential users in the home or remotely. It can be understood that smart home appliances are a component of smart homes.
[0031] The data query method based on the GPT model provided in this application can be applied to a user data demand question-and-answer device based on a large language model. During the application process, users can directly describe their needs (corresponding to data demand information) on the device, and the device will automatically convert the demand content into executable SQL statements, execute them in the database, and return the results, so that users can directly obtain data, reduce communication costs, improve efficiency, and reduce errors.
[0032] Figure 2 It is a flowchart of the data query method based on the GPT model provided in this application.
[0033] In order to further introduce the data query method based on the GPT model provided by this application, Figure 2 Provide explanation.
[0034] In an exemplary embodiment of the present application, Figure 2 It can be seen that the data query method based on the GPT model may include steps 210 to 240, and each step will be introduced below.
[0035] In step 210, a preset first GPT model interface is called to perform data query attribute analysis on the acquired data demand information to determine the data query attributes contained in the data demand information.
[0036] In step 220, metadata information of the to-be-queried data table that meets the preset matching condition is determined in a preset database according to the data query attribute.
[0037] In one embodiment, data demand information can be obtained, wherein the data demand information can be understood as a demand put forward by a user, and a suitable result is desired to be matched for the demand. During the application process, the user can put forward a demand, such as "I want to know the number of people who used air conditioners in Beijing last month".
[0038] In one embodiment, the data query attribute may be information that characterizes the characteristics of the data demand information, for example, it may be a keyword of the data demand information, or it may be descriptive characteristic information of the data demand information. In one example, the data demand characteristic information may include the attribute information dimension (also called dimension) of the data demand characteristic information, may include the data characterization result (also called indicator) of the data demand characteristic information, and may also include the characteristic label (also called label) of the data demand characteristic information. During the application process, the preset first GPT model interface may be called to perform data query attribute analysis on the data demand information to obtain the data query attributes contained in the data demand information, for example, information corresponding to the dimensions, indicators and labels in the user requirements may be included.
[0039] Continuing with the example of "I want to know the number of people who used air conditioners in Beijing last month", the attribute information dimensions can be last month and Beijing; the indicator can be the number of users; and the label can be air conditioner.
[0040] In another embodiment, metadata information of a data table to be queried that meets the preset matching condition can be determined in a preset database according to the data query attribute. The data table to be queried is an information table used to generate the final target data. In another embodiment, metadata information of a data table to be queried that meets the preset matching condition can be determined in a preset database through a matching operation with the data query attribute.
[0041] In step 230, a preset second GPT model is called to perform query statement analysis on the data query attributes and metadata information to obtain a target structured data query statement in a preset format.
[0042] In step 240, the target structured query statement is executed in the preset database to obtain target data corresponding to the data requirement information.
[0043] In one embodiment, a preset second GPT model can be called to perform query statement analysis on the data query attributes and metadata information to obtain a target structured data query statement (also known as an SQL statement) in a preset format, and the SQL statement is executed to obtain the target data corresponding to the data demand information. In this embodiment, query statement analysis is performed on the data query attributes and metadata information based on the second GPT model to obtain an SQL statement in a preset format, thereby avoiding manual writing of SQL statements, laying the foundation for the automated processing of the user's data demand information, and then automatically executing the SQL statement corresponding to the data demand information, thereby improving the processing efficiency and accuracy of generating query information.
[0044] The data query method based on the GPT model provided in the present application performs data query attribute analysis on the acquired data demand information by calling a preset first GPT model interface to determine the data query attributes contained in the data demand information; according to the data query attributes, metadata information of the data table to be queried that meets the preset matching conditions is determined in the preset database; and then a preset second GPT model is called to perform query statement analysis on the data query attributes and metadata information to obtain a target structured data query statement in a preset format, thereby avoiding manual writing of SQL statements, thereby laying the foundation for the automated processing of the user's data demand information, and then executing the target structured query statement in the preset database to obtain the target data corresponding to the data demand information, thereby improving the processing efficiency and accuracy of generating query information.
[0045] Figure 3 It is a flowchart provided by the present application for calling a preset second GPT model to perform query statement analysis on data query attributes and metadata information.
[0046] In another exemplary embodiment of the present invention, Figure 3 It can be seen that calling the preset second GPT model to perform query statement analysis on the data query attribute and metadata information may include step 310 and step 320, and each step will be introduced below.
[0047] In step 310, text prompt information is determined based on the data query attribute and metadata information, wherein the text prompt information is used to guide a preset second GPT model to generate prompt information of a structured query statement in a preset format.
[0048] In step 320, the text prompt is input into the preset second GPT model to implement query statement analysis of the data query attributes and metadata information. In one embodiment, the text prompt information is also called Prompt, wherein the text prompt information is used to guide the preset second GPT model to generate prompt information of a structured query statement in a preset format, and the second GPT model is a pre-trained language model for outputting structured query statements.
[0049] Prompt can be understood as a way to start a machine learning model (also known as a GPT model). Prompt can be a piece of text or a sentence used to guide the machine learning model to generate output of a specific type, theme or format. In this embodiment, it can be to guide the second GPT model to output a structured query statement. It should be noted that in the field of natural language processing, Prompt can usually consist of a question or task description, such as "Write me an article about artificial intelligence", "Translate this English sentence into French", etc.
[0050] In one embodiment, the text prompt information can be passed to the second GPT model, so that a target structured query statement in a preset format can be obtained based on the second GPT model. The SQL statement is executed to obtain the target query information corresponding to the query requirement information. Continuing to use the embodiment described above as an example, the query result of "I want to know the number of people who used air conditioners in Beijing last month" can be obtained.
[0051] The preset format can be adjusted according to actual conditions and is not specifically limited in this embodiment. In this embodiment, the data query attributes and metadata information are used to determine the text prompt information and used as the Prompt of the second GPT model. The second GPT model generates the corresponding SQL statement, so that the user's needs can be automatically processed, thereby laying a foundation for efficiently generating query information with high accuracy.
[0052] In another exemplary embodiment of the present application, continue with Figure 2 The embodiment shown is used as an example for explanation. According to the data query attribute, the metadata information of the data table to be queried that meets the preset matching condition is determined in the preset database (corresponding to step 220), which can be implemented in the following manner:
[0053] Based on the data query attributes, metadata information of the data table to be queried that meets the preset matching conditions is determined in the preset database through similarity matching.
[0054] In one embodiment, similarity matching can be performed on data query attributes to determine metadata information of a data table to be queried that meets preset matching conditions in a preset database. Furthermore, based on the data query attributes and metadata information, text prompt information is determined to lay the foundation for automatically obtaining a target structured data query statement in a preset format.
[0055] In another exemplary embodiment of the present application, the preset database may include multiple sets of mapping relationships, and the mapping relationships may include the table name of the data table to be queried and the data vector corresponding to the data table to be queried. The data table to be queried is data with a preset table structure corresponding to the data query information.
[0056] In one embodiment, the preset database may include multiple groups of mapping structures, and the mapping structure may include the table name of the data table to be queried and the data vector corresponding to the data table to be queried. For example, the mapping structure may be expressed as {table name: vector value}.
[0057] In an example, the data table to be queried can be described in conjunction with Table 1.
[0058] Table 1 Statistical data of basic information of air conditioner
[0059] Field Name Field Description Value date Date 20230501 province Province Beijing Municipality open_user_cnt Number of Booted Users 2000 total_use_time Total Usage Duration 10000 ... ... ...
[0060] Among them, the table name of the air conditioner basic information statistical data table can be AC_STAT. It can be understood that the metadata information of the data table to be queried can be the data feature information of the data information in the data table, such as date, province, number of people turned on, total usage time, air conditioning, summary, etc. Among them, date and province can be used as dimensions, number of people turned on and total usage time as indicators, and air conditioning and summary as labels. In the application process, the vector corresponding to the metadata of the aforementioned data table to be queried can be determined and used as the data vector.
[0061] Figure 4 It is a flow chart of determining metadata information of a data table to be queried that meets preset matching conditions in a preset database through similarity matching based on data query attributes provided by the present application.
[0062] In another exemplary embodiment of the present application, Figure 4 It can be seen that determining metadata information of a data table to be queried that meets a preset matching condition in a preset database through similarity matching based on data query attributes may include steps 410 to 430, and each step will be described below.
[0063] In step 410, the data query attribute is vector-converted to obtain a query attribute vector corresponding to the data query attribute.
[0064] In step 420, based on the query attribute vector and the data vector in the mapping relationship, a target mapping relationship matching the query attribute vector is determined in a preset database through similarity matching.
[0065] In step 430, the data table to be queried is obtained based on the table name of the data table to be queried in the target mapping relationship, and metadata information of the data table to be queried is obtained based on the data table to be queried.
[0066] In one embodiment, the data query attribute can be vectorized based on the GPT model to obtain a query attribute vector corresponding to the data query attribute. Then, based on the matching degree between the query attribute vector and the data vector in the mapping relationship, a target mapping relationship matching the query attribute vector is determined in the preset database. In this embodiment, the vector of the user's needs (corresponding to the query attribute vector) is matched with the preset database, so that the user's needs can be understood more accurately.
[0067] Since the target mapping relationship includes the table name of the data table to be queried, the data table to be queried that matches the data query attribute can be found based on the table name of the data table to be queried in the target mapping relationship, and the metadata of the data table to be queried can be obtained, which further lays the foundation for generating text prompts based on metadata.
[0068] In another example, let's continue with the user demand "I want to know the number of people who used air conditioners in Beijing last month". You can use the query attribute vector to search in the preset database and return data tables with high similarity. In this example, the data tables that may be returned include "AC_STAT" and "PROVINCE_INFO" (province dimension table).
[0069] Query the metadata of the data table returned in the previous step, for example: AC_STAT(date-date-20230501,province-province-Beijing,open_user_cnt-number of people turned on-2000,total_use_time-total usage time-10000...); PROVINCE_INFO(id-ID,province_name-province name...).
[0070] Furthermore, the basic requirements, table metadata, user requirements, and GPT parsing results can be spliced to generate a GPT prompt, where the GPT prompt can be expressed in the following ways:
[0071] Your task is to generate MySQL 5.7 SQL code as required based on user needs;
[0072] The SQL you generate needs to meet the following requirements:
[0073] Use only the provided table names and English field names, do not use other columns;
[0074] User demand: I want to know the number of people who used air conditioning in Beijing last month;
[0075] Dimensions: Last month, Beijing;
[0076] Indicator: Number of users;
[0077] Tags: air conditioning;
[0078] Table name and field name (field name-field comment):
[0079] AC_STAT(date-date-20230501,province-province-Beijing,open_user_cnt-number of users-2000,total_use_time-total usage time-10000...)
[0080] PROVINCE_INFO(id-ID,province_name-province name...).
[0081] The returned result should be a SQL statement starting with SELECT, meeting all the above requirements, and be clear, concise, and efficient.
[0082] In another example, GPT returns the result (corresponding to the target query information):
[0083] SELECT open_user_cnt FROMAC_STAT WHERE DATE_FORMAT(date,'%Y-%m')=DATE_FORMAT(CURRENT_DATE-INTERVAL 1MONTH,'%Y-%m');
[0084] Execute the SQL and return the SQL result, which is "2000" in this example.
[0085] Figure 5 It is a flowchart of building a preset database provided by this application.
[0086] The following will be combined Figure 5 Describes the process of building a provisioning database.
[0087] In an exemplary embodiment of the present application, the preset database may be pre-set and used to store the relevant description information vectors in the data table. Figure 5 It can be known that building a preset database may include steps 510 to 550, and each step will be introduced below.
[0088] In step 510, a data table is obtained.
[0089] In step 520, dimensions and indicators of the data table are labeled according to the information of the data table.
[0090] In step 530, the annotation information is vectorized and stored in a vector database.
[0091] In step 540, it is determined whether there are other data tables.
[0092] In step 550, if no other data tables exist, the process ends.
[0093] If there are other data tables, the process returns to step 510 and repeats the process.
[0094] During the application process, information of a data table can be obtained, wherein the data table can be shown in Table 1. Furthermore, according to the data information in the data table, the dimensions, indicators and other information that can be produced by the data table are marked and the data table is labeled. The obtained dimensions, indicators, labels and other information are then vectorized using a text conversion vector model, such as Word2Vec, BERT, etc. Thus, a mapping structure can be obtained, represented as {"AC_STAT":[0,1,0,1,0,1....]}. And multiple mapping structures are stored in a preset database. When it is determined that there are other data tables, steps 510 to 540 can be repeated. Thus, a more complete preset database can be obtained.
[0095] In another exemplary embodiment of the present application, continue with Figure 3 The above embodiment is used as an example to illustrate that before determining the text prompt information (step 310) based on the data query attribute and metadata information, the data query method based on the GPT model may further include the following steps:
[0096] Obtaining format requirements for generating structured query statements in a preset format, and obtaining data requirement information;
[0097] Data demand information, data query attributes, metadata information, and format requirements are spliced and processed to obtain text prompt information.
[0098] In one embodiment, the format requirements for generating a structured query statement in a preset format and the data requirement information can be obtained. Furthermore, the text prompt information is obtained based on the data requirement information, the data query attributes, the format requirements, and the metadata information. By using the metadata information to determine the text prompt, the corresponding SQL statement can be automatically generated based on the second GPT model, avoiding the manual writing of the SQL statement, thereby laying the foundation for the automatic processing of the user's data requirement information.
[0099] In another example, the format requirement may include the format of the generated SQL, the format requirement of the returned result (corresponding to the target data), etc. In this embodiment, the format requirement is not specifically limited and can be adjusted according to the actual situation.
[0100] In another exemplary embodiment of the present application, continue with Figure 2 Taking the above embodiment as an example, when the target structured query statement fails to be executed, the data query method based on the GPT model may further include the following steps:
[0101] Get the target structured query statement that failed to execute;
[0102] Based on the data query attributes, metadata information, and the target structured query statement that failed to execute, the text prompt is obtained again, and the obtained text prompt is input into the preset second GPT model to obtain the target data corresponding to the query requirement information. In other words, after the text prompt is obtained again, the above steps 230 and 240 are executed again.
[0103] In this embodiment, the text prompt is obtained again in combination with the target structured query statement that failed to execute, which can ensure that the newly generated text prompt can provide a prompt for eliminating the target structured query statement that failed to execute, laying a foundation for the second GPT model to obtain a more accurate target structured query statement.
[0104] Figure 6 It is a schematic diagram of the application scenario of the data query method based on the GPT model provided in this application.
[0105] In order to further introduce the data query method based on the GPT model provided by this application, Figure 6 Provide explanation.
[0106] In an exemplary embodiment of the present application, Figure 6 It can be seen that the data query method based on the GPT model may include steps 610 to 690, and each step will be introduced below.
[0107] In step 610, the user makes a request.
[0108] In step 620, GPT is used to analyze the dimensions, indicators, and labels in the user requirements.
[0109] In step 630, a match is performed in the vector database to obtain a corresponding data table.
[0110] In step 640, the original data of the data table is obtained.
[0111] In step 650, user requirements, GPT parsing results, table metadata, and other necessary information are combined to generate a GPT prompt.
[0112] In step 660, the GPT is used to generate SQL according to the Prompt.
[0113] In step 670, the SQL is executed.
[0114] In step 680, the execution is successful and the result is returned.
[0115] In step 690 , the execution fails, error information is collected, and the process returns to step 650 .
[0116] In one embodiment, the user can input data requirements (corresponding to data requirement information). Furthermore, the GPT model can be used to parse the data requirement information and obtain the dimensions, indicators and label information required by the user. The parsed dimensions, indicators and label information are vectorized using the text conversion vector model to obtain the query attribute vector. Then, a match is performed in the vector database to obtain a data table with a high similarity, and metadata information in the data table is obtained, such as table name, table comment, table field, field comment and field sample.
[0117] Furthermore, the user input content (corresponding to the data requirement information), the GPT parsing results (corresponding to the dimensions, indicators and label information required by the user), the metadata information of the data table, and other necessary information can be spliced to obtain the GPT prompt. Among them, other necessary information may include the format of generating SQL, the requirements for returning results, etc.
[0118] Then pass the prompt information to the GPT model, wait for the SQL to be returned, and execute the SQL to obtain the return result (corresponding to the target data).
[0119] In yet another example, if the SQL execution fails, the failed SQL and log information are obtained and used as part of the Prompt and re-executed in step 650 .
[0120] The query information generation method provided by the present application understands and locates user needs through vectorization of user needs and database matching, thereby improving the accuracy of demand analysis. This is a significant optimization for the problem of low efficiency and low accuracy in manual understanding of user needs in the prior art. Furthermore, the automation of this understanding method can greatly increase the speed of processing user needs, thereby increasing processing efficiency. In addition, the present application uses a large language model to convert metadata information into SQL statements, avoiding the process of manually writing SQL in traditional methods, greatly reducing human errors and time consumption. This solution provides a new solution to the problem of error-prone and low efficiency of manually writing SQL statements in the prior art.
[0121] The data query device based on the GPT model provided in the present application is described below. The data query device based on the GPT model described below and the data query method based on the GPT model described above can be referenced to each other.
[0122] Figure 7 It is a structural diagram of a data query device based on a GPT model provided in this application.
[0123] In an exemplary embodiment of the present application, Figure 7It can be seen that the data query device based on the GPT model may include a calling module 710, a determining module 720, an analyzing module 730, and a generating module 740, and each module will be introduced below.
[0124] The calling module 710 may be configured to call a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information to determine the data query attribute contained in the data demand information;
[0125] The determination module 720 may be configured to determine metadata information of a to-be-queried data table that meets a preset matching condition in a preset database according to a data query attribute;
[0126] The analysis module 730 may be configured to call a preset second GPT model to perform query statement analysis on the data query attribute and metadata information to obtain a target structured data query statement in a preset format;
[0127] The generating module 740 may be configured to execute a target structured query statement in a preset database to obtain target data corresponding to the data requirement information.
[0128] In another exemplary embodiment of the present application, the analysis module 730 may implement calling the preset second GPT model to perform query statement analysis on the data query attribute and metadata information in the following manner:
[0129] Based on the data query attribute and metadata information, determine text prompt information, wherein the text prompt information is used to guide the preset second GPT model to generate prompt information of a structured query statement in a preset format;
[0130] The text prompt is input into the preset second GPT model to implement query statement analysis on data query attributes and metadata information.
[0131] In another exemplary embodiment of the present application, the determination module 720 may determine metadata information of a to-be-queried data table that meets a preset matching condition in a preset database according to a data query attribute in the following manner:
[0132] Based on the data query attributes, metadata information of the data table to be queried that meets the preset matching conditions is determined in the preset database through similarity matching.
[0133] In another exemplary embodiment of the present application, the preset database includes multiple groups of mapping relationships, and the mapping relationships include table names of data tables to be queried and data vectors corresponding to the data tables to be queried;
[0134] The determination module 720 can determine the metadata information of the to-be-queried data table that meets the preset matching condition in the preset database by similarity matching based on the data query attribute in the following manner:
[0135] Perform vector conversion on the data query attribute to obtain a query attribute vector corresponding to the data query attribute;
[0136] Based on the query attribute vector and the data vector in the mapping relationship, a target mapping relationship matching the query attribute vector is determined in a preset database through similarity matching;
[0137] Based on the table name of the data table to be queried in the target mapping relationship, the data table to be queried is obtained, and based on the data table to be queried, metadata information of the data table to be queried is obtained.
[0138] In another exemplary embodiment of the present application, the determination module 720 may also be configured to:
[0139] Obtain format requirements for generating structured query statements in a preset format, and obtain data requirement information;
[0140] The determination module 720 may determine the text prompt information based on the data query attribute and metadata information in the following manner:
[0141] Data demand information, data query attributes, metadata information, and format requirements are spliced and processed to obtain text prompt information.
[0142] In another exemplary embodiment of the present application, when the target structured query statement fails to be executed, the generating module 740 may also be configured to:
[0143] Get the target structured query statement that failed to execute;
[0144] Based on data query attributes, metadata information, and the target structured query statement that failed to execute, text prompts are obtained again, and the obtained text prompts are input into the preset second GPT model to obtain the target data corresponding to the query requirement information.
[0145] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute a data query method based on the GPT model, the method comprising: calling a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determining the data query attribute contained in the data demand information; according to the data query attribute, determining metadata information of the data table to be queried that meets the preset matching condition in the preset database; calling a preset second GPT model to perform query statement analysis on the data query attribute and the metadata information, and obtaining a target structured data query statement in a preset format; executing the target structured query statement in the preset database to obtain target data corresponding to the data demand information.
[0146] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0147] On the other hand, the present application also provides a computer program product, which includes a computer program, and the computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data query method based on the GPT model provided by the above methods, and the method includes: calling a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determine the data query attributes contained in the data demand information; according to the data query attributes, determine the metadata information of the data table to be queried that meets the preset matching conditions in the preset database; calling a preset second GPT model to perform query statement analysis on the data query attributes and the metadata information, and obtain a target structured data query statement in a preset format; execute the target structured query statement in the preset database to obtain the target data corresponding to the data demand information.
[0148] On the other hand, the present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, the data query method based on the GPT model provided by the above-mentioned methods is executed, and the method includes: calling a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determine the data query attributes contained in the data demand information; according to the data query attributes, determine the metadata information of the data table to be queried that meets the preset matching conditions in the preset database; calling a preset second GPT model to perform query statement analysis on the data query attributes and the metadata information, and obtain a target structured data query statement in a preset format; executing the target structured query statement in the preset database to obtain the target data corresponding to the data demand information.
[0149] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0150] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data query method based on the GPT model, It is characterized in that The method comprises: Calling a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determining the data query attribute contained in the data demand information; According to the data query attribute, metadata information of the data table to be queried that meets the preset matching condition is determined in the preset database; Calling a preset second GPT model to perform query statement analysis on the data query attribute and the metadata information to obtain a target structured data query statement in a preset format; The target structured query statement is executed in the preset database to obtain target data corresponding to the data demand information.
2. According to the data query method based on the GPT model in claim 1, It is characterized in that The calling of the preset second GPT model to perform query statement analysis on the data query attribute and the metadata information specifically includes: Based on the data query attribute and the metadata information, determining text prompt information, wherein the text prompt information is used to guide the preset second GPT model to generate prompt information of a structured query statement in a preset format; The text prompt is input into the preset second GPT model to implement query statement analysis on the data query attribute and the metadata information.
3. The data query method based on the GPT model according to claim 1, It is characterized in that Determining metadata information of a data table to be queried that meets a preset matching condition in a preset database according to the data query attribute specifically includes: Based on the data query attribute, metadata information of the data table to be queried that meets the preset matching condition is determined in the preset database through similarity matching.
4. According to the data query method based on the GPT model of claim 3, It is characterized in that The preset database includes multiple groups of mapping relationships, and the mapping relationships include table names of data tables to be queried and data vectors corresponding to the data tables to be queried; The step of determining metadata information of a data table to be queried that meets a preset matching condition in a preset database through similarity matching based on the data query attribute specifically includes: Performing vector conversion on the data query attribute to obtain a query attribute vector corresponding to the data query attribute; Based on the query attribute vector and the data vector in the mapping relationship, determining a target mapping relationship matching the query attribute vector in the preset database through similarity matching; Based on the table name of the data table to be queried in the target mapping relationship, the data table to be queried is obtained, and based on the data table to be queried, metadata information of the data table to be queried is obtained.
5. The data query method based on the GPT model according to claim 2, It is characterized in that Before determining text prompt information based on the data query attribute and the metadata information, the method further includes: Obtaining format requirements for generating a structured query statement in a preset format, and obtaining the data requirement information; The determining of text prompt information based on the data query attribute and the metadata information specifically includes: The data demand information, the data query attribute, the metadata information, and the format requirement are concatenated to obtain the text prompt information.
6. The data query method based on the GPT model according to claim 1, It is characterized in that In the case where the target structured query statement fails to be executed, the method further includes: Get the target structured query statement that failed to execute; Based on the data query attributes, the metadata information, and the target structured query statement that failed to execute, text prompts are obtained again, and the obtained text prompts are input into the preset second GPT model to obtain the target data corresponding to the query requirement information.
7. A data query device based on the GPT model, It is characterized in that The device comprises: A calling module, used to call a preset first GPT model interface to perform data query attribute analysis on the acquired data demand information, and determine the data query attribute contained in the data demand information; A determination module, used to determine metadata information of a data table to be queried that meets a preset matching condition in a preset database according to the data query attribute; An analysis module, used for calling a preset second GPT model to perform query statement analysis on the data query attribute and the metadata information to obtain a target structured data query statement in a preset format; A generating module is used to execute the target structured query statement in the preset database to obtain target data corresponding to the data demand information.
8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored program, wherein the program, when running, executes the data query method based on the GPT model described in any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, It is characterized in that A computer program is stored in the memory, and the processor is configured to execute the GPT model-based data query method according to any one of claims 1 to 6 through the computer program.
10. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the GPT model-based data query method as described in any one of claims 1 to 6 is implemented.