Data analysis method and device, storage medium and program product

By generating and correcting the closed-loop link of database operation statements, and using pre-trained language models to automatically tune, the problem of insufficient self-repair mechanism in the traditional data analysis process is solved, and an efficient and robust data analysis process is achieved.

CN120030039APending Publication Date: 2025-05-23ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510193838.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The lack of effective self-repair mechanisms in traditional data analysis processes leads to complex writing of database operation statements, frequent problems such as syntax errors, logical errors or mismatch in database table structures, increasing labor costs and response delays.

Method used

A data analysis method is proposed, by receiving data analysis tasks, generating prompt words and inputting pre-trained language models, and generating database operation statements that meet the task. If the execution error is performed, use the pre-trained language model to refer to the error information correction statement, and form a closed-loop link to automatically tune the database operation statement.

Benefits of technology

Significantly reduce the execution failure rate caused by syntax errors, mismatch of table structures or logical defects, and realize a highly robust data analysis process without manual intervention and avoid resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030039A_ABST
    Figure CN120030039A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a data analysis method and device, a storage medium and a program product. The data analysis method comprises the steps of receiving a data analysis task for a target database; generating a cue word according to the data analysis task and data table information in the target database, and inputting the cue word into a pre-training language model, so that the pre-training language model generates a database operation statement conforming to the data analysis task; triggering an operation statement submission event based on the database operation statement to enable the target database to execute the received database operation statement; if execution error information returned by the target database is received, the pre-training language model is used for referring to the execution error information to modify the database operation statement, the modified database operation statement is obtained, and the operation statement submission event is triggered again based on the modified database operation statement; and if the target database is successfully executed, outputting a data analysis result of the data analysis task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of data analysis technology, and in particular, to a data analysis method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In the traditional data analysis process, it is usually necessary to write database operation statements according to actual needs to operate the database. Due to the complexity of writing database operation statements, problems such as syntax errors, logical errors, or database table structure mismatches frequently occur, and related technologies lack an effective self-repair mechanism. When a database operation statement fails to execute, the database system can only return basic error information, and subsequent troubleshooting, debugging, and modification rely entirely on manual intervention. Analysts need to repeatedly verify the cause of the error and manually adjust the database operation statements, which not only increases labor costs, but may also cause response delays. Especially when faced with large-scale databases or complex analysis scenarios, the technical defects of traditional methods are further amplified, seriously restricting the real-time nature of data analysis. Summary of the invention

[0003] In view of this, one or more embodiments of the present specification provide a data analysis method, an electronic device, a computer-readable storage medium, and a computer program product.

[0004] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions:

[0005] According to a first aspect of one or more embodiments of this specification, a data analysis method is proposed, comprising:

[0006] Receive data analysis tasks for the target database;

[0007] Generate prompt words according to the data analysis task and the data table information in the target database, and input the prompt words into a pre-trained language model so that the pre-trained language model generates a database operation statement that meets the data analysis task;

[0008] triggering an operation statement submission event based on the database operation statement, so that the target database executes the received database operation statement;

[0009] If execution error information returned by the target database is received, modify the database operation statement by using the pre-trained language model with reference to the execution error information to obtain a modified database operation statement, and trigger the operation statement submission event again based on the modified database operation statement;

[0010] If the target database is executed successfully, the data analysis result of the data analysis task is output.

[0011] According to a second aspect of the embodiments of this specification, an electronic device is provided, including:

[0012] processor;

[0013] a memory for storing processor-executable instructions;

[0014] When the processor executes the executable instructions, it is used to implement the method described in the first aspect.

[0015] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0017] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:

[0018] In an embodiment of the present specification, a data analysis task is combined with data table information of a target database to generate prompt words, and a pre-trained language model is driven to generate database operation statements, which are submitted to the target database for execution. When the target database returns an execution error message, the pre-trained language model can be driven to refer to the execution error message to correct the database operation statement, thereby forming a closed-loop link of "database operation statement generation-database operation statement execution-execution error information feedback-database operation statement correction", thereby significantly reducing the execution failure rate caused by syntax errors, table structure mismatches or logical defects; at the same time, the closed-loop link can realize automatic tuning of operation statements without human intervention, which can avoid the waste of resources caused by manual trial and error in traditional methods, and realize a highly robust data analysis process.

[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the architecture of a data analysis service system provided by an exemplary embodiment.

[0021] Figure 2 It is a flow chart of a data analysis method provided by an exemplary embodiment.

[0022] Figure 3 The present invention is a schematic diagram of a processing flow of a database operation statement error provided by an exemplary embodiment.

[0023] Figure 4 It is a schematic diagram of a data analysis process provided by an exemplary embodiment.

[0024] Figure 5 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment. DETAILED DESCRIPTION

[0025] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0026] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0028] Here is an explanation of the relevant terms:

[0029] Pre-trained language models, such as Large Language Model (LLM), refer to artificial intelligence models based on deep learning technology, especially those trained with a large corpus, designed to understand and generate text similar to human language, with strong natural language understanding and generation capabilities. The goal of LLM is to achieve a variety of applications through natural language processing capabilities, such as text generation, translation, summarization, question-answering, and dialogue systems, thereby helping to improve the efficiency and automation of human-computer interaction.

[0030] Figure 1 FIG. 1 is a schematic diagram of the architecture of a data analysis service system provided by an exemplary embodiment. Figure 1As shown, the system may include a server 11, several terminals 12, a first database 13 belonging to the server 11, and a second database 14 belonging to the terminals 12. Figure 1 Terminal A and terminal B are used as examples.

[0031] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server carried by a host cluster. During operation, the server 11 may run a server-side program of a data analysis application to realize a corresponding data analysis service platform.

[0032] Different terminals 12 are used to represent different data analysis demand parties, such as individuals, enterprises or institutions, etc., but not limited to this. Terminals 12 include but are not limited to PCs (Personal Computers), mobile phones, tablet devices, laptops, desktop computers, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.) or other devices. During operation, the terminal 12 can run the client-side program of the data analysis application and implement it as a corresponding client. Among them, the client-side program of the data analysis application can be started and run on the terminal 12. The client-side program can be a native application installed on the terminal, or the client-side program can be a small program, a quick application or other similar forms. Of course, when using web page technologies such as HTML5 or similar, the relevant functions can be implemented through the page displayed by the browser. The browser here can be an independent browser application or a browser module embedded in certain applications.

[0033] In a data analysis scenario, terminal A uploads the data table it generates to server 11. After receiving the data table, server 11 stores it in the first database 13 maintained by itself. At this time, the management right of the first database 13 belongs entirely to server 11, and the storage of the data table and subsequent analysis work are all the responsibility of server 11. Terminal A interacts with server 11 through the client application and submits the data analysis task. Server 11 performs further analysis based on the data analysis task and the data table belonging to terminal A stored in the first database 13 to obtain the data analysis result and provide it to terminal A. Terminal A views the analysis result or performs subsequent operations through the client program.

[0034] In another data analysis scenario, terminal B maintains its own independent second database 14, which contains data tables required for data analysis. Terminal B provides the login information (such as user name, password) of the second database 14 and the database URL or access interface to the server 11. The server 11 establishes a connection with the second database 14 based on the provided login information and URL, and can remotely access the second database 14 of terminal B. Terminal B interacts with the server 11 through the client application and submits a data analysis task. The server 11 performs further analysis based on the data analysis task and the data tables stored in the second database 14 to obtain the data analysis results and provides them to terminal B. Terminal B views the analysis results or performs subsequent operations through the client program.

[0035] In some embodiments, based on the problem that traditional data analysis processes lack effective self-repair mechanisms, please refer to Figure 2 , the embodiment of this specification provides a data analysis method, which can be Figure 1 The method is performed by the server 11 in the embodiment of the present invention, and includes:

[0036] In S201, a data analysis task for a target database is received.

[0037] In this step, the server receives a data analysis task from a terminal (such as terminal A or terminal B) or other data analysis demander. This task may include data query, processing or analysis requirements for a target database (such as the target database of terminal A is the first database maintained by the server, and the target database of terminal B is the second database maintained by itself). Through this step, the server clarifies the specific goals of the data analysis task, such as calculation, aggregation, filtering or other specific data processing operations.

[0038] In S202, prompt words are generated according to the data analysis task and the data table information in the target database, and the prompt words are input into the pre-trained language model so that the pre-trained language model generates database operation statements that meet the data analysis task.

[0039] In this step, the server generates prompt words according to the specific requirements of the data analysis task and the data table information in the target database, and inputs them into the pre-trained language model. The pre-trained language model generates database operation statements (such as SQL queries, data updates or insert operations, etc.) that meet the intention of the data analysis task based on the prompt words. In this way, the server automatically generates database operation statements without human intervention, which can eliminate the complexity of manually writing statements and reduce human errors (such as grammatical errors, logical errors, etc.), thereby improving execution efficiency and accuracy.

[0040] In S203, an operation statement submission event is triggered based on the database operation statement, so that the target database executes the received database operation statement.

[0041] In this step, the server 11 submits the generated database operation statement to the target database for execution. There are two execution situations, one is that an error occurs during the execution of the database operation statement (corresponding to S204), and the other is that the database operation statement is executed successfully (corresponding to S205).

[0042] In S204, if an execution error message returned by the target database is received, the database operation statement is modified using the pre-trained language model with reference to the execution error message to obtain a modified database operation statement, and the operation statement submission event is triggered again based on the modified database operation statement.

[0043] In this step, if an error occurs during the execution of a database operation statement (such as a syntax error, data mismatch, etc.), the server inputs the received execution error information into the pre-trained language model, so that the pre-trained language model can refer to the execution error information and intelligently modify or optimize the database operation statement to obtain a modified database operation statement. Then, the server resends the modified operation statement to the target database for execution. This step effectively solves the problem of the lack of a self-repair mechanism in traditional data analysis. It automatically repairs errors through a pre-trained language model, reduces the need for manual intervention, improves the robustness and adaptability of the data analysis service system, ensures the smooth progress of data analysis tasks, and reduces the time and cost of troubleshooting.

[0044] In some possible implementations, in order to prevent falling into an infinite retry state, a preset number of retries can be set, that is, see Figure 3After triggering the operation statement submission event based on the database operation statement (S301), the server determines whether the execution error information returned by the target database is received (S302); if the execution error information returned by the target database is received, the server further determines whether the number of executions of the target database reaches the preset number of retries (S303). If the execution error information returned by the target database is received and the number of executions of the target database does not reach the preset number of retries, the database operation statement generated last time is modified by referring to the execution error information using the pre-trained language model to obtain the modified database operation statement, and the operation statement submission event is triggered again based on the modified database operation statement (S304). If the execution error information returned by the target database is received and the number of executions of the target database reaches the preset number of retries, a prompt message of data analysis failure is output (S305). By setting the preset number of retries, this embodiment can enhance fault tolerance in the face of temporary failures, improve the execution success rate of tasks, and avoid endless retries to consume the computing and storage resources of the server; when the execution error exceeds the preset number of retries, a prompt message of "data analysis failure" can be output to ensure that the user obtains clear feedback.

[0045] The specific value of the preset number of retries may be set according to the actual application scenario, and this embodiment does not impose any limitation on this.

[0046] In S205, if the target database is executed successfully, the data analysis result of the data analysis task is output.

[0047] In this step, if the target database successfully executes the database operation statement submitted by the server, the server will generate and output the data analysis results of the data analysis task based on the execution result. At this point, the final data analysis results of the data analysis task can be provided to the data analysis demander, such as data summary, chart display, statistical data, etc.

[0048] It is understandable that this embodiment does not impose any restrictions on the specific output method of the data analysis results, and can be specifically set according to the actual application scenario, such as text display, display in the form of charts, or display in the form of audio and video, etc.

[0049] Exemplarily, the server may receive the execution result returned by the target database when the execution is successful; and then convert the execution result returned by the target database into a visual output by converting natural language to icons, so as to provide users with a more intuitive data information presentation.

[0050] In one possible implementation, the server can use the pre-trained language model to process the execution result, generate rendering code that can render the execution result as a visual chart, and then call the rendering component to execute the rendering code, generate the visual chart and output it. This embodiment uses the pre-trained language model to automatically generate rendering code that meets the needs of the analysis task, eliminating the tedious work of manually writing rendering code.

[0051] Specifically, the server can input the description information of the data analysis task, the execution results returned by the target database, and the database operation statements corresponding to the execution results into the pre-trained language model. The pre-trained language model first understands the description information of the data analysis task, clarifies the chart types that need to be displayed (such as bar charts, pie charts, line charts, etc.), and then matches the execution results returned by the target database with the display requirements of the data analysis task, confirms which fields in the execution results should be used for rendering, and finally generates appropriate rendering code based on the display requirements of the data analysis task. The generated rendering code is, for example, a Chart.js configuration object, an ECharts object, or other commonly used visualization framework code. This implementation method automatically generates appropriate rendering code by inputting data analysis tasks, execution results, and their corresponding database operation statements into the pre-trained language model, and then generates and displays visualization charts through rendering components. This can effectively improve the automation of data visualization, reduce development workload, and ensure that the final output charts meet user needs, thereby improving the intelligence and user experience of the system.

[0052] Among them, the description information of the data analysis task can provide the display requirements of the data analysis task, and help the pre-trained language model understand how to match the execution results with the display requirements of the data analysis task. For example, the data analysis task can be "displaying the sales of each product" or "analyzing the user growth trend by region". The execution result returned by the target database contains the actual query data, which is the data basis required to generate a chart. For example, the execution result can be a result set containing multiple data fields (such as product name, sales, time, etc.). The database operation statement corresponding to the execution result can help the pre-trained language model confirm the source and structure of the data. In this embodiment, the efficiency and accuracy of data analysis can be improved by automatically generating database operation statements, intelligently repairing execution errors, and timely outputting analysis results. Compared with the traditional method of manually repairing database operation statements, it can effectively reduce manual intervention, improve the execution success rate of analysis tasks, and ensure the efficiency and accuracy of data analysis services.

[0053] In some embodiments, when the target database is executed successfully, the server receives the execution result returned by the target database when the execution is successful. In order to improve the execution accuracy of the data analysis task, a pre-trained language model can be used to determine whether the execution result returned by the target database meets the expected goal of the data analysis task.

[0054] Exemplarily, the server may input the description information of the data analysis task, the execution result returned by the target database, and the database operation statement corresponding to the execution result into the pre-trained language model, so that the pre-trained language model can analyze the execution result returned by the database and the corresponding database operation statement to determine whether it meets the expected goal of the data analysis task. For example, if the expected goal of the data analysis task is to "query the maximum value of a certain field", but the execution result does not return the maximum value or the returned data format is incorrect, the pre-trained language model will identify the problem.

[0055] If the output of the pre-trained language model does not conform to the conclusion, the server further uses the pre-trained language model to generate adjustment suggestions for the database operation statements corresponding to the execution results. The adjustment suggestions include but are not limited to: (1) statement correction, such as field selection in SQL queries, filter condition modification, etc.; (2) execution logic optimization, such as optimizing database operations by redefining query conditions, adjusting connection order, etc.; (3) data structure adjustment, such as adjusting fields or indexes in database tables, etc. The server then uses the pre-trained language model to refer to the adjustment suggestions to adjust the database operation statements corresponding to the execution results to obtain the adjusted database operation statements; then the server triggers the operation statement submission event again based on the adjusted database operation statement. In this embodiment, if the returned execution result does not meet the expected goal of the task, it will automatically be corrected and the operation will be re-executed instead of simply reporting an error or stopping the task. This mechanism ensures the stable execution of the data analysis task and reduces the risk of task failure.

[0056] Exemplarily, in order to avoid falling into an infinite retry state, a preset maximum number of times can be set. That is, if the received pre-trained language model output does not conform to the conclusion, the server further determines whether the number of times the pre-trained language model output does not conform to the conclusion reaches the preset maximum number of times. If the output of the pre-trained language model does not conform to the conclusion this time, and the number of times the pre-trained language model output does not conform to the conclusion does not reach the preset maximum number of times, the server further uses the pre-trained language model to generate adjustment suggestions for the database operation statement corresponding to the execution result, and uses the pre-trained language model to refer to the adjustment suggestions to adjust the database operation statement corresponding to the execution result to obtain the adjusted database operation statement. If the output of the pre-trained language model does not conform to the conclusion this time, and the number of times the pre-trained language model output does not conform to the conclusion has reached the preset maximum number of times, a prompt message of data analysis failure is output. By setting the preset maximum number of times, it is possible to enhance fault tolerance and improve the execution accuracy of the task when the execution result does not conform to the expected goal of the task, and also avoid the server from entering an infinite retry state when facing an unsolvable error; when the number of outputs that do not conform to the conclusion exceeds the preset maximum number of times, a prompt message of "data analysis failure" can be output to ensure that the user obtains clear feedback.

[0057] If the output of the pre-trained language model meets the conclusion, the server can output the data analysis results of the data analysis task based on the execution results returned by the target database. For example, the server can use the pre-trained language model to process the execution results, generate rendering code that can render the execution results into a visual chart, and then call the rendering component to execute the rendering code, generate a visual chart and output it.

[0058] This embodiment uses a pre-trained language model to judge the execution results returned by the target database, and can promptly discover situations where the execution results do not match the expected goals of the data analysis tasks, so as to make adjustments and ensure the accuracy of the analysis results. This can effectively avoid inaccurate analysis results caused by errors in the generation of database operation statements and improve the quality of data analysis.

[0059] In some embodiments, Figure 3In step S302, when generating a database operation statement, if the number of data tables in the target database is large, and the prompt word contains information about all the data tables in the target database, the amount of data in the prompt word may be large, exceeding the maximum input length specified by the pre-trained language model, and the pre-trained language model cannot process all the information at one time, resulting in failure to generate the database operation statement. Alternatively, if the prompt word contains information about all the data tables in the target database, multiple data tables in the target data table may interfere with the pre-trained language model, and the pre-trained language model may be affected by information overload, making it difficult to correctly understand the relationship between the tables, resulting in inaccurate database operation statements generated by the pre-trained language model.

[0060] Based on this, in order to improve the accuracy of database operation statements, the server can pre-build a knowledge base corresponding to any database. Taking the target database as an example, the server can query the metadata in the target database to obtain the table creation statements corresponding to all data tables from the target database. The table creation statements of each data table describe the structure of each data table, including field names, data types, constraints, etc., which helps the server to accurately understand the meaning and structure of the data table and make more accurate analysis decisions. The server stores the table creation statements corresponding to all data tables in the target database in the knowledge base corresponding to the target database to improve the intelligence of data analysis services and make the analysis process more efficient and accurate.

[0061] In some embodiments, the knowledge base corresponding to any database may store at least one of the following information in addition to storing table creation statements corresponding to all data tables in the database:

[0062] (1) The column value of a specified column in at least one data table of the target database. The column value of the specified column is the actual data of certain key columns in the database table, such as character type data. By storing these column values, the server can better understand the meaning and role of these columns in actual data analysis. When generating database operation statements, the model can use the context of these column values ​​to make more accurate judgments, such as helping to select filtering conditions. When generating database operation statements, the server can set query conditions based on the actual column values ​​(such as specific categories, names, times, etc.) to ensure that the generated query statements meet the requirements of the data analysis task.

[0063] For example, if a column value of a specified column in the table is "xxx Technology Co., Ltd." and is described as "xxx Company" in the data analysis task, if the database operation statement is generated only according to the description information of the data analysis task, an incorrect database operation statement such as "SELECT*FROM company WHERE name='xxx Company'" will be generated. This database operation statement cannot find the correct result. When the knowledge base stores "xxx Technology Co., Ltd." and is input to the pre-trained language model during the generation of the database operation statement, the pre-trained language model can refer to the column value to write the correct database operation statement "SELECT*FROM company WHERE name='xxx Technology Co., Ltd.", thereby improving the generation accuracy of the database operation statement.

[0064] (2) Multiple database operation statement examples. Database operation statement examples help the pre-trained language model understand common database operation patterns and structures. By referring to these examples, the pre-trained language model can better infer how to generate database operation statements that meet data analysis tasks, especially when faced with complex queries. Database operation statement examples can include various query types (such as aggregation queries, join queries, conditional queries, etc.), providing operation guidance for the pre-trained language model in different scenarios.

[0065] (3) Multiple data analysis experiences; multiple data analysis experiences are used to at least describe the implementation methods of different types of data analysis tasks. Data analysis experience refers to the rules, preferences and strategies accumulated in the process of historical data analysis. These experiences may include common data analysis patterns, optimization strategies and preset analysis rules. When generating database operation statements, the pre-trained language model can refer to these experiences to improve the execution efficiency and accuracy of the analysis tasks. For example, by referring to the data analysis experience of historical data analysis tasks of the same type, the pre-trained language model can select more appropriate data analysis methods or algorithms (such as statistical analysis, trend analysis, etc.) to improve the depth and quality of the analysis. For another example, data analysis experience can also include personalized preference settings to help the pre-trained language model generate customized database operation statements that meet the needs of users or organizations.

[0066] Then in a data analysis scenario, after receiving a data analysis task for a target database, the server first obtains a pre-built knowledge base corresponding to the target database, the knowledge base at least including table creation statements corresponding to all data tables in the target database, and then obtains the target table creation statement that meets the intention of the data analysis task from the knowledge base. For example, if the data analysis task is to query sales data, the server will obtain the table creation statement of the table related to the sales data from the knowledge base; then based on the description information of the data analysis task and the target table creation statement, a prompt word is generated, and the prompt word is input into the pre-trained language model so that the pre-trained language model generates a database operation statement that meets the intention of the data analysis task. In this embodiment, the generated prompt word includes the target table creation statement that meets the intention of the data analysis task, rather than all the data table information in the target database, which effectively avoids the problem of lengthy prompt words or irrelevant data table interference caused by too many data tables in traditional methods; and the target table creation statement describes the structural information of the data table related to the intention of the data analysis task, which can reduce the risk of incorrectly generating database operation statements and improve the accuracy of database operation statements.

[0067] Exemplarily, in the process of obtaining the target table creation statement, the server can recall multiple candidate table creation statements related to the intention of the data analysis task from the knowledge base. It can be understood that this embodiment does not impose any restrictions on the specific recall method, and can be specifically set according to the actual application scenario. For example, the table creation statements corresponding to all data tables in the target database are stored in the knowledge base in the form of embedded vectors. The server can convert the description information of the data analysis task into an embedded vector, and then calculate the similarity between the embedded vector of the data analysis task and the embedded vector of the table creation statement of each data table, and determine the table creation statement whose similarity exceeds a preset threshold as the table creation statement. Other recall schemes can also be used, such as recall based on graph structure. Through the recall process, relevant candidate table creation statements can be retrieved from the knowledge base to narrow the search scope, ensuring that only those statements that are most relevant to the task objectives are used.

[0068] The server then uses the pre-trained language model to refer to the intention of the data analysis task and filters out the target table creation statement from multiple candidate table creation statements. For example, the server inputs the description information of the data analysis task and multiple candidate table creation statements into the pre-trained language model, so that the pre-trained language model combines the description information of the data analysis task, analyzes and understands multiple candidate table creation statements through natural language processing technology, and identifies and screens out the target table creation statements related to the data analysis task. Through the reasoning ability of the language model, the most relevant and representative target table creation statements can be intelligently filtered out from multiple candidate statements. In this embodiment, by recalling multiple candidate table creation statements when obtaining the target table creation statement, and using the pre-trained language model for intelligent filtering, the server can effectively reduce the problems of model misunderstanding or calculation errors caused by excessive input data or too many data tables.

[0069] If the filtering result is empty, it means that there is no table creation statement related to the data analysis task in the recalled knowledge, then the server outputs a prompt message indicating that the data analysis failed, for example, it may prompt that the data analysis task is not related to the existing data or guide the user to re-enter the data analysis task.

[0070] In another data analysis scenario, in addition to obtaining the target table creation statement from the knowledge base, at least one of the target column values, target database operation statement examples, and target data analysis experience that meet the intent of the data analysis task can also be recalled from the knowledge base. The server can generate prompt words based on the target column values, target database operation statement examples, and at least one of the target data analysis experience that meet the data analysis task, the description information of the data analysis task, and the target table creation statement, and input the prompt words into the pre-trained language model so that the pre-trained language model generates a database operation statement that meets the intent of the data analysis task.

[0071] Among them, the target column value can help the server understand the meaning of the field and ensure the accuracy of the query conditions. The target database operation statement example can provide an operation statement generation template to help the pre-trained language model quickly generate database operation statements that meet the requirements. The target data analysis experience can provide rules and optimization strategies based on historical data analysis to help the pre-trained model generate accurate operation statements more efficiently. By using this information in the process of generating database operation statements, the server can significantly improve the intelligence level and accuracy of data analysis, ensure that the generated database operation statements are more in line with the requirements of data analysis tasks, and improve the efficiency and quality of queries.

[0072] In some possible implementations, if the pre-trained language model has acquired knowledge of the grammatical rules of the target database through learning, there is no need to explicitly add grammatical rules to the input prompt words. The ability to generate grammatical rules may be internalized in the pre-trained language model itself and is an implicit capability of the model.

[0073] Alternatively, if the pre-trained language model has not fully learned the grammatical rules of the target database, or the grammar of the target database is relatively special, the above-mentioned prompt words for generating database operation statements can also include the grammatical rules of the target database, so that the database operation statements generated by the pre-trained language model conform to the grammatical rules. In this case, the prompt words will contain grammatical rule information related to the target database (such as table name, field type, query syntax, etc.) to help the model understand how to construct operation statements that conform to the target database. Grammatical rules can include how to correctly use specific keywords, functions, operators, etc., thereby reducing errors and improving efficiency.

[0074] In some possible implementations, after determining the target table creation statement, the server may also generate a query statement for the target data table indicated by the target table creation statement based on a preset query statement template and the data table information described by the target table creation statement; send the query statement to the target database for execution; and receive the random content in the target data table returned by the target database executing the query statement. Among them, the above-mentioned prompt words for generating database operation statements may also include random content in the target data table, so that the random content in the target data table is input into the pre-trained language model, which helps the pre-trained language model better understand the specific form and structure of the data, which can help the pre-trained language model identify the rules, features and potential relationships in the data, so as to better generate database operation statements that conform to the target data table.

[0075] It is understandable that the above-mentioned prompt words for generating database operation statements at least include description information of the data analysis task and the target table creation statement, and on this basis, at least one other information mentioned above is added according to the actual application scenario. This implementation does not impose any restrictions on this.

[0076] In an exemplary embodiment, see Figure 4 , provides a data analysis process, including a knowledge recall sub-process, a table creation statement filtering sub-process, a database operation statement generation sub-process, a database operation statement execution sub-process, an execution result evaluation sub-process and a result output sub-process.

[0077] (1) In the knowledge recall sub-process, the server recalls multiple candidate table creation statements, target column values, target database operation statement examples, and target data analysis experience related to the data analysis task from the knowledge base corresponding to the target database based on the received data analysis task for the target database. Through the knowledge recall process, the table creation statements and related knowledge related to the data analysis task can be located to avoid interference from other irrelevant information. Then the table creation statement filtering sub-process is carried out.

[0078] (2) In the table creation statement filtering subprocess, the description information of the data analysis task and multiple candidate table creation statements are used as the input of the pre-trained language model, so that the pre-trained language model refers to the intention of the data analysis task and filters out the target table creation statement from multiple candidate table creation statements. If the filtering result is empty, it means that there are no table creation statements related to the data analysis task in the recalled knowledge, and a prompt message indicating that the data analysis failed is output. Through the table creation statement filtering subprocess, it is ensured that the target table creation statements related to the data analysis task are accurately located and screened out, the accuracy and effectiveness of subsequent processes are optimized and improved, and the system response efficiency is improved. If the filtering result is not empty, the database operation statement generation subprocess is then carried out.

[0079] (3) In the database operation statement generation sub-process, the server can use the grammatical rules of the target database, the target table creation statement, the target column value, the target database operation statement example, the target data analysis experience, the description information of the data analysis task, and the random content in the target data table indicated by the target table creation statement as the input of the pre-trained language model, so that the pre-trained language model generates the database operation statement according to the input information. The model refers to multiple information, which helps to improve the generation accuracy of the database operation statement and the subsequent execution success rate. Then proceed to the database operation statement execution sub-process.

[0080] (4) In the database operation statement execution sub-process, the server submits the database operation statement to the target database for execution.

[0081] If an execution error message returned by the target database is received and the number of executions of the target database does not reach the preset number of retries, the database operation statement generation sub-process is returned, so that in the database operation statement generation sub-process, at least the execution error message is used as the input of the pre-trained language model (in addition to the execution error message, at least one of the above-mentioned grammatical rules of the target database, target table creation statements, target column values, target database operation statement examples, target data analysis experience, description information of the data analysis task, and random content in the target data table indicated by the target table creation statement can be used as the input of the pre-trained language model again), so that the pre-trained language model modifies the last generated database operation statement with reference to the execution error message, obtains the modified database operation statement, and performs the database operation statement execution sub-process again based on the modified database operation statement.

[0082] If an execution error message is received from the target database and the number of executions of the target database reaches the preset number of retries, a prompt message indicating that the data analysis failed is output.

[0083] If the target database is executed successfully, the execution result evaluation sub-process will be carried out.

[0084] (5) In the execution result evaluation process, the description information of the data analysis task, the execution result returned by the target database and the corresponding database operation statement are used as the input of the pre-trained language model, so that the pre-trained language model determines whether the execution result meets the expected goal of the data analysis task. If the output of the pre-trained language model does not meet the conclusion this time, and the number of times the pre-trained language model outputs do not meet the conclusion does not reach the preset maximum number, the pre-trained language model is further used to generate adjustment suggestions for the database operation statement corresponding to the execution result, and then the database operation statement generation sub-process is returned, so that in the database operation statement generation sub-process, at least the database operation statement corresponding to the execution result and the adjustment suggestions for the database operation statement corresponding to the execution result are used as the input of the pre-trained language model (for example, the grammatical rules of the target database, the target table creation statement, the target column value, the target database operation statement example, the target data analysis experience, the description information of the data analysis task, and the random content in the target data table indicated by the target table creation statement can also be used as the input of the pre-trained language model again), so that the pre-trained language model adjusts the database operation statement corresponding to the execution result with reference to the adjustment suggestions to obtain the adjusted database operation statement. Then the database operation statement execution sub-process is performed again.

[0085] If the output of the pre-trained language model does not conform to the conclusion this time, and the number of times the pre-trained language model output does not conform to the conclusion reaches the preset maximum number, a prompt message indicating that the data analysis has failed is output.

[0086] If the output of the pre-trained language model this time meets the conclusion, the result output sub-process will be carried out.

[0087] (6) In the result output subprocess, the server uses the description information of the data analysis task, the execution results returned by the target database, and the corresponding database operation statements as inputs to the pre-trained language model, so that the pre-trained language model generates rendering code that can render the execution results into visual charts. Next, the server calls the rendering component to execute the rendering code, generates visual charts, and outputs them, providing users with intuitive and easy-to-understand graphical data representations, improving data insight and analysis efficiency.

[0088] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope of this specification.

[0089] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements any of the above methods by running the executable instructions.

[0090] Figure 5 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 5 At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and may also include hardware required for other functions. One or more embodiments of this specification may be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0091] The data analysis device can be used for Figure 5 The device shown in the figure is used to implement the technical solution of this specification. The data analysis device may include:

[0092] A data analysis task receiving module, used for receiving data analysis tasks for a target database;

[0093] A database operation statement generation module, used to generate prompt words according to the data analysis task and the data table information in the target database, and input the prompt words into a pre-trained language model so that the pre-trained language model generates a database operation statement that meets the data analysis task;

[0094] A database operation statement submission module, used to trigger an operation statement submission event based on the database operation statement, so that the target database executes the received database operation statement;

[0095] The database operation statement generating module is further used to modify the database operation statement by using the pre-trained language model and referring to the execution error information to obtain a modified database operation statement if execution error information returned by the target database is received;

[0096] The database operation statement submission module is further used to trigger the operation statement submission event again based on the modified database operation statement;

[0097] The data analysis result output module is used to output the data analysis result of the data analysis task if the target database is executed successfully.

[0098] In some embodiments, the database operation statement generation module is specifically used to modify the last generated database operation statement by using the pre-trained language model with reference to the execution error information if an execution error message returned by the target database is received and the number of executions of the target database has not reached a preset number of retries.

[0099] The device also includes a prompt information output module, which is used to output a prompt information of data analysis failure if the execution error information returned by the target database is received and the execution times of the target database reaches the preset retry times.

[0100] In some embodiments, the device also includes an execution result evaluation module, which is used to receive the execution result returned by the target database when the execution is successful; use the pre-trained language model to determine whether the execution result meets the expected goal of the data analysis task; if the output of the pre-trained language model does not meet the conclusion, further use the pre-trained language model to generate adjustment suggestions for the database operation statement corresponding to the execution result.

[0101] The database operation statement generation module is also used to modify the database operation statement corresponding to the execution result by using the pre-trained language model and referring to the adjustment suggestion to obtain a modified database operation statement.

[0102] The database operation statement submission module is further used to trigger the operation statement submission event again based on the modified database operation statement.

[0103] In some embodiments, the execution result evaluation module is specifically used to further use the pre-trained language model to generate adjustment suggestions for the database operation statement corresponding to the execution result if the output of the pre-trained language model this time does not conform to the conclusion and the number of times the pre-trained language model outputs the conclusion that does not conform to the conclusion does not reach a preset maximum number.

[0104] The device also includes a prompt information output module, which is used to output a prompt information of data analysis failure if the output of the pre-trained language model does not conform to the conclusion this time, and the number of times that the pre-trained language model outputs the conclusion that does not conform to the conclusion reaches a preset maximum number.

[0105] In some embodiments, the data analysis result output module is specifically used to receive the execution result returned by the target database when the execution is successful; use the pre-trained language model to process the execution result to generate a rendering code that can render the execution result into a visual chart; call the rendering component to execute the rendering code, generate a visual chart and output it.

[0106] In some embodiments, the device further includes a knowledge acquisition module, which is used to acquire a pre-built knowledge base corresponding to the target database, wherein the knowledge base at least includes table creation statements corresponding to all data tables in the target database; and to acquire a target table creation statement that meets the data analysis task from the knowledge base. The data table information in the target database includes the target table creation statement.

[0107] Exemplarily, the knowledge acquisition module is specifically used to recall multiple candidate table creation statements related to the data analysis task from the knowledge base; use the pre-trained language model to refer to the data analysis task, and filter out the target table creation statement from the multiple candidate table creation statements;

[0108] The device also includes a prompt information output module, which is used to output a prompt information of data analysis failure if the filtering result is empty.

[0109] In some embodiments, the knowledge base also includes at least one of the following: column values ​​of specified columns in at least one data table of the target database, multiple database operation statement examples, and multiple data analysis experiences; the multiple data analysis experiences are at least used to describe the implementation methods of different types of data analysis tasks; the prompt words also include at least one of the following: target column values, target database operation statement examples, and target data analysis experiences that are recalled from the knowledge base and meet the data analysis tasks.

[0110] In some embodiments, it also includes a table content acquisition module, which is used to generate a query statement for the target data table indicated by the target table creation statement based on a preset query statement template and the data table information described by the target table creation statement; send the query statement to the target database for execution; and receive the random content in the target data table returned by the target database executing the query statement. Wherein, the prompt word also includes the random content in the target data table.

[0111] In some embodiments, the prompt word also includes the grammatical rules of the target database, so that the database operation statement generated by the pre-trained language model conforms to the grammatical rules.

[0112] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0113] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.

[0114] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0115] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0116] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.

[0117] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.

Claims

1. A data analysis method, comprising: Receive data analysis tasks for the target database; Generate prompt words according to the data analysis task and the data table information in the target database, and input the prompt words into a pre-trained language model so that the pre-trained language model generates a database operation statement that meets the data analysis task; triggering an operation statement submission event based on the database operation statement, so that the target database executes the received database operation statement; If execution error information returned by the target database is received, modify the database operation statement by using the pre-trained language model with reference to the execution error information to obtain a modified database operation statement, and trigger the operation statement submission event again based on the modified database operation statement; If the target database is executed successfully, the data analysis result of the data analysis task is output.

2. The method according to claim 1, wherein if execution error information returned by the target database is received, modifying the database operation statement by using the pre-trained language model with reference to the execution error information comprises: If an execution error message returned by the target database is received and the number of executions of the target database has not reached a preset number of retries, modify the last generated database operation statement by using the pre-trained language model with reference to the execution error message; The method further comprises: If the execution error information returned by the target database is received and the execution times of the target database reaches the preset retry times, a prompt message indicating that the data analysis has failed is output.

3. The method according to claim 1, before outputting the data analysis result of the data analysis task, further comprising: Receiving the execution result returned by the target database when the execution is successful; Using the pre-trained language model to determine whether the execution result meets the expected goal of the data analysis task; If the output of the pre-trained language model does not conform to the conclusion, further using the pre-trained language model to generate an adjustment suggestion for the database operation statement corresponding to the execution result, and using the pre-trained language model to adjust the database operation statement corresponding to the execution result with reference to the adjustment suggestion to obtain an adjusted database operation statement; The operation statement submission event is triggered again based on the adjusted database operation statement.

4. The method according to claim 3, if the output of the pre-trained language model does not conform to the conclusion, further using the pre-trained language model to generate an adjustment suggestion for the database operation statement corresponding to the execution result, and using the pre-trained language model to refer to the adjustment suggestion to adjust the database operation statement corresponding to the execution result to obtain the adjusted database operation statement, comprising: If the pre-trained language model outputs a non-conforming conclusion this time, and the number of times the pre-trained language model outputs the non-conforming conclusion does not reach a preset maximum number, further using the pre-trained language model to generate an adjustment suggestion for the database operation statement corresponding to the execution result, and using the pre-trained language model to adjust the database operation statement corresponding to the execution result with reference to the adjustment suggestion, to obtain an adjusted database operation statement; The method further comprises: If the output of the pre-trained language model this time does not conform to the conclusion, and the number of times the pre-trained language model outputs the conclusion that does not conform to the conclusion reaches a preset maximum number, a prompt message indicating that the data analysis has failed is output.

5. The method according to claim 1, wherein outputting the data analysis result of the data analysis task comprises: Receiving the execution result returned by the target database when the execution is successful; Processing the execution result using the pre-trained language model to generate rendering code capable of rendering the execution result into a visual chart; The rendering component is called to execute the rendering code, generate a visual chart and output it.

6. The method according to claim 1, before generating prompt words according to the data analysis task and the data table information in the target database, and inputting the prompt words into a pre-trained language model so that the pre-trained language model generates a database operation statement that meets the data analysis task, further comprises: Acquire a pre-built knowledge base corresponding to the target database, wherein the knowledge base at least includes table creation statements corresponding to all data tables in the target database; Acquire a target table creation statement that meets the data analysis task from the knowledge base; The data table information in the target database includes the target table creation statement.

7. According to the method of claim 6, obtaining a target table creation statement that meets the data analysis task from the knowledge base comprises: Recalling a plurality of candidate table creation statements related to the data analysis task from the knowledge base; Using the pre-trained language model and referring to the data analysis task, the target table creation statement is filtered out from the plurality of candidate table creation statements; The method further comprises: If the filtered result is empty, a prompt message indicating that the data analysis failed is output.

8. According to the method of claim 6, the knowledge base further comprises at least one of the following: column values ​​of a specified column in at least one data table of the target database, a plurality of database operation statement examples, and a plurality of data analysis experiences; the plurality of data analysis experiences are at least used to describe data analysis rules for different types of data analysis tasks; The prompt words also include at least one of the following: target column values, target database operation statement examples, and target data analysis experience that are recalled from the knowledge base and meet the data analysis task.

9. The method according to claim 6, further comprising: Based on a preset query statement template and the data table information described by the target table creation statement, a query statement for the target data table indicated by the target table creation statement is generated; Sending the query statement to the target database for execution; Receiving random content in the target data table returned by the target database executing the query statement; Wherein, the prompt word also includes random content in the target data table.

10. According to the method of claim 6, the prompt word also includes the grammatical rules of the target database, so that the database operation statement generated by the pre-trained language model conforms to the grammatical rules.

11. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.

12. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

13. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Information processing method, information processing device, computer equipment and storage medium

    CN117273017A

  • Database text query method, device and equipment based on data examples

    CN117743374A

  • Natural language intelligent query method and device based on multi-agent interaction

    CN118012900A

  • Method and system for automatically generating sql and replying questions

    CN118093634A

  • Database automatic man-machine interaction method based on large language model

    CN118363984A