Index processing method and device, storage medium and program product

Through a two-stage processing strategy, the pre-trained language model is used to generate non-relational database query statements and SQL statements, which solves the problem of flexible query and personalized analysis of indicator data in non-relational databases, and realizes efficient and accurate indicator processing.

CN120429401APending Publication Date: 2025-08-05ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510515264.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing technology is difficult to meet the indicator data needs of flexible query and personalized analysis, especially in non-relational databases. Users need to master complex query statement writing or rely on complex pre-trained language model conversion, resulting in high operational thresholds and low efficiency.

Method used

A two-stage processing strategy is adopted: in the first stage, a simple query statement that conforms to a non-relational database is generated through a pre-trained language model to ensure the accuracy of the basic query function; in the second stage, the query results are stored in a database that supports structured queries, and SQL statements are generated using the pre-trained language model for further analysis and processing.

Benefits of technology

It simplifies user interaction, reduces operational complexity, improves query accuracy and analysis accuracy, reduces syntax errors and resource consumption, and improves the ease of use and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429401A_ABST
    Figure CN120429401A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide an index processing method and device, a storage medium and a program product. The method comprises the following steps: receiving an index processing demand which is input by a user and is described in a natural language; based on the index processing demand and a non-relational database used for storing index data, determining a to-be-queried index indicated by the index processing demand and a query condition; utilizing the pre-training language model to generate a query statement which conforms to the grammar specification of the non-relational database and queries the data of the to-be-queried index according to the query condition, and submitting the query statement to the non-relational database for execution to obtain a query result; if the index processing requirement further comprises an analysis requirement for the query result, the query result is stored in a specified database supporting structured query, a pre-training language model is used for generating an SQL statement used for analyzing and processing the query result according to the analysis requirement, the SQL statement is submitted to the specified database to be executed, and an index processing result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of data processing technology, and in particular, to an indicator processing method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In various information systems, data management and indicator monitoring are crucial components supporting the entire lifecycle of these systems, encompassing core functions such as resource management, anomaly alerts, performance optimization, and security assurance. Systems typically collect and store various indicator data, which can be used to measure system health, performance, and potential areas for improvement. Because many indicator data exhibit time series characteristics, non-relational databases are often used to store this data. Each record associated with each indicator is accompanied by a precise timestamp, facilitating historical trend tracking and predictive analysis.

[0003] To meet user indicator query needs, the traditional approach is to visualize data based on predefined indicator dashboards. However, in the face of complex and changing usage scenarios, such static solutions are difficult to meet the needs of flexible query and personalized analysis. There are two solutions in related technologies:

[0004] One is to adopt a self-service query mode, where users manually search the indicator catalog or manually write query statements adapted to non-relational databases to obtain the indicator data of interest. This method has the limitations of high operation threshold and low response efficiency.

[0005] The other approach is to rely on pre-trained language models to achieve intelligent conversion (Text-to-SQL) of natural language into SQL (Structured Query Language) statements. However, since non-relational databases use a time-series data storage structure rather than a traditional relational table structure, a dedicated SQL execution engine must be developed. This engine converts SQL statements into query instructions suitable for non-relational databases through syntax parsing optimization and query plan rewriting, making the development process relatively complex. Summary of the Invention

[0006] In view of this, one or more embodiments of this specification provide an indicator processing method, an electronic device, a computer-readable storage medium, and a computer program product.

[0007] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions:

[0008] According to a first aspect of one or more embodiments of this specification, a method for processing an indicator is proposed, including:

[0009] Receive indicator processing requirements described in natural language from users;

[0010] Based on the indicator processing requirement and a non-relational database for storing indicator data, determining the indicator to be queried and the query condition indicated by the indicator processing requirement;

[0011] Generate a query statement that conforms to the syntax specification of the non-relational database and queries the data of the indicator to be queried according to the query condition by using the pre-trained language model, and submit the query statement to the non-relational database for execution to obtain a query result;

[0012] If the indicator processing requirements also include analysis requirements for the query results, the query results are stored in a designated database that supports structured queries, and the pre-trained language model is used to generate an SQL statement for analyzing and processing the query results according to the analysis requirements, and the SQL statement is submitted to the designated database for execution to obtain the indicator processing results.

[0013] According to a second aspect of the embodiments of this specification, an electronic device is provided, including:

[0014] processor;

[0015] a memory for storing processor-executable instructions;

[0016] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.

[0017] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0018] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0019] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:

[0020] In the embodiments of this specification, the user only needs to input the indicator processing requirements described in natural language, which simplifies the complexity of the user's interaction with the electronic device. The user does not need to master complex database knowledge. The electronic device determines the indicator to be queried and the query conditions indicated by the indicator processing requirements. The indicator processing process is divided into two stages. The first stage is the query stage, which is to allow the pre-trained language model to generate a query statement that conforms to the grammatical specifications of the non-relational database and queries the data of the indicator to be queried according to the query conditions, and submits it to the non-relational database for execution to ensure the accuracy of the basic query function. First, allowing the pre-trained language model to generate a relatively simple statement with only basic query functions can effectively improve the query accuracy and avoid grammatical specification errors or query failures caused by directly processing complex analysis. The second stage is the analysis stage. If the user's indicator processing requirements include further analysis requirements, the query results obtained in the query stage are first stored in a database that supports structured queries. Then, the pre-trained language model is used to generate an SQL statement to analyze and process the query results according to the analysis requirements for data analysis and processing, making full use of the mature ability of the pre-trained language model in SQL statement generation, thereby achieving more accurate indicator analysis.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a schematic diagram of the architecture of an indicator processing service system provided by an exemplary embodiment.

[0023] Figure 2A It is a flowchart of an indicator processing method provided by an exemplary embodiment.

[0024] Figure 2B It is a flowchart of another indicator processing method provided by an exemplary embodiment.

[0025] Figure 3 It is a schematic diagram of an exemplary embodiment providing a method of using a pre-trained language model to determine the indicators to be queried and the query conditions indicated by the indicator processing requirements.

[0026] Figure 4 This is a schematic diagram of generating a query statement using a pre-trained language model provided by an exemplary embodiment.

[0027] Figure 5 This is a schematic diagram of generating SQL statements using a pre-trained language model provided by an exemplary embodiment.

[0028] Figure 6 It is a structural diagram of an electronic device provided by an exemplary embodiment. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.

[0030] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.

[0031] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0032] Here are some explanations of the terms mentioned in this manual:

[0033] 1. Pretrained language models, such as the Large Language Model (LLM) and the BERT model, are machine learning models pretrained using large amounts of data and complex algorithms. They can understand and generate data in various forms, such as human language and images, and are used in a wide range of applications, such as text generation, translation, and sentiment analysis. These models absorb vast amounts of information to improve the accuracy of their predictions and decisions, enabling them to demonstrate human-level or higher cognitive capabilities in specific tasks.

[0034] 2. Prompts are tips, instructions, or directions. In the field of artificial intelligence (AI), prompts are used to guide and direct pre-trained language models to produce specific outputs. Prompts can be text snippets or questions that inspire the pre-trained language model to think and generate relevant content.

[0035] 3. Non-relational databases are databases that don't use the traditional relational model (such as tables and rows and columns) to store data. They are more flexible in handling large amounts of unstructured or semi-structured data. Unlike relational databases, non-relational databases typically don't use SQL as a query language. Instead, they offer a variety of data storage models, such as key-value pairs, documents, column families, or graphs, tailored to different application scenarios. They offer advantages in scalability, performance, and high-concurrency processing, making them particularly well-suited for big data, real-time analytics, and distributed systems.

[0036] Data management and metric monitoring are crucial components of any information system's lifecycle, encompassing core functions such as resource management, anomaly alerts, performance optimization, and security assurance. Systems typically collect and store various metrics, which can be used to measure system health, performance, and potential areas for improvement.

[0037] Taking the operation and maintenance system as an example, the operation and maintenance system is a technical system that supports the full life cycle management of IT (Information Technology) infrastructure, applications and services. Its core functions cover resource monitoring, abnormal alarms, performance optimization and security assurance, aiming to ensure the high availability, stability and security of the IT environment. As a core component of the operation and maintenance system, the operation and maintenance database specializes in storing operation and maintenance indicator data with time series characteristics. Operation and maintenance indicators are a set of metrics used to measure and evaluate the operation and maintenance effects of IT systems. Through operation and maintenance indicators, operation and maintenance personnel can understand the health status and performance of IT systems in real time, as well as whether optimization or troubleshooting is needed. These operation and maintenance indicators are stored in the operation and maintenance database in the form of time series data. Each piece of data has a precise timestamp, which supports historical trend backtracking and predictive analysis.

[0038] Exemplary operation and maintenance indicators include, but are not limited to: ① System performance indicators, such as CPU utilization, memory utilization, disk usage, network bandwidth and traffic, etc. ② Service availability indicators, such as system uptime, fault recovery time, and fault interval time, etc. ③ Performance and response time indicators, such as request response time, throughput, and latency, etc. ④ Security indicators, such as intrusion detection rate, vulnerability repair time, and number of security incidents, etc. ⑤ Operation and maintenance efficiency indicators, such as incident response time (the time from incident occurrence to operation and maintenance personnel response), change management efficiency, and automated task ratio (the proportion of automated operation and maintenance tasks to total operation and maintenance tasks).

[0039] For example, if the operation and maintenance system is a Prometheus system, the operation and maintenance database supports querying and aggregating indicator data through the PromQL (Prometheus Query Language) query language. The basic unit of PromQL query is "indicator". Each indicator is usually a time series, representing a specific monitoring data, such as CPU usage, memory usage, number of requests, etc. In other words, the data stored in the operation and maintenance database is time series data, and each time series data is usually composed of the following parts: ① Indicator name, which describes the core meaning of the data, similar to the column name in the SQL database. ② Label, the label is a key-value pair used to further distinguish different time series data, and each time series data can have multiple labels. ③ Timestamp, which records the timestamp of each data point, indicating the recording time of the time series data. ④ Value, the actual value of the data point, is usually a numeric type. For counter-type indicators, it represents the current count value; for instrument-type indicators, it represents the measurement value.

[0040] Of course, in addition to operation and maintenance scenarios, there is also a need to manage and analyze indicator data in the form of time series data in areas such as financial monitoring, industrial equipment status assessment, and environmental monitoring.

[0041] Based on the problems mentioned in the above background technology, the embodiments of this specification provide an indicator processing method, an electronic device, a computer-readable storage medium, and a computer program product.

[0042] In one possible application scenario, based on the user's privacy protection requirements, a non-relational database for storing indicator data is deployed in the user's electronic device, which means that all data storage and query operations will be performed on the user's electronic device. The core purpose of this approach is to ensure that the user's data does not leave their device, thereby avoiding the risk of data leakage. Under this architecture, the indicator processing method provided in the embodiments of this specification is designed to be executed directly on the user's local electronic device, which includes but is not limited to physical servers, server clusters, cloud servers, smart phones / mobile phones, tablet computers, personal digital assistants (PDAs), laptop computers, and desktop computers.

[0043] In another possible application scenario, such as Figure 1 As shown, Figure 1 1 is a schematic diagram of an architecture of an indicator processing service system provided by an exemplary embodiment. The system may include a server 11, a network 12, and several user terminals, such as a PC (Personal Computer) 13, a mobile phone 14, and the like.

[0044] The server 11 can be a physical server containing an independent host, or the server 11 can be a virtual server carried by a host cluster. During operation, the server 11 can run the server-side program of the indicator processing application to realize the corresponding indicator processing service platform. Exemplarily, the server 11 is a database server, such as maintaining one or more non-relational databases for storing indicator data of each client to ensure efficient data access and query optimization. The functions provided by the server 11 include: ① Data storage, such as using a non-relational database to store indicator data, supporting the writing, updating and long-term storage of indicator data. ② Data query, returning relevant indicator data based on the query requirements input by the user.

[0045] PC 13 and mobile phone 14 are only some types of user terminals that can be used by users. In fact, users can obviously also use user terminals such as the following types: tablet devices, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), etc., and one or more embodiments of this specification do not limit this. During operation, the user terminal can run the client-side program of the indicator processing application, which can be implemented as a client of the indicator processing service. Among them, the client application of the above-mentioned indicator processing service can be started and run on the user terminal. The client-side program can be a native application installed on the user terminal, or the client-side program can be a small program, a quick application or other similar forms. Of course, when using web page technologies such as HTML5 or similar, the relevant functions can be implemented through the page displayed by the browser. The browser here can be an independent browser application or a browser module embedded in certain applications.

[0046] Regarding the network 12 for interaction between user terminals such as PC 13 and mobile phone 14 and server 11, communication can be achieved using a wired or wireless network based on the communication methods supported by the corresponding user terminals, and this specification does not limit this. For example, if PC 13 supports both wired and wireless communication, then communication can be achieved using either a wired or wireless network as needed, while mobile phone 14 generally only supports wireless communication and thus can achieve communication using a wireless network.

[0047] The indicator processing method provided in the embodiment of this specification can be executed by the server 11, or by the user terminal, or part of the indicator processing method can be executed by the user terminal and the other part can be executed by the server 11. This embodiment does not impose any restrictions on this.

[0048] In the process of implementing the embodiments of this specification, the inventors further considered that the pre-trained language models of related technologies are already relatively mature in Text-to-SQL tasks and can convert natural language into SQL statements with high accuracy; however, for non-relational databases, due to the diversity and complexity of their query syntax specifications, the conversion technology from natural language to non-relational database query statements is not yet perfect. In particular, when complex analysis requirements are involved, the accuracy of directly generating non-relational database query statements that have both query and analysis functions is low. Based on this, the data query method provided in the embodiments of this specification adopts a two-stage processing strategy to improve the accuracy and adaptability of the overall query:

[0049] Phase 1: Query Phase—The pre-trained language model generates query statements that only have query functionality and conform to the grammatical specifications of the non-relational database. These statements are then submitted to the non-relational database for execution to ensure the accuracy of basic query functionality. Initially, having the pre-trained language model generate relatively simple statements that only have basic query functionality can effectively improve query accuracy and avoid grammatical errors or query failures that may result from directly processing complex analysis.

[0050] Phase 2: Analysis Phase - If the user's indicator processing requirements include further analysis requirements, the query results are first stored in a database that supports structured queries (such as a relational database or an in-memory database). Then, the pre-trained language model is used to generate SQL statements for data analysis and processing, fully utilizing the pre-trained language model's mature capabilities in SQL statement generation to achieve more accurate indicator analysis.

[0051] See also Figure 2A Next, a data query method provided in an embodiment of this specification is exemplarily described. The method can be applied to an electronic device (such as the above-mentioned server or other electronic device connected to a database). The method includes:

[0052] In S200 , an indicator processing requirement described in natural language and input by a user is received.

[0053] This step simplifies the complexity of user interaction with the electronic device by accepting the user's indicator processing requirements expressed in natural language. Users do not need to master complex database knowledge; they simply express their requirements in natural language, and the electronic device intelligently interprets and translates them into specific indicator processing tasks. This approach improves the usability of electronic devices, lowers the user barrier to entry, and is particularly friendly to non-technical users, while also enhancing the user experience.

[0054] In S202 , based on the indicator processing requirement and the non-relational database for storing indicator data, the indicator to be queried and the query condition indicated by the indicator processing requirement are determined.

[0055] This step mainly involves parsing the indicator processing requirements input by the user and matching the indicator catalog. By determining the indicators to be queried and their query conditions in advance, it ensures that subsequent query operations only target the data of the indicators to be queried, thereby avoiding redundant data processing.

[0056] Among them, the query conditions can be characterized according to different user needs and data characteristics. For example, the query conditions include but are not limited to: ① Time range conditions, which limit the time window for data query to ensure that only data within the time period of interest to the user is extracted to avoid redundant information, such as querying data for the past 24 hours, the last 7 days, or a specific start and end time period. ② Numerical filtering conditions, which enable fine-grained screening of data indicators, focusing on abnormal or specific data, such as querying data with a CPU utilization rate exceeding 80%, or records with a memory usage rate below 30%. ③ Status conditions, based on the system operation status, filter out data in a specific state to help users quickly locate problems, such as querying indicator data in an "abnormal" or "warning" state.

[0057] In one possible implementation, see Figure 3 The electronic device can generate a first prompt word based on the indicator processing requirement and all indicators included in the non-relational database. The first prompt word is then input into a pre-trained language model. The pre-trained language model then determines the indicator to be queried indicated by the indicator processing requirement from all indicators included in the non-relational database, and extracts the query conditions for the indicator to be queried from the indicator processing requirement. This embodiment utilizes a pre-trained language model to automatically identify the indicator to be queried and its query conditions that match the indicator processing requirement, reducing the risk of manual intervention and misjudgment.

[0058] For example, suppose the metric processing requirement is "Query data on server CPU utilization exceeding 80% over the past seven days." A non-relational database contains all the metrics in the following list: [CPU utilization, memory usage, disk read / write, network traffic, etc.]. The user requirement and the metric list are combined into a first prompt, for example: "From the following list of metrics, identify the metrics relevant to the requirement 'Query data on server CPU utilization exceeding 80% over the past seven days': metric list: [CPU utilization, memory usage, disk read / write, network traffic, etc.]. Please enter the name of the metric to be queried and the corresponding query conditions." This first prompt guides the pre-trained language model to filter out the truly relevant metrics from all metrics and parse the corresponding query conditions based on the user's requirements.

[0059] Exemplarily, the first prompt word may also include at least one of the following prompt information: (1) The role played by the pre-trained language model: setting a specific role so that the model can generate a response based on the role. (2) Background information: providing sufficient background information so that the model can understand the context of the question. (3) The task to be performed by the pre-trained language model: letting the model know clearly what needs to be achieved, such as indicator matching. (4) At least one reference example; representative and differentiated examples can be added. Providing diverse examples helps the model better understand the task requirements and imitate the example style to complete the task.

[0060] In another possible implementation, the electronic device can convert the indicator processing requirement into a first embedding vector that can capture the key information and semantic features of the requirement; and the electronic device converts each indicator contained in the non-relational database into a second embedding vector that reflects the potential semantic content of each indicator. It is understandable that to improve processing efficiency, each indicator contained in the non-relational database can be converted into a second embedding vector in advance and stored for use in each indicator processing task. Then, based on the similarity between the first embedding vector and each second embedding vector (for example, using a metric such as cosine similarity or Euclidean distance), the query indicator indicated by the indicator processing requirement is determined. In this process, indicators with high similarity (for example, similarity greater than a preset threshold) are considered to be query indicators, while indicators with low similarity are excluded. In this embodiment, semantic matching of requirements and indicators through embedding vectors can automatically identify the most relevant query indicators, reduce the need for manual intervention, and improve screening efficiency. Moreover, for complex or ambiguous indicator processing requirements, vector similarity calculation can better understand the semantics behind the requirements, thereby screening out query indicators that better meet the requirements.

[0061] For query conditions, electronic devices can use natural language processing technology to extract query conditions from the indicator processing requirements. Alternatively, predefined condition templates can be introduced to match common query conditions. For example, for time conditions, templates such as "Last 7 Days" and "Last Month" can be predefined and the description in the indicator processing requirement text can be matched with the condition template. Alternatively, the indicator processing requirements can be input into a pre-trained language model, which can then extract the query conditions for the indicator to be queried from the indicator processing requirements.

[0062] In S208, a pre-trained language model is used to generate a query statement that complies with the grammatical specifications of the non-relational database and queries the data of the to-be-queried indicator according to the query conditions, and the query statement is submitted to the non-relational database for execution to obtain a query result.

[0063] For example, see Figure 4, the electronic device can generate a third prompt word based on the indicator to be queried, the query condition of the indicator to be queried, and the grammatical specification of the non-relational database, and input the third prompt word into the pre-trained language model, so that the pre-trained language model generates a corresponding query statement according to the instruction of the third prompt word. The generated query statement conforms to the grammatical specification of the non-relational database and only contains basic query functions. This embodiment utilizes the excellent performance of the pre-trained language model in natural language processing, and can convert the user's query intention into a database query statement, reducing the complexity of manually writing query statements, and by generating simple statements that focus on data queries, it can avoid errors caused by directly generating complex query statements, thereby ensuring the accuracy of data extraction.

[0064] For example, suppose the metric to be queried is CPU usage, and the query condition is: Query CPU usage records for the past 7 days. The third prompt word could be "Please generate a query statement to query the metric 'CPU usage' from a non-relational database, with a time range of the past 7 days. The syntax must conform to the query requirements of the non-relational database."

[0065] It will be understood by those skilled in the art that if the pre-trained language model has already learned the relevant knowledge of the grammatical specifications of the non-relational database, there is no need to input the grammatical specifications of the non-relational database into the pre-trained language model.

[0066] Exemplarily, the third prompt word may also include at least one of the following prompt information: (1) The role played by the pre-trained language model: setting a specific role so that the model can generate a response based on the role. (2) Background information: providing sufficient background information so that the model can understand the context of the question. (3) The task to be performed by the pre-trained language model: letting the model know clearly what needs to be achieved, such as generating a query statement. (4) At least one reference example; representative and differentiated examples can be added. Providing diverse examples helps the model better understand the task requirements and imitate the example style to complete the task.

[0067] Next, the electronic device may pass the generated query statement to the database for execution. After the query statement is submitted, the non-relational database parses and executes the statement and returns the query result.

[0068] In one possible scenario, after submitting a query statement to a non-relational database for execution, if a first execution error message (such as a syntax error, data mismatch, etc.) is received from the non-relational database, and the number of executions of the non-relational database does not reach the first retry count, the electronic device inputs the received first execution error message and the query statement to be modified into the pre-trained language model, so that the pre-trained language model can be used to modify the query statement to be modified with reference to the first execution error message, to obtain a modified query statement, and to resubmit the modified query statement to the non-relational database for execution. This embodiment effectively solves the problem of the lack of a self-repair mechanism in traditional data analysis, automatically repairs errors through a pre-trained language model, reduces the need for manual intervention, improves adaptive capabilities, ensures the smooth progress of query tasks, and reduces the time and cost of troubleshooting; and by setting the first retry count, it can avoid endless retries that consume computing and storage resources of the electronic device. The specific value of the first retry count can be set according to the actual application scenario.

[0069] If an execution error message is received from the non-relational database and the number of executions of the non-relational database reaches the first retry count, a prompt message indicating that the indicator processing failed is output to ensure that the user receives clear feedback.

[0070] In S206, if the indicator processing requirements also include analysis requirements for the query results, the query results are stored in a designated database that supports structured queries, and a pre-trained language model is used to generate an SQL statement for analyzing and processing the query results according to the analysis requirements. The SQL statement is submitted to the designated database for execution to obtain the indicator processing results.

[0071] For example, see Figure 2B The electronic device can input the indicator processing requirements and the query results into the pre-trained language model to use the pre-trained language model to determine whether the indicator processing requirements also include analysis requirements for the query results (S206). For example, analysis requirements include but are not limited to aggregate calculations, statistical analysis, trend prediction, etc.

[0072] In one scenario, if the metric processing requirements also include analysis requirements for query results, the query results are stored in a designated database that supports structured queries. This allows for flexible invocation of various SQL analysis functions to meet diverse analysis needs. For example, the designated database can be an in-memory database or another relational database. In-memory databases typically support a variety of data structures and query syntaxes, facilitating seamless integration with SQL statements generated by pre-trained language models, thereby streamlining data conversion and processing.

[0073] Next, see Figure 5The electronic device can generate an SQL table creation statement for the query results. The SQL table creation statement is used to describe the data organization structure of the query results in a specified database. Then, based on the SQL table creation statement and the indicator processing requirements, a second prompt word is generated. The second prompt word is input into the pre-trained language model so that the pre-trained language model generates an SQL statement for analyzing and processing the query results according to the analysis requirements in the indicator processing requirements. In this embodiment, the SQL table creation statement describes in detail the data organization structure of the query results in the specified database, such as field names, data types, indexes, etc., providing clear context information for the pre-trained language model. Based on this information, the pre-trained model can reduce ambiguity or errors that may occur when generating SQL statements and more accurately generate SQL statements that conform to the actual data storage method. Through the second prompt word, the pre-trained language model can automatically generate an analysis SQL statement that reflects the actual data structure, realizing the automated conversion from natural language requirements to data structure descriptions and then to complex query statements, reducing the workload of manual intervention and manual SQL writing.

[0074] Exemplarily, the second prompt word may also include at least one of the following prompt information: (1) The role played by the pre-trained language model: setting a specific role so that the model can generate a response based on the role. (2) Background information: providing sufficient background information so that the model can understand the context of the question. (3) The task to be performed by the pre-trained language model: letting the model know clearly what needs to be achieved, such as generating SQL statements. (4) At least one reference example; representative and differentiated examples can be added. Providing diverse examples helps the model better understand the task requirements and imitate the example style to complete the task.

[0075] In one possible implementation, the query results are stored in a designated database in the form of a data table. In the process of generating an SQL table creation statement for the query results, the electronic device can obtain the table name, column name, and column data type of the query results stored in the form of a data table from the designated database; embed the table name, column name, and column data type of the query results into a preset SQL table creation statement template to obtain an SQL table creation statement for the query results. The SQL table creation statement template contains SQL keywords and syntax structures, such as CREATE TABLE IF NOT EXISTS, brackets, commas, and semicolons, and other fixed SQL syntax elements; and also has placeholders for representing the data table name, column name, and the data type of each column. String replacement can be used to replace each placeholder in the preset SQL table creation statement template with the corresponding data table name, column name, and data type of each column. This embodiment automatically obtains relevant metadata of the query results from the specified database and embeds it into the template, ensuring that the generated SQL table creation statement can accurately reflect the actual data structure of the query results, avoiding manual errors, and using a templating method to achieve automatic generation, which can reduce manual intervention and ensure that the SQL table creation statements generated in different scenarios are consistent in format and standardized, thereby improving the stability of the entire indicator processing process.

[0076] In another possible implementation, the query results are stored in a designated database in a non-table format, such as in CSV (Comma-Separated Values) format. CSV is a common text format for storing tabular data that uses a comma (,) as a delimiter to separate data fields, with each row representing a data record. When generating an SQL table creation statement for the query results, the electronic device can extract CSV metadata from the designated database and extract column names from the first row of the CSV file, which is typically the header row and contains the names of each column. The electronic device then uses the data in subsequent rows of the CSV file to infer the data type of each column. For example, the data format (number, date, text, etc.) can be detected to determine the data type for each column. If the CSV file does not directly provide a table name, the table name can be determined based on preset rules (e.g., file name or user input). After determining the table name, column name, and column data type, the query result's table name, column name, and column data type can be embedded into a pre-set SQL table creation statement template to generate an SQL table creation statement for the query results.

[0077] After obtaining the SQL statement, the electronic device can pass the SQL statement to the designated database for execution. After submitting the SQL statement, the designated database parses and calculates the SQL statement, and returns the indicator processing results after analysis and processing. Finally, the indicator processing results can be fed back to the user or displayed and used later.

[0078] This embodiment divides the indicator processing process into two stages: basic query and further SQL analysis and processing. It fully utilizes the advantages of pre-trained language models in natural language conversion and SQL generation, while avoiding the limitations of directly generating complex non-relational database query statements. The phased execution ensures that each step focuses on a single function, reducing the error rate in the overall operation, and automatically generates query statements and SQL statements, reducing the reliance on manually written complex query logic. There is no need to customize complex syntax conversion models for each non-relational database, reducing the development workload. For a variety of non-relational databases, such as time series databases, document databases, or columnar databases, the correctness of data acquisition can be ensured through the first stage of basic query, and then in-depth analysis can be carried out through the second stage of SQL analysis and processing.

[0079] In some embodiments, after submitting the SQL statement to the designated database for execution, if a second execution error message (such as syntax error, data mismatch, etc.) is received from the designated database, and the number of executions of the designated database does not reach the second retry number, the electronic device inputs the received second execution error message and the SQL statement to be modified into the pre-trained language model, so as to modify the SQL statement to be modified by the pre-trained language model with reference to the second execution error message, obtain the modified SQL statement, and resubmit the modified SQL statement to the designated database for execution. This embodiment effectively solves the problem of the lack of a self-repair mechanism in traditional data analysis, automatically repairs errors through a pre-trained language model, reduces the need for manual intervention, improves adaptive capabilities, ensures the smooth progress of query tasks, and reduces the time and cost of troubleshooting; and by setting the second retry number, it can avoid endless retries that consume the computing and storage resources of the electronic device. The specific value of the second retry number can be set according to the actual application scenario.

[0080] If an execution error message is received from the specified database and the execution count of the specified database reaches the second retry count, a prompt message indicating that the indicator processing failed is output to ensure that the user receives clear feedback.

[0081] In another case, see Figure 2B If the indicator processing requirement does not include an analysis requirement for the query result, the electronic device directly outputs the query result (S210).

[0082] For example, if a user enters "Calculate the average CPU usage over the past 7 days," the system first queries the time series data of CPU usage. Based on the analysis requirement (average value), the data is stored in the in-memory database, and a SQL statement is generated to execute AVG(cpu_usage) and return the result.

[0083] If a user enters "Query CPU usage data for the past 7 days," the system first queries the time series data for CPU usage. Because further analysis of this time series data is required, no additional calculations are performed and the required CPU usage time series data is directly returned.

[0084] In other words, whether to perform SQL analysis is determined by the user's performance requirements. If the query results themselves can meet the requirements, the electronic device directly returns the data, reducing computing costs, improving query efficiency, and avoiding unnecessary resource consumption.

[0085] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of the various technical features in the above embodiments also falls within the scope of disclosure of this specification.

[0086] In some embodiments, an embodiment of this specification further provides an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor implements any of the above methods by running the executable instructions.

[0087] Figure 6 This is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 6 At the hardware level, the device includes a processor 602, an internal bus 604, a network interface 606, a memory 608, and a non-volatile memory 610. Of course, it may also include hardware required for other functions. One or more embodiments of this specification can be implemented based on software, such as the processor 602 reading the corresponding computer program from the non-volatile memory 610 into the memory 608 and then running it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0088] In some embodiments, the indicator processing device can be applied to Figure 6 The device shown in the figure is used to implement the technical solution of this specification. The indicator processing device may include:

[0089] A demand receiving module is used to receive the indicator processing demand described in natural language input by the user;

[0090] A demand parsing module is used to determine the indicators to be queried and the query conditions indicated by the indicator processing requirements based on the indicator processing requirements and the non-relational database used to store indicator data;

[0091] The query module is used to use the pre-trained language model to generate a query statement that conforms to the syntax specification of the non-relational database and queries the data of the query indicator according to the query conditions, and submit the query statement to the non-relational database for execution to obtain the query result;

[0092] The analysis module is used to store the query results in a designated database that supports structured queries if the indicator processing requirements also include analysis requirements for the query results, use the pre-trained language model to generate SQL statements for analyzing and processing the query results according to the analysis requirements, and submit the SQL statements to the designated database for execution to obtain the indicator processing results.

[0093] Exemplarily, the requirement parsing module is specifically used to generate a first prompt word based on the indicator processing requirements and all indicators contained in the non-relational database; input the first prompt word into the pre-trained language model, so that the pre-trained language model can determine the indicator to be queried indicated by the indicator processing requirements from all indicators contained in the non-relational database, and extract the query conditions of the indicator to be queried from the indicator processing requirements.

[0094] Exemplarily, the requirement parsing module is specifically used to convert the indicator processing requirement into a first embedding vector, and convert each indicator contained in the non-relational database into a second embedding vector; based on the similarity between the first embedding vector and each second embedding vector, determine the indicator to be queried indicated by the indicator processing requirement.

[0095] Exemplarily, the indicator processing device also includes an analysis requirement judgment module, which is used to use a pre-trained language model to determine whether the indicator processing requirement also includes an analysis requirement for the query result.

[0096] The indicator processing device also includes an output module, which is used to output the query result if there is no analysis requirement for the query result in the indicator processing requirement.

[0097] Exemplarily, the analysis module is specifically used to generate an SQL table creation statement for the query results, where the SQL table creation statement is used to describe the data organization structure of the query results in a specified database; based on the SQL table creation statement and the indicator processing requirements, a second prompt word is generated; the second prompt word is input into the pre-trained language model, so that the pre-trained language model generates an SQL statement for analyzing and processing the query results according to the analysis requirements in the indicator processing requirements.

[0098] Exemplarily, the analysis module is specifically used to obtain the table name, column name and column data type of the query results stored in the form of a data table from a specified database; embed the table name, column name and column data type of the query results into a preset SQL table creation statement template to obtain an SQL table creation statement for the query results.

[0099] Exemplarily, the indicator processing device also includes a modification module, which is used to modify the query statement using a pre-trained language model with reference to the first execution error information if a first execution error message returned by the non-relational database is received and the number of executions of the non-relational database has not reached the first retry number, to obtain a modified query statement, and resubmit the modified query statement to the non-relational database for execution.

[0100] The indicator processing device also includes an output module, which is used to output a prompt message of indicator processing failure if an execution error message returned by the non-relational database is received and the execution number of the non-relational database reaches the first retry number.

[0101] Exemplarily, the indicator processing device also includes a modification module, which is used to modify the SQL statement with reference to the second execution error information using a pre-trained language model if a second execution error message returned by the specified database is received and the number of executions of the specified database has not reached the second retry number, to obtain a modified SQL statement, and resubmit the modified SQL statement to the specified database for execution.

[0102] The indicator processing device also includes an output module, which is used to output a prompt message of indicator processing failure if an execution error message returned by the designated database is received and the execution number of the designated database reaches the second retry number.

[0103] Exemplarily, the designated database includes an in-memory database.

[0104] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0105] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor implements the steps of the method described in any of the above embodiments by running the executable instructions.

[0106] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0107] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0108] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.

[0109] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of one or more embodiments of this specification.

Claims

1. An indicator processing method, comprising: Receive indicator processing requirements described in natural language from users; Based on the indicator processing requirement and a non-relational database for storing indicator data, determining the indicator to be queried and the query condition indicated by the indicator processing requirement; Generate a query statement that conforms to the syntax specification of the non-relational database and queries the data of the indicator to be queried according to the query condition by using the pre-trained language model, and submit the query statement to the non-relational database for execution to obtain a query result; If the indicator processing requirements also include analysis requirements for the query results, the query results are stored in a designated database that supports structured queries, and the pre-trained language model is used to generate an SQL statement for analyzing and processing the query results according to the analysis requirements, and the SQL statement is submitted to the designated database for execution to obtain the indicator processing results.

2. The method according to claim 1, wherein determining the to-be-queried indicator and query condition indicated by the indicator processing requirement based on the indicator processing requirement and a non-relational database for storing indicator data comprises: generating a first prompt word based on the indicator processing requirement and all indicators included in the non-relational database; The first prompt word is input into the pre-trained language model so that the pre-trained language model determines the indicator to be queried indicated by the indicator processing requirement from all indicators contained in the non-relational database, and extracts the query condition of the indicator to be queried from the indicator processing requirement.

3. The method according to claim 1, wherein determining the indicator to be queried indicated by the indicator processing requirement based on the indicator processing requirement and a non-relational database for storing indicator data comprises: Converting the indicator processing requirement into a first embedding vector, and converting each indicator included in the non-relational database into a second embedding vector; Based on the similarities between the first embedding vector and each of the second embedding vectors, the to-be-queried indicator indicated by the indicator processing requirement is determined.

4. The method according to claim 1, further comprising, after obtaining the query result: Using the pre-trained language model, determining whether the indicator processing requirement also includes an analysis requirement for the query result; The method further comprises: If the indicator processing requirement does not have an analysis requirement for the query result, the query result is output.

5. The method according to claim 1, wherein the step of using the pre-trained language model to generate an SQL statement for analyzing and processing the query result according to the analysis requirement comprises: Generate an SQL table creation statement for the query result, wherein the SQL table creation statement is used to describe the data organization structure of the query result in the specified database; Generate a second prompt word based on the SQL table creation statement and the indicator processing requirement; The second prompt word is input into the pre-trained language model so that the pre-trained language model generates an SQL statement for analyzing and processing the query result according to the analysis requirements in the indicator processing requirements.

6. The method according to claim 5, wherein generating an SQL table creation statement for the query result comprises: Obtaining, from the designated database, the table name, column name, and column data type of the query result stored in a data table format; The table name, column name and column data type of the query result are embedded into a preset SQL table creation statement template to obtain an SQL table creation statement for the query result.

7. The method according to claim 1, after submitting the query statement to the non-relational database for execution, further comprising: If a first execution error message returned by the non-relational database is received and the number of executions of the non-relational database has not reached a first retry number, modifying the query statement by using the pre-trained language model with reference to the first execution error message to obtain a modified query statement, and resubmitting the modified query statement to the non-relational database for execution; If an execution error message returned by the non-relational database is received and the number of executions of the non-relational database reaches the first number of retries, a prompt message indicating that the indicator processing has failed is output.

8. The method according to claim 1, after submitting the SQL statement to the designated database for execution, further comprising: If a second execution error message returned by the designated database is received and the number of executions of the designated database has not reached a second retry number, modifying the SQL statement by using the pre-trained language model with reference to the second execution error message to obtain a modified SQL statement, and resubmitting the modified SQL statement to the designated database for execution; If an execution error message returned by the designated database is received and the number of executions of the designated database reaches the second number of retries, a prompt message indicating that the indicator processing has failed is output.

9. The method according to any one of claims 1 to 8, wherein the designated database comprises an in-memory database.

10. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 9 by running the executable instructions.

11. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Network traffic query processing method and device, equipment, storage medium and program product

    CN121166986A