A log query statement generation method, device, equipment and storage medium

CN118152341BActive Publication Date: 2026-09-29BEIJING YOUTEJIE INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410296636.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2026-09-29
Estimated Expiration
2044-03-15

AI Technical Summary

Technical Problem

[0003]目前对大模型的利用,自然语言理解技术可能无法完全理解复杂或模糊的用户查询,特别是当涉及到特定领域的术语或非常特定的查询要求时

Benefits of technology

[0023]根据本发明的另一方面,提供了一种计算机可读存储介质,所述计算机可读存储介质存储有计算机指令,所述计算机指令用于使处理器执行时实现本发明任一实施例所述的一种日志查询语句生成方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118152341B_ABST
    Figure CN118152341B_ABST
Patent Text Reader

Abstract

The application discloses a log query statement generation method and device, equipment and a storage medium. Including: obtaining the user input question text, dividing the question text through the pre-trained word segmentation model to generate the word segmentation text; obtain the vector database, generate each enhanced field according to the vector database and the word segmentation text; combine each enhanced field and the question text to generate the log query statement. By dividing the user input question text to generate the word segmentation text, and combining the vector database for field enhancement, finally splicing each enhanced field and the question text, outputting the log query statement through the large model, which helps to more accurately understand the user's query intention. The user's question text is combined with historical data and scene classification, so that the generated log query statement is more accurate and effective. Through dynamic grammar checking and revision, periodic quality inspection and optimization of historical samples, the long-term effectiveness and accuracy of the system are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of log query technology, and in particular to a method, apparatus, device and storage medium for generating log query statements. Background Technology

[0002] With the rapid development of artificial intelligence technology, current log analysis systems can greatly reduce the difficulty of use for users by leveraging new technologies to generate high-quality log query statements. Generating log query statements from user-provided text information is a multi-step process involving understanding the user's query intent, processing log data in specific formats, and generating accurate query commands.

[0003] Currently, when utilizing large models, natural language understanding technology may not be able to fully comprehend complex or fuzzy user queries, especially when domain-specific terminology or very specific query requirements are involved. For applications requiring real-time or near-real-time queries, the inference time of large models is also a bottleneck limiting real-time response. Processing large or complex log data consumes significant computing resources. Furthermore, as log formats, system architectures, or business needs change, large models also require regular updates and maintenance to maintain their accuracy and relevance. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and storage medium for generating log query statements, so as to accurately understand the user's query intent and generate accurate log query statements.

[0005] According to one aspect of the present invention, a method for generating log query statements is provided, the method comprising:

[0006] The system obtains the user's input question text and divides it into segmented text using a pre-trained word segmentation model. The segmented text includes the input time range, input data source, input filtering conditions, and input statistical scenarios.

[0007] Obtain the vector database, and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements and their corresponding scenario classifications.

[0008] Combine the enhanced fields and the query text to generate a log query statement.

[0009] Optionally, various enhancement fields are generated based on the vector database and the segmented text, including: data structure enhancement based on the vector database, input data source, and input filtering conditions to obtain the first enhancement field; keyword enhancement based on the input data source and input filtering conditions to obtain the second enhancement field; and scenario enhancement based on the vector database and input statistical scenarios to obtain the third enhancement field.

[0010] Optionally, data structure enhancement is performed based on the vector database, input data source, and input filtering conditions to obtain the first enhanced field, including: obtaining log index information from the log query database and determining the historical data source corresponding to each log index information; filtering each historical data source through the input data source and taking the historical data source matching the input data source as the target data source; taking the log index information corresponding to the target data source as the target index information and obtaining the key field corresponding to the target index information through a first preset interface; segmenting the input filtering conditions through a word segmentation model to generate each word segmentation condition; matching each historical query statement through each word segmentation condition to obtain the matching query statement corresponding to each word segmentation condition, taking the specified field in the matching query statement as the matching field; and taking the key field and the matching field as the first enhanced field.

[0011] Optionally, keyword enhancement is performed based on the input data source and input filtering conditions to obtain a second enhanced field, including: obtaining the original log text corresponding to the input data source through a second preset interface, and extracting each fixed text from the original log text in a specified manner; determining the first similarity between the input filtering conditions and each fixed text; and using the fixed text with the first similarity greater than a first threshold as the second enhanced field.

[0012] Optionally, scene enhancement is performed based on the vector database and the input statistical scenario to obtain a third enhancement field, including: determining the target scene category corresponding to the input statistical scenario; taking each historical query statement in the vector database corresponding to the target scene category as the target query statement; determining the second similarity between the question text and each target query statement, and arranging each target query statement in descending order of the second similarity to generate a historical sample list; and extracting the target query statements from the historical sample list in a specified number of steps as the third enhancement field.

[0013] Optionally, after combining each enhanced field and the query text to generate a log query statement, the method further includes: obtaining preset syntax validation rules; determining whether the log query statement meets the syntax validation rules; if so, determining that the syntax validation result is valid; otherwise, determining that the syntax validation result is invalid; generating a prompt message based on the syntax validation result; and sending an alarm to the prompt message in a specified manner.

[0014] Optionally, the method further includes: sequentially taking each scene category in the vector database as the scene category to be optimized; determining the third similarity corresponding to each historical query statement in the scene category to be optimized according to a preset period; taking the historical query statements with a third similarity greater than the second threshold as the query statements to be optimized; and optimizing the vector database based on the query statements to be optimized.

[0015] According to another aspect of the present invention, a log query statement generation apparatus is provided, the apparatus comprising:

[0016] The word segmentation text generation module is used to obtain the question text input by the user, and to segment the question text into word segments using a pre-trained model. The word segmentation text includes the input time range, input data source, input filtering conditions, and input statistical scenarios.

[0017] The enhanced field generation module is used to obtain the vector database and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements.

[0018] The log query statement generation module is used to combine various enhanced fields and query text to generate log query statements.

[0019] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0020] At least one processor; and

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute a log query statement generation method according to any embodiment of the present invention.

[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement a log query statement generation method according to any embodiment of the present invention.

[0024] The technical solution of this invention involves segmenting the user-input query text to generate word-segmented text, enhancing fields using a vector database, and finally concatenating the enhanced fields with the query text. This is then used to output a log query statement through a large model, which helps to more accurately understand the user's query intent. Combining the user's query text with historical data and scenario classification makes the generated log query statement more accurate and effective. Dynamic syntax validation and revision, along with regular quality checks and optimizations of historical examples, ensure the long-term effectiveness and accuracy of the system.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of a log query statement generation method provided in Embodiment 1 of the present invention;

[0028] Figure 2 This is a flowchart of another log query statement generation method provided in Embodiment 1 of the present invention;

[0029] Figure 3 This is a flowchart of another log query statement generation method provided in Embodiment 2 of the present invention;

[0030] Figure 4 This is a schematic diagram of a log query statement generation device according to Embodiment 3 of the present invention;

[0031] Figure 5 This is a schematic diagram of the structure of an electronic device that implements a log query statement generation method according to an embodiment of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] Example 1

[0035] Figure 1 This is a flowchart illustrating a log query statement generation method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where users perform log queries. The method can be executed by a log query statement generation device, which can be implemented in hardware and / or software and can be configured in a computer controller. Figure 1 As shown, the method includes:

[0036] S110. Obtain the question text input by the user, and divide the question text into segmented texts using a pre-trained word segmentation model. The segmented texts include the input time range, input data source, input filtering conditions, and input statistical scenarios.

[0037] In this context, the query text refers to the text content of a question or request entered by the user in response to a log query requirement. In the technical solution of this invention embodiment, the query text is the text input into the log query system. The system analyzes and understands it, and attempts to answer the user's question or provide relevant information. Word segmentation is a natural language processing technique used to divide text into meaningful lexical units for subsequent processing and analysis. Word segmentation models are typically based on statistical machine learning or deep learning techniques. Through learning and training on large amounts of text data, the model can identify word boundaries in Chinese text and segment the text into individual words. A log query statement is a statement used to query log files or log records in a database. It is typically used for monitoring system operation status, diagnosing faults, and analyzing performance.

[0038] Specifically, the controller needs to acquire the user's input query text. The query text may contain a variety of information, such as the input time range, input data source, input filtering conditions, and input statistical scenarios. To better understand this text, a pre-trained word segmentation model can be used to segment it to generate word-segmented text. The controller refers to a computer controller that includes a large language model.

[0039] In one specific implementation, the word segmentation text corresponding to the question text "the number of times Windows login failed today" is: today, Windows, login failed and number of times.

[0040] S120. Obtain the vector database and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements and their corresponding scenario classifications.

[0041] Specifically, the controller retrieves a vector database and generates enhanced fields based on the vector database and the segmented text. The vector database contains historical query statements and their corresponding scenario classifications. By analyzing and learning from this historical data, the controller can generate enhanced fields relevant to the user's query to improve the accuracy and comprehensiveness of the query.

[0042] Figure 2 This invention provides a flowchart of a log query statement generation method according to Embodiment 1. Step S120 mainly includes the following steps S121 to S123:

[0043] S121. Perform data structure enhancement based on the vector database, input data source, and input filtering conditions to obtain the first enhanced field.

[0044] Specifically, the controller can perform structured processing on the data in the vector database based on the characteristics and filtering conditions of the input data to extract information relevant to the input text. For example, if the input text contains specific data sources and filtering conditions, we can extract the corresponding data structure from the vector database to generate the first enhanced field.

[0045] Optionally, data structure enhancement is performed based on the vector database, input data source, and input filtering conditions to obtain the first enhanced field, including: obtaining log index information from the log query database and determining the historical data source corresponding to each log index information; filtering each historical data source through the input data source and taking the historical data source matching the input data source as the target data source; taking the log index information corresponding to the target data source as the target index information and obtaining the key field corresponding to the target index information through a first preset interface; segmenting the input filtering conditions through a word segmentation model to generate each word segmentation condition; matching each historical query statement through each word segmentation condition to obtain the matching query statement corresponding to each word segmentation condition, taking the specified field in the matching query statement as the matching field; and taking the key field and the matching field as the first enhanced field.

[0046] Specifically, the controller needs to connect to the log query database and retrieve the log index information. Log index information can include the log's timestamp, type, and content. Then, it needs to determine the historical data source corresponding to each log index. These historical data sources may come from different data sources, such as different servers or different applications. Furthermore, the controller can categorize the data sources, including standardized sources and custom sources specific to the current environment. Standardized sources include operating systems, firewalls, switches, middleware, etc., and the corresponding appnames / tags can be built into the large model through fine-tuning training. Custom sources are typically the Chinese and English names of business systems, requiring the large model to retrieve the corresponding appnames from a dictionary based on similarity matching.

[0047] For example, a Cisco firewall corresponds to the appname:firewall tag:cisco. The risk control system obtains the appname:rkm based on the similarity of the custom "Risk Management Platform / Risk Management / rkm".

[0048] Furthermore, the controller performs data structure acquisition, comparing the input data source with various historical data sources. If a historical data source matches the input data source, it is used as the target data source. For example, if the input data source is a specific server, historical data sources related to that server can be used as the target data source. Then, the log index information corresponding to the target data source is used as the target index information, and key fields corresponding to the target index information are obtained through a first preset interface. These key fields may include the log timestamp, log type, and log content. By classifying data sources and extracting metadata, a deeper understanding of the log data is achieved, which helps generate more accurate query statements.

[0049] It should be noted that the data structure includes "fields of interest" and "fields of favorites" corresponding to the appname / tag; and a list of possible values ​​for text-type fields with fewer than 50 unique entries. Once the data source is identified, the field list for that source is obtained through the first preset interface, with the "favorites" field having a higher weight; and for non-high-cardinality fields, all possible values ​​are then retrieved. For example, firewall logs include fields such as firewall.src_ip, firewall.dst_ip, and firewall.dst_port. Access logs include fields such as apache.method for GET and POST requests.

[0050] Specifically, the controller segments the filtering conditions and matches them against historical queries. Each segmentation condition is compared with a historical query; if a historical query contains that segmentation condition, it is used as the matching query. Finally, specified fields from the matching queries are used as the matching fields; these fields can be the data source or field names. This method of using a vector database for reverse filtering and real-time metadata retrieval provides a new and efficient way to process unstructured log data.

[0051] In one specific implementation, the controller can segment the "filter criteria" into words and look up the corresponding data source and field name in the vector database. For example, if the user's query only contains "UDP514", the words are segmented as UDP and 514, and then the vector query finds that the protocol field value includes UDP and the dst_port field value includes 514.

[0052] S122. Perform keyword enhancement based on the input data source and input filtering conditions to obtain a second enhanced field.

[0053] Optionally, keyword enhancement is performed based on the input data source and input filtering conditions to obtain a second enhanced field, including: obtaining the original log text corresponding to the input data source through a second preset interface, and extracting each fixed text from the original log text in a specified manner; determining the first similarity between the input filtering conditions and each fixed text; and using the fixed text with the first similarity greater than a first threshold as the second enhanced field.

[0054] Specifically, the second preset interface can be used to obtain the original log text corresponding to the input data source. Then, fixed text is extracted from the original log text according to a specified method. The specified method can be clustering learning and template extraction with dynamic parameters removed. The fixed text may be certain keywords or specific strings in the original log text. The method for extracting the fixed text can be selected according to the specific situation, such as using regular expressions or string matching. The controller will perform similarity matching between the filtering conditions and these fixed texts to determine the first similarity and obtain the actual original text query keywords.

[0055] The first similarity refers to the degree of similarity between the input filter condition and each fixed text. A higher similarity indicates a closer relationship between the input filter condition and each fixed text. The controller will use fixed text with a first similarity greater than a first threshold as the second enhancement field. The first threshold is a preset threshold used to determine the level of similarity. If the first similarity is greater than the first threshold, it indicates a relatively close relationship between the input filter condition and that fixed text, and therefore that fixed text can be used as the second enhancement field.

[0056] For example, the template corresponding to "Task Did Not Load:123" is "Task Did Not Load:<*>", with the enhanced keyword "task did not load". Only then will the user's question "task failed" not generate "taskfailure".

[0057] S123. Perform scene enhancement based on the vector database and input statistical scenarios to obtain the third enhancement field.

[0058] Specifically, to enhance the scenarios, developers first categorize them based on expert experience, dividing them into common scenarios. (There are hundreds of log query statement functions, making syntactic categorization unsuitable.) For example, scenarios are categorized based on statistical methods: group statistics, trend statistics, percentile statistics, geographical distribution statistics, and year-on-year / month-on-month statistics, etc. If a scenario falls outside these fixed usage scenarios, it is classified as another scenario. Simultaneously, existing query history records are retrieved. (The log query statements in this system are in pipe form, using "|" to separate statements, chaining multiple commands together, breaking down complex tasks into simpler ones for divide-and-conquer solutions, making it ideal for large models to interpret step-by-step into corresponding statement descriptions.) This yields question-and-answer pairs. The large model then performs scenario categorization, storing all the results in a vector database.

[0059] Optionally, scene enhancement is performed based on the vector database and the input statistical scenario to obtain a third enhancement field, including: determining the target scene category corresponding to the input statistical scenario; taking each historical query statement in the vector database corresponding to the target scene category as the target query statement; determining the second similarity between the question text and each target query statement, and arranging each target query statement in descending order of the second similarity to generate a historical sample list; and extracting the target query statements from the historical sample list in a specified number of steps as the third enhancement field.

[0060] Specifically, the controller determines the target scenario category corresponding to the input statistical scenario. The target scenario category refers to the scenario class to which the input statistical scenario belongs. Then, it uses historical query statements from the vector database corresponding to the target scenario category as target query statements. These historical query statements are those previously executed under that scenario category and can be used as references. Next, it determines the second similarity between the query text and each target query statement. The second similarity refers to the degree of similarity between the query text and each target query statement. Finally, the controller arranges the target query statements in descending order of second similarity to generate a historical sample list. The historical sample list is a list containing all target query statements, arranged in descending order of similarity. The specified quantity refers to the number of target query statements to be extracted, as defined by the user. Users can choose the number of target query statements to extract based on specific needs. For example, the specified quantity can be 3, in which case the controller will search the historical samples for the three most similar samples under the same scenario category.

[0061] S130. Combine the enhanced fields and the query text to generate a log query statement.

[0062] Specifically, the controller will combine the preceding enhancement fields with the user's query to form the final log query statement. If the preceding steps fail and the scenario classification is "other," then the query relies solely on the zero-shot capability of the large model and will alert the user that the query is not highly credible.

[0063] Optionally, after combining each enhanced field and the query text to generate a log query statement, the method further includes: obtaining preset syntax validation rules; determining whether the log query statement meets the syntax validation rules; if so, determining that the syntax validation result is valid; otherwise, determining that the syntax validation result is invalid; generating a prompt message based on the syntax validation result; and sending an alarm to the prompt message in a specified manner.

[0064] The syntax validation rules are user-defined predefined rules used to check whether log query statements conform to the syntax rules. The controller determines whether the log query statement meets the syntax validation rules. If the log query statement meets the syntax validation rules, the syntax validation result is determined to be valid; otherwise, the syntax validation result is determined to be invalid. Then, a prompt message is generated based on the syntax validation result. The prompt message can be an error message indicating a syntax error in the log query statement. Finally, the prompt message is sent to the system in a specified manner. The alert method can be sending an email, etc. If an error is reported, the controller will return the error message for modification. If repeated modifications still fail the validation, the controller will report the result to the user, explicitly prompting the user to make adjustments themselves.

[0065] The technical solution of this invention involves segmenting the user-input query text to generate word-segmented text, enhancing fields using a vector database, and finally concatenating the enhanced fields with the query text. This is then used to output a log query statement through a large model, which helps to more accurately understand the user's query intent. Combining the user's query text with historical data and scenario classification makes the generated log query statement more accurate and effective. Dynamic syntax validation and revision, along with regular quality checks and optimizations of historical examples, ensure the long-term effectiveness and accuracy of the system.

[0066] Example 2

[0067] Figure 3 This is a flowchart of a log query statement generation method provided in Embodiment 2 of the present invention. This embodiment adds a vector database optimization process based on Embodiment 1 above. Figure 3 As shown, the method includes:

[0068] S210. Sequentially classify each scene category in the vector database as the scene category to be optimized.

[0069] It should be noted that during system operation, as the query history accumulates, the distribution of samples across different scenarios may become unbalanced. This imbalance can negatively impact vector retrieval results and the few-shot learning performance of the large model. To address this issue, an external optimization task can be added. This task periodically performs large model classification on the historical records to better understand query statements in different scenarios.

[0070] S220. Determine the third similarity of each historical query statement in the scenario category to be optimized according to the preset period.

[0071] Specifically, historical query statements from the scene classification to be optimized can be collected from the vector database at preset intervals, and meaningful features can be extracted from these historical queries. These features can be used to describe each historical query statement. The extracted features are then used to calculate the similarity between the historical queries, such as cosine similarity, Euclidean distance, or Jaccard similarity. The historical queries are then sorted according to their similarity to select those with higher similarity.

[0072] S230. Select historical query statements with a similarity greater than the second threshold in the third similarity category as query statements to be optimized.

[0073] Specifically, historical query statements with a third similarity greater than the second threshold can be marked as query statements to be optimized. Query statements to be optimized are those with high similarity that may have optimization potential.

[0074] S240. Optimize the vector database based on the query statement to be optimized.

[0075] In one specific implementation, for each scene category, query statements can be converted into vectors, and the distances between them can be calculated. Then, only 5-10 results with large distances are retained, while redundant statements with too small a similarity distance are deleted. The technical solution of this invention reduces redundant information in historical records, improves the accuracy of vector retrieval results, and enhances the few-sample learning effect of large models. It also better balances the sample distribution across different scenes, thereby improving the overall performance and efficiency of the system.

[0076] For example, of the three statements firewall.action:deny|stats count()by firewall.dst_ip, *|statspct(apache.resp_time,50) and *|stats pct(apache.resp_time,75), the latter two are too similar, so only one needs to be kept.

[0077] The technical solution of this invention involves segmenting the user-input query text to generate word-segmented text, enhancing fields using a vector database, and finally concatenating the enhanced fields with the query text. This is then used to output a log query statement through a large model, which helps to more accurately understand the user's query intent. Combining the user's query text with historical data and scenario classification makes the generated log query statement more accurate and effective. Dynamic syntax validation and revision, along with regular quality checks and optimizations of historical examples, ensure the long-term effectiveness and accuracy of the system.

[0078] Example 3

[0079] Figure 4 This is a schematic diagram of a log query statement generation device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: a word segmentation text generation module 310, which is used to acquire the question text input by the user, and to segment the question text by a pre-trained model to generate word segmentation text, wherein the word segmentation text includes the input time range, input data source, input filtering conditions and input statistical scenario;

[0080] The enhanced field generation module 320 is used to obtain a vector database and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements.

[0081] The log query statement generation module 330 is used to combine the enhanced fields and the query text to generate log query statements.

[0082] Optionally, the enhanced field generation module 320 specifically includes: a data structure enhancement unit, used to: enhance the data structure according to the vector database, the input data source and the input filtering conditions to obtain a first enhanced field; a keyword enhancement unit, used to: enhance keywords according to the input data source and the input filtering conditions to obtain a second enhanced field; and a scene enhancement unit, used to: enhance the scene according to the vector database and the input statistical scene to obtain a third enhanced field.

[0083] Optionally, the data structure enhancement unit is specifically used for: obtaining log index information from the log query database and determining the historical data source corresponding to each log index; filtering each historical data source by input data source and taking the historical data source matching the input data source as the target data source; taking the log index information corresponding to the target data source as the target index information and obtaining the key field corresponding to the target index information through a first preset interface; segmenting the input filtering conditions into words using a word segmentation model to generate each word segmentation condition; matching each historical query statement with each word segmentation condition to obtain the matching query statement corresponding to each word segmentation condition, taking the specified field in the matching query statement as the matching field; and taking the key field and the matching field as the first enhancement field.

[0084] Optionally, the keyword enhancement unit is specifically used to: obtain the original log text corresponding to the input data source through a second preset interface, and extract each fixed text from the original log text in a specified manner; determine the first similarity between the input filtering conditions and each fixed text; and use the fixed text with the first similarity greater than the first threshold as the second enhancement field.

[0085] Optionally, the scene enhancement unit is specifically used for: determining the target scene category corresponding to the input statistical scene; taking each historical query statement corresponding to the target scene category in the vector database as the target query statement; determining the second similarity between the query text and each target query statement, and arranging each target query statement in descending order of the second similarity to generate a historical sample list; and extracting the target query statements from the historical sample list in a specified number of steps as the third enhancement field.

[0086] Optionally, the device also includes: a syntax verification module, used to obtain preset syntax verification rules after combining each enhanced field and the query text to generate a log query statement; determine whether the log query statement meets the syntax verification rules; if so, determine that the syntax verification result is verified as passed; otherwise, determine that the syntax verification result is verified as failed; generate a prompt message based on the syntax verification result; and send an alarm to the prompt message in a specified manner.

[0087] Optionally, the device further includes a vector database optimization module, used to: sequentially classify each scene category in the vector database as a scene category to be optimized; determine the third similarity corresponding to each historical query statement in the scene category to be optimized according to a preset period; take the historical query statements with a third similarity greater than a second threshold as query statements to be optimized; and optimize the vector database based on the query statements to be optimized.

[0088] The technical solution of this invention involves segmenting the user-input query text to generate word-segmented text, enhancing fields using a vector database, and finally concatenating the enhanced fields with the query text. This is then used to output a log query statement through a large model, which helps to more accurately understand the user's query intent. Combining the user's query text with historical data and scenario classification makes the generated log query statement more accurate and effective. Dynamic syntax validation and revision, along with regular quality checks and optimizations of historical examples, ensure the long-term effectiveness and accuracy of the system.

[0089] The log query statement generation device provided in this embodiment of the invention can execute a log query statement generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0090] Example 4

[0091] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0092] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0093] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0094] Processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a log query statement generation method. That is: acquiring the query text input by the user; segmenting the query text using a pre-trained word segmentation model to generate word segmented text, wherein the word segmented text includes the input time range, input data source, input filtering conditions, and input statistical scenario; acquiring a vector database; generating various enhanced fields based on the vector database and the word segmented text, wherein the vector database includes historical query statements and their corresponding scenario classifications; and combining the enhanced fields and the query text to generate a log query statement.

[0095] In some embodiments, a log query statement generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the log query statement generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute a log query statement generation method by any other suitable means (e.g., by means of firmware).

[0096] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0097] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0098] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0099] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0100] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0101] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0102] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for generating log query statements, characterized in that, include: The system obtains the question text input by the user, and divides the question text into segmented text using a pre-trained word segmentation model to generate segmented text. The segmented text includes the input time range, input data source, input filtering conditions, and input statistical scenario. Obtain a vector database, and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements and their corresponding scenario classifications. The scenario classifications include group statistics, trend statistics, percentile statistics, geographical distribution statistics, and year-on-year and month-on-month statistics. The enhanced fields and the query text are combined to generate a log query statement; The step of generating each enhanced field based on the vector database and the segmented text includes: Data structure enhancement is performed based on the vector database, the input data source, and the input filtering conditions to obtain a first enhanced field; Based on the input data source and the input filtering conditions, keyword enhancement is performed to obtain a second enhanced field; Scene enhancement is performed based on the vector database and the input statistical scenario to obtain a third enhancement field; The step of performing scene enhancement based on the vector database and the input statistical scenario to obtain the third enhancement field includes: Determine the target scene category corresponding to the input statistical scene; Use the historical query statements in the vector database that correspond to the target scene classification as the target query statements; Determine the second similarity between the query text and each of the target query statements, and arrange the target query statements in descending order of the second similarity to generate a historical sample list; The target query statement is extracted sequentially from the historical sample list according to a specified number and used as the third enhanced field.

2. The method according to claim 1, characterized in that, The step of performing data structure enhancement based on the vector database, the input data source, and the input filtering conditions to obtain the first enhanced field includes: Obtain the log index information of each log query database and determine the historical data source corresponding to each log index information; The historical data sources are filtered by the input data sources, and the historical data sources that match the input data sources are taken as the target data sources; The log index information corresponding to the target data source is used as the target index information, and the key fields corresponding to the target index information are obtained through the first preset interface; The input filtering conditions are segmented using the word segmentation model to generate each segmentation condition. Each of the historical query statements is matched using the segmentation conditions to obtain the matching query statement corresponding to each segmentation condition, and the specified field in the matching query statement is used as the matching field. The key field and the matching field are used as the first enhanced field.

3. The method according to claim 1, characterized in that, The step of enhancing keywords based on the input data source and the input filtering conditions to obtain a second enhanced field includes: The log text corresponding to the input data source is obtained through the second preset interface, and each fixed text in the log text is extracted in a specified manner. Determine the first similarity between the input filtering conditions and each of the fixed texts; The fixed text with a first similarity greater than a first threshold is used as the second enhanced field.

4. The method according to claim 1, characterized in that, After combining the enhanced fields and the query text to generate a log query statement, the method further includes: Get the preset syntax validation rules; Determine whether the log query statement meets the syntax validation rules; if so, determine that the syntax validation result is valid. Otherwise, if the syntax verification result is determined to be a failure, a prompt message is generated based on the syntax verification result, and the prompt message is used to trigger an alarm in a specified manner.

5. The method according to claim 1, characterized in that, The method further includes: Each scene classification in the vector database is sequentially used as the scene classification to be optimized. The third similarity of each historical query statement in the scenario classification to be optimized is determined according to a preset period. Historical query statements with a third similarity greater than the second threshold are considered as query statements to be optimized. Optimize the vector database based on the query statement to be optimized.

6. A log query statement generation device, characterized in that, include: The word segmentation text generation module is used to obtain the question text input by the user, and to divide the question text into word segments using a pre-trained word segmentation model to generate word segmented text. The word segmented text includes the input time range, input data source, input filtering conditions, and input statistical scenarios. An enhanced field generation module is used to obtain a vector database and generate various enhanced fields based on the vector database and the segmented text. The vector database includes historical query statements and their corresponding scene classifications. The scene classifications include group statistics, trend statistics, percentile statistics, geographical distribution statistics, and year-on-year and month-on-month statistics. The log query statement generation module is used to combine the enhanced fields and the query text to generate a log query statement. The enhanced field generation module specifically includes: a data structure enhancement unit, used to: enhance the data structure according to the vector database, the input data source, and the input filtering conditions to obtain a first enhanced field; a keyword enhancement unit, used to: enhance keywords according to the input data source and the input filtering conditions to obtain a second enhanced field; and a scene enhancement unit, used to: enhance the scene according to the vector database and the input statistical scene to obtain a third enhanced field. Specifically, the scene enhancement unit is used to: determine the target scene category corresponding to the input statistical scene; take each historical query statement in the vector database corresponding to the target scene category as the target query statement; determine the second similarity between the question text and each of the target query statements, and arrange each of the target query statements in descending order of the second similarity to generate a historical sample list; and extract the target query statements from the historical sample list in a specified number of steps as the third enhancement field.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A computer storage medium, characterized in that, The computer storage medium stores computer instructions that are used to cause a processor to execute the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Knowledge query method and device for vertical domain knowledge graph, computer equipment and storage medium

    CN117194616A

  • Method and apparatu for querying

    US20180365257A1