Log retrieval method and device based on natural language, medium and product

By employing a natural language-based log retrieval method and utilizing a large language model and a multi-level knowledge base matching mechanism, the natural language log retrieval problem is transformed into a professional query statement, solving the complexity and accuracy issues of traditional log retrieval methods and achieving efficient and accurate log retrieval.

CN120973899APending Publication Date: 2025-11-18BEIJING YOUTEJIE INFORMATION TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511089089.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18

Smart Images

  • Figure CN120973899A_ABST
    Figure CN120973899A_ABST
Patent Text Reader

Abstract

The invention discloses a log retrieval method and device based on a natural language, a medium and a product, and the method comprises the steps: in response to a natural language log retrieval problem of a user, extracting at least one piece of retrieval key information from the natural language log retrieval problem; matching each piece of retrieval key information with a knowledge base and historical environment data to generate corresponding alternative query field sets, and inputting the alternative query field sets into a large language model for screening; screening each alternative query field set by the large language model, determining a target query field corresponding to each piece of retrieval key information, and generating a log retrieval analysis statement aiming at a natural language log retrieval problem according to the target query field; according to the log retrieval analysis statement, retrieval is carried out in at least one log database, and user feedback is carried out on the target log set obtained through retrieval, according to the technical scheme of the embodiment of the invention, intelligence and automation of log retrieval are realized, and the use threshold of a user is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a log retrieval method, device, medium, and product based on natural language. Background Technology

[0002] As enterprises deepen their digital transformation, the scale and complexity of enterprise log data are growing exponentially. Therefore, achieving efficient log retrieval and analysis has become particularly important.

[0003] In existing technologies, traditional log retrieval methods typically employ command-line queries or fixed templates: command-line queries require users to master professional query syntax and pipe commands, and obtain target logs by combining filter conditions, time ranges, and statistical functions; fixed template methods require pre-configuring standard query templates, and users generate query statements by filling in parameters.

[0004] In developing this invention, the inventors discovered that while traditional command-line queries can retrieve logs, they are costly to learn and complex to operate, resulting in low user adoption. While fixed templates lower the barrier to entry, they struggle to meet the diverse needs of various business scenarios. Furthermore, due to differences in log formats and business scenarios, directly applying generic query solutions often leads to inaccurate search results, increasing the workload of operations and maintenance analysis. Summary of the Invention

[0005] This invention provides a log retrieval method, device, medium, and product based on natural language, which can realize the intelligent conversion of natural language into professional log query statements.

[0006] According to one aspect of the present invention, a log retrieval method based on natural language is provided, the method comprising:

[0007] In response to a user's natural language log retrieval question, extract at least one key retrieval information from the natural language log retrieval question;

[0008] Each key retrieval information is matched with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information, and each set of candidate query fields is provided to the large language model.

[0009] After filtering the set of candidate query fields by the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field.

[0010] The system performs a search in at least one log database based on the log retrieval and analysis statement, and provides the target log set obtained from the search to the user.

[0011] According to another aspect of the present invention, a log retrieval device based on natural language is provided, the device comprising:

[0012] The natural language processing module is used to respond to a user's natural language log retrieval question and extract at least one key retrieval information from the natural language log retrieval question.

[0013] The knowledge matching module is used to match each key search information with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key search information, and to provide each set of candidate query fields to the big language model.

[0014] The query statement generation module is used to filter the set of candidate query fields through a large language model, determine the target query fields corresponding to each key information to be retrieved, and then generate log retrieval analysis statements for natural language log retrieval problems based on each target query field.

[0015] The log retrieval module is used to perform retrieval in at least one log database based on log retrieval analysis statements, and to provide feedback to the user with the target log set obtained from the retrieval.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform a natural language-based log retrieval method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement a natural language-based log retrieval method as described in any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the steps of the method as described in any embodiment of the present invention.

[0022] The technical solution of this invention first responds to a user's natural language log retrieval question by extracting at least one key retrieval information from the question. Then, each key retrieval information is matched with knowledge in a knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information. These candidate query field sets are then provided to a large language model. Next, the large language model filters these candidate query field sets to determine the target query fields corresponding to each key retrieval information. Based on these target query fields, a log retrieval analysis statement is formed for the natural language log retrieval question. Finally, the log retrieval analysis statement is used to perform a retrieval in at least one log database, and the retrieved target log set is provided back to the user. This novel log retrieval method significantly lowers the user threshold through natural language interaction, allowing complex log retrieval without requiring expertise in query syntax. By utilizing the semantic understanding capabilities of a large language model combined with a multi-level knowledge base matching mechanism, the business accuracy and technical feasibility of the query statement are ensured.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a log retrieval method based on natural language according to Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of another log retrieval method based on natural language according to Embodiment 2 of the present invention;

[0027] Figure 3 This is a flowchart of another log retrieval method based on natural language according to Embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of a log retrieval device based on natural language according to Embodiment 4 of the present invention;

[0029] Figure 5 This is a schematic diagram of the structure of an electronic device that implements a log retrieval method based on natural language according to an embodiment of the present invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] Example 1

[0033] Figure 1 This is a flowchart of a log retrieval method based on natural language provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where efficient log retrieval and analysis can be achieved through natural language. The method can be executed by a log retrieval device based on natural language, which can be implemented in hardware and / or software and is generally configured in an electronic device.

[0034] Correspondingly, such as Figure 1 As shown, the method includes:

[0035] S110. In response to the user's natural language log retrieval question, extract at least one key retrieval information from the natural language log retrieval question.

[0036] Natural language log retrieval questions can be understood as log analysis requests submitted by users using everyday language (such as "show the services with the most errors in the last hour") rather than professional query syntax. These questions typically include elements such as business entities, time ranges, and query conditions. Retrieving key information can be understood as extracting the core query elements from the user's natural language query, such as business elements, time ranges, and analysis dimensions.

[0037] In this embodiment, when a log retrieval request expressed in natural language is received from a user, considering that directly processing the raw natural language request by a large language model may lead to inaccurate answers, the request is first processed in a structured manner by a semantic parsing engine. This process extracts the core data elements that constitute the query intent and converts them into standardized question statements that the large language model can accurately understand, thereby ensuring the accuracy of subsequent processing results.

[0038] Specifically, the semantic parsing engine extracts four main categories of structured elements from natural language requests: business entities (e.g., "order service"), query conditions (e.g., "500 errors"), time ranges (e.g., "past 1 hour"), and analytical dimensions (e.g., "number of occurrences"). Taking a user query "Please count the number of 500 errors that occurred in the order service within the past 1 hour" as an example, it is deconstructed into a clear combination of elements: the business object refers to the order service, the time window is limited to the most recent 1 hour, the filter condition is set to 500 errors, and a number of occurrences is requested. This structured element information is transformed into standardized input that the large language model can accurately process, effectively avoiding comprehension biases caused by natural language ambiguity.

[0039] S120. Match each key retrieval information with the knowledge in the knowledge base and historical environment data to obtain a set of candidate query fields corresponding to each key retrieval information, and provide each set of candidate query fields to the large language model.

[0040] The knowledge base can be understood as a hierarchical database containing common log fields, business terminology mappings, and session context. Historical environment data can be understood as dynamic context information accumulated during runtime, including user query habits, service node status, and intermediate session results. For example, it records frequently queried field combinations by users recently, or automatically excludes log sources from faulty nodes, providing real-time optimization basis for the current query.

[0041] The set of candidate query fields can be understood as a list of candidate fields generated through intelligent matching, including standard fields extracted from the knowledge base, derived fields inferred from historical behavior, and extended fields suggested by the large model. Each field is labeled with a matching score, data source, and field type, providing multi-dimensional selection criteria for query generation. The large language model can be understood as an artificial intelligence system trained on massive amounts of text, capable of understanding natural language and generating coherent text. It analyzes the input question or instruction, combines it with the language rules and knowledge learned during training, and outputs a response that conforms to semantic logic.

[0042] Furthermore, each key retrieval information corresponds to a set of candidate query fields, which contains one or more candidate query fields. These candidate query fields can be further understood as data columns or attributes to be searched in the database. For example, by querying the "name" column in the database, various member names can be obtained, such as "Zhang San" or "Li Si," etc. Continuing the previous example, if a user wants to query the number of times the error type "500 Error" occurred within a certain period, they need to search for the specific key retrieval information "500 Error" under the query field matching the "error type" in the database. In various databases, commonly used query fields matching error types can be "Fault" or "Error," etc. When the specific value of the query field cannot be clearly determined, both "Fault" and "Error" can be added to the set of candidate query fields matching the key retrieval information "500 Error," providing them to the large language model for subsequent filtering. In this embodiment, each key retrieval information extracted from the user's question (such as time range, business entity, error type, etc.) is matched against predefined fields in the knowledge base. Simultaneously, real-time information such as user query habits and system operating status recorded in historical environment data is combined to generate a set of candidate query fields. For example, for the key information "login failed," multiple related fields such as auth_status and login_result may be matched. These candidate field sets retain the original matching source and confidence score, and are ultimately submitted to the large language model in a structured data format for intelligent filtering and optimization.

[0043] S130. After filtering the set of candidate query fields through the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field.

[0044] In this embodiment, the semantic understanding capabilities of a large language model are used to perform in-depth analysis of candidate fields. By evaluating dimensions such as the matching degree between fields and user intent, relevance to business scenarios, and technical feasibility, the optimal field combination is intelligently selected. Based on the selection results, a search statement conforming to the syntax specifications of the target log system is automatically constructed. This statement can fully include core elements such as time limits, filtering conditions, and statistical requirements, ensuring that the generated query accurately expresses user needs and can be directly executed in the log system. For example, the system may intelligently combine key fields such as the selected time range, error code, and service name into a professional search instruction that conforms to a specific query syntax.

[0045] Specifically, in S120, a set of candidate query fields corresponding to each key retrieval information has been determined. For example, the set of candidate query fields corresponding to "the past hour" is {DATE; TIME}; the set of candidate query fields corresponding to "500 error" is {Fault; Error}, etc. Through parsing and reasoning of the large language model, the target query fields corresponding to each key retrieval information can be selected from the set of candidate query fields for database retrieval.

[0046] Furthermore, the target query field can be used as the key name, and the key information retrieved can be used as the key value to construct a query statement that conforms to the database query rules, that is, a log retrieval and analysis statement.

[0047] S140. Perform a search in at least one log database based on the log search and analysis statement, and provide the target log set obtained from the search to the user.

[0048] In this embodiment, the standardized query statement generated in the preceding steps is submitted to one or more log databases in the target log storage system to perform a retrieval operation. After obtaining the result dataset that meets the conditions, the original log data is standardized and encapsulated. Finally, the obtained target log set is returned as structured query results through a preset interface. This process completes the conversion and delivery from the query statement to the actual log data.

[0049] The technical solution of this invention first responds to a user's natural language log retrieval question by extracting at least one key retrieval information from the question. Then, each key retrieval information is matched with knowledge in a knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information. These candidate query field sets are then provided to a large language model. Next, the large language model filters these candidate query field sets to determine the target query fields corresponding to each key retrieval information. Based on these target query fields, a log retrieval analysis statement is formed for the natural language log retrieval question. Finally, the log retrieval analysis statement is used to perform a retrieval in at least one log database, and the retrieved target log set is provided back to the user. This novel log retrieval method significantly lowers the user threshold through natural language interaction, allowing complex log retrieval without requiring expertise in query syntax. By utilizing the semantic understanding capabilities of a large language model combined with a multi-level knowledge base matching mechanism, the business accuracy and technical feasibility of the query statement are ensured.

[0050] Example 2

[0051] Figure 2This is a flowchart of another log retrieval method based on natural language provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiments. Specifically, the operation of "extracting at least one key retrieval information in response to the user's natural language log retrieval question" has been refined.

[0052] Correspondingly, such as Figure 2 As shown, the method includes:

[0053] S210. Upon receiving a natural language question input by the user, obtain a set of natural language retrieval question templates, and construct a first prompt word template using the set of natural language retrieval question templates.

[0054] The natural language retrieval question template set can be understood as a pre-built standardized query pattern library, containing typical question frameworks for various log retrieval scenarios. The first prompt word template can be understood as a dynamically generated guidance instruction based on the matched question template, explicitly requiring the large language model to process the input according to a specific structure. That is, this natural language retrieval question template set defines example templates of log retrieval questions that the large language model can answer. It can be obtained by collecting, summarizing, processing, and handling various high-frequency log retrieval questions submitted by multiple users throughout history. By using the natural language retrieval question template set as prompt word templates (i.e., prompts) and inputting them along with the user's natural language question into the large language model, the large language module can quickly determine whether the natural language question falls within the answerable range.

[0055] In this embodiment, user input is initially categorized using a predefined set of natural language retrieval question templates. After selecting a matching template framework, a first prompt word template containing standardized parameter placeholders and format requirements is constructed. This template transforms the original natural language question into an input format that can be processed by a large language model. For example, "recent error reports" is transformed into a structured prompt such as "please complete the query conditions in the format of [time range] [error level]".

[0056] S220. The natural language question and the first prompt word template are provided to the large language model to obtain feedback results on whether the natural language question falls within the scope of the answer.

[0057] In this embodiment, the user's natural language question is integrated with the pre-built first prompt template and submitted to a large language model for intent recognition and scope determination. The large language model analyzes the combined input and evaluates whether the query content falls within the designed range of responsive questions. If the question is determined to be within the answerable range, the subsequent search process will continue; if the question is determined to be outside the range, the processing will be terminated immediately and a prompt message will be returned.

[0058] S230. If it is determined that the natural language question falls within the scope of the answer, a natural language log retrieval question is generated based on the natural language question, and at least one key retrieval information is extracted from the natural language log retrieval question.

[0059] In this embodiment, when it is confirmed that the user's question falls within the answerable range, a standard log query question is directly generated based on the original question content, and the key elements to be retrieved are extracted from this question. In a specific example, these core element units may include basic query conditions such as business objects, time parameters, and filtering conditions, providing structured input materials for subsequent query field matching and log retrieval analysis statement generation.

[0060] S240. Match each key retrieval information with the knowledge in the knowledge base and historical environment data to obtain a set of alternative query fields corresponding to each key retrieval information, and provide each set of alternative query fields to the large language model.

[0061] S250. After filtering the set of candidate query fields through the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field.

[0062] S260. Perform a search in at least one log database based on the log search and analysis statement, and provide the target log set obtained from the search to the user.

[0063] The technical solution of this invention, upon receiving a natural language question input by a user, acquires a set of natural language retrieval question templates and constructs a first prompt word template from them. The natural language question and the first prompt word template are jointly provided to a large language model to obtain feedback on whether it falls within the scope of the answer. If it is determined to fall within the scope, a natural language log retrieval question is generated based on the natural language question, and at least one key retrieval information is extracted. Each key retrieval information is matched with knowledge and historical environment data in the knowledge base to obtain a corresponding set of candidate query fields, which is then submitted to the large language model. The large language model filters each set of candidate query fields to determine the target query field, forming a log retrieval analysis statement. Finally, the log database is searched based on this statement, and the target log set is returned to the user. This novel log retrieval method improves the accuracy of natural language understanding through template-based preprocessing, ensures query compliance through multi-level filtering, and lays the foundation for subsequent retrieval through structured extraction of key information, significantly reducing the technical threshold. In particular, the template matching and scope verification mechanisms efficiently identify legitimate requests, and the accurately extracted retrieval elements ensure the quality of knowledge base matching, making the entire process both flexible and reliable.

[0064] Optionally, based on the above embodiments, a natural language log retrieval question can be generated according to the natural language query, which may specifically include:

[0065] Detect whether there is any prior historical dialogue information associated with the natural language question;

[0066] If so, obtain the result summary information that matches the previous historical dialogue information; wherein, the result summary information includes the previous historical natural language log retrieval question, and the information summary of the historical log set retrieved based on the previous historical natural language log retrieval question;

[0067] The natural language query is rewritten based on the results summary information to obtain the natural language log retrieval query;

[0068] If not, directly identify the natural language question as a natural language log retrieval question.

[0069] The preceding historical dialogue information can be understood as past dialogue records that are contextually related to the current query, including relevant questions previously asked by the user and the returned log retrieval results. For example, a dialogue chain formed when continuously querying different dimensional metrics of the same business module. The result summary information can be understood as a condensed expression of the historical dialogue, including the original question text of the previous query and key feature extraction from the corresponding log results. For example, core statistical indicators such as high-frequency error codes and time distribution extracted from a large number of error logs.

[0070] Generally, the process begins by checking if there are any related historical dialogue records for the current natural language question. This process analyzes the dialogue context to determine if the user's current question has a logical connection to previously discussed topics. When a relevant preceding dialogue is confirmed, the corresponding historical result summary information is extracted. This summary includes the content of the previously submitted natural language query (i.e., the preceding historical natural language log retrieval question) and a summary of the key information of the log results obtained after executing the query for that preceding historical natural language log retrieval question; in other words, an information summary of the historical log set.

[0071] Based on the result summary information obtained from previous historical dialogues, the current natural language question can be semantically expanded and context-adapted through question rewriting. This rewriting process automatically supplements necessary contextual information while maintaining the user's original intent, ensuring that the final generated natural language log retrieval question has complete contextual relevance.

[0072] If the detection confirms that there is no historical dialogue information that can be linked, the user's original question will be kept unchanged and directly input into the subsequent process as a natural language log retrieval question to be processed. This situation usually occurs in the initial query or in the case of a new question after a topic change.

[0073] Furthermore, based on the above embodiments, after providing user feedback on the retrieved target log set, the process may further include:

[0074] Information summarization is performed on the natural language log retrieval problem and the target log set to obtain a result summary of the current round of dialogue information for use in the next round of dialogue.

[0075] Generally, after the log retrieval results are returned, the entire interaction process is condensed to extract the core semantic elements of the natural language question and the key feature indicators of the target log set, generating a structured dialogue summary. This summary is equivalent to the result summary information of the previous historical dialogue information matching, containing a textual representation of the query intent and statistical features of the log results, such as high-frequency field values, abnormal time point distribution, and other core data features, and is stored in a standardized format for subsequent dialogue context association.

[0076] Example 3

[0077] Figure 3 This is a flowchart of another log retrieval method based on natural language provided in Embodiment 3 of the present invention. This embodiment is based on and optimized from the above embodiments. Specifically, the operation of "matching each key retrieval information with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information" has been refined.

[0078] Correspondingly, such as Figure 3 As shown, the method includes:

[0079] S310. In response to the user's natural language log retrieval question, extract at least one key retrieval information from the natural language log retrieval question.

[0080] S320. Sequentially obtain one key information for retrieval, use it as the current key information for retrieval, and initialize and construct the current candidate query field set that matches the current key information for retrieval.

[0081] In this embodiment, a single key piece of information is selected sequentially from the extracted set of key retrieval information as the current processing object, and an empty set of candidate query fields is initialized for this key piece of information. This newly created set will be used to store candidate query fields subsequently obtained through different matching channels, establishing possible query field mapping relationships for the current key piece of information. For example, when processing the key information "login failed," a corresponding empty set will be created, waiting to be filled with possible matching query fields through channels such as business knowledge base and general knowledge base.

[0082] S330. Match the current key information being retrieved against the pre-built business knowledge base. If a first query field is successfully matched, add the first query field to the current set of candidate query fields.

[0083] The business knowledge base can be understood as storing mapping relationships of professional fields strongly related to specific business scenarios. It includes the correspondence rules between internal business terms and technical fields (e.g., mapping "Gold Member" defined by the business department to "vip_level=3" in the database). These mapping rules are customized specifically for business scenarios and may differ from general mapping rules. The first query field can be understood as a professional field matched from the business knowledge base, reflecting the terminology mapping relationship under a specific business scenario. These fields have business-specific characteristics.

[0084] In this embodiment, the currently processed key information is matched against a pre-built business knowledge base. If a first query field associated with the key information is found in the business knowledge base (e.g., mapping "login failed" to the "auth_result=fail" field defined in the business knowledge base), the successfully matched field is added to the current set of candidate query fields. This process ensures that professional terms in the business scenario can be preferentially mapped to the system's predefined business fields.

[0085] S340. If a match is not found with the business knowledge base, the current key information to be retrieved will be matched with the pre-built general knowledge base. If a second query field is successfully matched, the second query field will be added to the current set of candidate query fields.

[0086] The general knowledge base can be understood as containing standard log field definitions that are common across industries, providing basic terminology-to-field mapping rules. This type of knowledge base is suitable for default matching scenarios when business-specific fields are missing, ensuring that all basic query elements receive field mapping support. The second query field can be understood as a standard field matched from the general knowledge base, reflecting the basic mapping rules that are common across industries. This type of field serves as a supplementary solution when business fields are missing, ensuring that all basic query elements are supported.

[0087] In this embodiment, when the business knowledge base fails to match, the currently retrieved key information is automatically redirected to the general knowledge base for matching. If a corresponding second query field is successfully matched in the general knowledge base (e.g., mapping "error" to the "error_level=ERROR" standard field in the general knowledge base), the general field is added to the current set of candidate query fields as a supplementary solution when the business knowledge base fails to match.

[0088] S350. In the operation and maintenance rule set, obtain the target operation and maintenance rule that the current retrieval key information hits, and according to the target operation and maintenance rule, obtain at least one target historical retrieval data in the historical retrieval dataset;

[0089] Each historical search data entry includes: a historical search question, a historical search analysis statement constructed based on the historical search question, and a set of historical search logs that match the historical search analysis statement.

[0090] In this embodiment, rule matching is performed on the preset operation and maintenance rule set based on the current key search information. After identifying the applicable target operation and maintenance rule, relevant records are extracted from the historical search dataset based on the rule. Each historical record contains three complete elements: the original historical search question text, the search analysis statement generated accordingly, and the historical log result set obtained by executing the statement. These structured historical data provide supplementary references for the current query.

[0091] S360. Match the current key information with each historical search analysis statement in the target historical search data, and add the third query field to the current candidate query field set when a third query field is successfully matched.

[0092] The third field can be understood as an experience field mined from historical retrieval data. It comes from verified fields in past successful query statements. These fields carry actual operation and maintenance experience and have scenario adaptability and practical verification.

[0093] In this embodiment, the current key information is matched with the search analysis statement in the target historical search data obtained in the previous step. When a third query field that is semantically related to the key information is found in the historical statement, the historically verified field is added as a candidate to the current set of candidate query fields.

[0094] S370. Check whether the processing of all key information has been completed: if yes, proceed to S380; otherwise, return to S320.

[0095] In this embodiment, the process of retrieving key information is retried through a loop control mechanism, returning to execute the operation of "sequentially obtaining the next key information as the current processing object" until all key information extracted from natural language queries has completed the processing flow of business knowledge base matching, general knowledge base matching, and historical data matching, ensuring that each key information generates a corresponding set of alternative query fields before entering the subsequent large language model processing stage.

[0096] S380. Provide the set of each alternative query field to the large language model.

[0097] S390. After filtering the set of candidate query fields through the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field.

[0098] S3100. Perform a search in at least one log database based on the log retrieval and analysis statement, and provide the target log set obtained from the search to the user.

[0099] The technical solution of this invention responds to a user's natural language log retrieval question and extracts at least one key retrieval information. It sequentially obtains one key retrieval information as the current key retrieval information and initializes a current candidate query field set. The current key retrieval information is matched against a business knowledge base, and if successful, a first query field is added to the current set. If the business knowledge base match fails, a match is made against a general knowledge base, and if successful, a second query field is added. Based on the operation and maintenance rule set, the target operation and maintenance rule and its corresponding target historical retrieval data are obtained. The current key retrieval information is matched against historical retrieval analysis statements, and if successful, a third query field is added. The process involves returning the processed data until all key information is retrieved. The set of candidate query fields is then provided to the large language model. The large language model filters and determines the target query field, generating a log retrieval and analysis statement. Finally, this statement is used to retrieve data from the log database, and the target log set is returned to the user. This novel log retrieval method significantly improves the accuracy of converting natural language to professional queries through multi-level knowledge base collaborative matching. It employs a multi-level matching mechanism that prioritizes business knowledge bases, supplements with general knowledge bases, and verifies historical data. This ensures that field mapping conforms to business characteristics and has been tested in practice, effectively reducing manual configuration costs while improving the reliability and processing efficiency of query statements.

[0100] Optionally, based on the above embodiments, providing each set of alternative query fields to the large language model may include:

[0101] A second prompt word template is constructed based on the business knowledge base, the general knowledge base, and at least one target historical retrieval data obtained from the historical retrieval dataset;

[0102] The set of alternative query fields corresponding to each key retrieval information, along with the second prompt word template, are provided to the large language model.

[0103] Generally, a structured suggestion framework is constructed based on the mapping of proprietary fields in the business knowledge base, the standard terminology comparison in the general knowledge base, and the effective query patterns in historical search records. This framework clearly marks the source of each field and its applicable scenarios, providing standardized guidance for subsequent processing.

[0104] Generally, this framework intelligently combines the candidate field sets corresponding to each key retrieval information to form a complete model input. This combination preserves the diversity of the original query elements while ensuring that the large language model can accurately understand the processing requirements through standardized templates, ultimately outputting query statements that conform to technical specifications.

[0105] Optionally, based on the above embodiments, after filtering the set of candidate query fields using a large language model to determine the target query fields corresponding to each key retrieval information, a log retrieval analysis statement for the natural language log retrieval problem is formed based on each target query field, which may include:

[0106] By using a large language model to filter the set of candidate query fields, the target query fields corresponding to each key piece of retrieval information are determined.

[0107] Based on the field type of each target query field, a local query statement corresponding to each target query field is constructed using a large language model.

[0108] By constructing templates based on preset retrieval and analysis statements using a large language model, the various local query statements are assembled to form log retrieval and analysis statements for natural language log retrieval problems.

[0109] Generally, large language models will evaluate the set of candidate fields from multiple dimensions, taking into account factors such as semantic relevance, business adaptability, and historical usage frequency, to select the most suitable target query field for each key piece of information to be retrieved.

[0110] Generally, based on field type characteristics (such as time fields, status fields, or statistical fields), the large language model will generate structured local query statement fragments according to the corresponding syntax rules, ensuring that each field can be accurately converted into executable query conditions.

[0111] Generally, by using predefined statement assembly templates, the large language model combines and arranges various local query statements according to logical relationships, ultimately generating a complete query statement that conforms to the syntax specifications of the target log database and fully expresses the user's search intent.

[0112] Example 4

[0113] Figure 4 This is a schematic diagram of a log retrieval device based on natural language, provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes:

[0114] Natural Language Processing Module 410 is used to extract at least one key search information from a natural language log retrieval question in response to the user's natural language log retrieval question.

[0115] The knowledge matching module 420 is used to match each key search information with knowledge in the knowledge base and historical environment data to obtain a set of candidate query fields corresponding to each key search information, and to provide each set of candidate query fields to the big language model.

[0116] The query statement generation module 430 is used to filter the set of candidate query fields through a large language model, determine the target query fields corresponding to each key information to be retrieved, and then form log retrieval analysis statements for natural language log retrieval problems based on each target query field.

[0117] The log retrieval module 440 is used to perform retrieval in at least one log database based on the log retrieval analysis statement, and to provide the user with the target log set obtained from the retrieval.

[0118] The technical solution of this invention first responds to a user's natural language log retrieval question by extracting at least one key retrieval information from the question. Then, each key retrieval information is matched with knowledge in a knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information. These candidate query field sets are then provided to a large language model. Next, the large language model filters these candidate query field sets to determine the target query fields corresponding to each key retrieval information. Based on these target query fields, a log retrieval analysis statement is formed for the natural language log retrieval question. Finally, the log retrieval analysis statement is used to perform a retrieval in at least one log database, and the retrieved target log set is provided back to the user. This novel log retrieval method significantly lowers the user threshold through natural language interaction, allowing complex log retrieval without requiring expertise in query syntax. By utilizing the semantic understanding capabilities of a large language model combined with a multi-level knowledge base matching mechanism, the business accuracy and technical feasibility of the query statement are ensured.

[0119] Based on the above embodiments, the natural language processing module 410 may further include:

[0120] The template matching submodule is used to obtain a set of natural language retrieval question templates when a natural language question is received from the user, and to construct the first prompt word template through the set of natural language retrieval question templates;

[0121] The range determination submodule is used to provide the natural language question and the first prompt word template to the large language model to obtain feedback results on whether the natural language question belongs to the answer range;

[0122] The information extraction submodule is used to generate a natural language log retrieval question based on the natural language question if it is determined that the natural language question falls within the scope of the answer, and to extract at least one key retrieval information from the natural language log retrieval question.

[0123] Based on the above embodiments, the knowledge matching module 420 is specifically used for:

[0124] Retrieve one key information item at a time as the current key information item, and initialize and construct a set of candidate query fields that match the current key information item.

[0125] The current key information to be retrieved is matched against a pre-built business knowledge base. If the first query field is successfully matched, the first query field is added to the current set of candidate query fields.

[0126] If a match is not found in the business knowledge base, the current key information to be retrieved will be matched in the pre-built general knowledge base, and if a second query field is successfully matched, the second query field will be added to the current set of candidate query fields.

[0127] In the set of operation and maintenance rules, obtain the target operation and maintenance rule that the current key information of the search hits, and according to the target operation and maintenance rule, obtain at least one target historical search data in the historical search dataset; wherein, each historical search data includes: historical search question, historical search analysis statement constructed based on historical search question, and historical search log set matching historical search analysis statement;

[0128] The current key search information is matched with each historical search analysis statement in the target historical search data, and when a third query field is successfully matched, the third query field is added to the current candidate query field set.

[0129] Return to the previous step and execute the operation of retrieving one key information item at a time as the current key information item, until all key information items have been processed.

[0130] Based on the above embodiments, the knowledge matching module 420 is specifically used for:

[0131] A second prompt word template is constructed based on the business knowledge base, the general knowledge base, and at least one target historical retrieval data obtained from the historical retrieval dataset;

[0132] The set of alternative query fields corresponding to each key retrieval information, along with the second prompt word template, are provided to the large language model.

[0133] Based on the above embodiments, the query statement generation module 430 is specifically used for:

[0134] By using a large language model to filter the set of candidate query fields, the target query fields corresponding to each key piece of retrieval information are determined.

[0135] Based on the field type of each target query field, a local query statement corresponding to each target query field is constructed using a large language model.

[0136] By constructing templates based on preset retrieval and analysis statements using a large language model, the various local query statements are assembled to form log retrieval and analysis statements for natural language log retrieval problems.

[0137] Based on the above embodiments, the information extraction submodule is specifically used for:

[0138] Detect whether there is any prior historical dialogue information associated with the natural language question;

[0139] If so, obtain the result summary information that matches the previous historical dialogue information; wherein, the result summary information includes the previous historical natural language log retrieval question, and the information summary of the historical log set retrieved based on the previous historical natural language log retrieval question;

[0140] The natural language query is rewritten based on the results summary information to obtain the natural language log retrieval query;

[0141] If not, directly identify the natural language question as a natural language log retrieval question.

[0142] Based on the above embodiments, it may further include: a summary generation module, wherein:

[0143] The summary generation module is used to process the natural language log retrieval question and the target log set after receiving user feedback, and to obtain the result summary information of the current round of dialogue information for use in the next round of dialogue.

[0144] The log retrieval device based on natural language provided in the embodiments of the present invention can execute the log retrieval method based on natural language provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0145] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0146] The information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0147] Provide users with a corresponding entry point to choose whether to agree to or reject the automated decision-making result; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0148] Example 5

[0149] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0150] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0151] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0152] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a natural language-based log retrieval method, namely:

[0153] In response to a user's natural language log retrieval question, extract at least one key retrieval information from the natural language log retrieval question;

[0154] Each key retrieval information is matched with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information, and each set of candidate query fields is provided to the large language model.

[0155] After filtering the set of candidate query fields by the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field.

[0156] The system performs a search in at least one log database based on the log retrieval and analysis statement, and provides the target log set obtained from the search to the user.

[0157] In some embodiments, a natural language-based log retrieval method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the natural language-based log retrieval method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a natural language-based log retrieval method by any other suitable means (e.g., by means of firmware).

[0158] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0159] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0160] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0162] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0163] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0164] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0165] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A log retrieval method based on natural language, characterized in that, include: In response to a user's natural language log retrieval question, extract at least one key retrieval information from the natural language log retrieval question; Each key retrieval information is matched with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key retrieval information, and each set of candidate query fields is provided to the large language model. After filtering the set of candidate query fields by the large language model and determining the target query fields corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field. The system performs a search in at least one log database based on the log retrieval and analysis statement, and provides the target log set obtained from the search to the user.

2. The method according to claim 1, characterized in that, In response to a user's natural language log retrieval question, at least one key retrieval information is extracted from the natural language log retrieval question, including: When a user inputs a natural language question, a set of natural language retrieval question templates is obtained, and a first prompt word template is constructed using the set of natural language retrieval question templates. The natural language question and the first prompt word template are provided together to the large language model to obtain feedback results on whether the natural language question falls within the scope of the answer; If it is determined that the natural language question falls within the scope of the answer, a natural language log retrieval question is generated based on the natural language question, and at least one key retrieval information is extracted from the natural language log retrieval question.

3. The method according to claim 1, characterized in that, Each key search term is matched with knowledge in the knowledge base and historical environmental data to obtain a set of candidate query fields corresponding to each key search term, including: Retrieve one key information item at a time as the current key information item, and initialize and construct a set of candidate query fields that match the current key information item. The current key information to be retrieved is matched against a pre-built business knowledge base. If the first query field is successfully matched, the first query field is added to the current set of candidate query fields. If a match is not found in the business knowledge base, the current key information to be retrieved will be matched in the pre-built general knowledge base, and if a second query field is successfully matched, the second query field will be added to the current set of candidate query fields. In the set of operation and maintenance rules, obtain the target operation and maintenance rule that the current retrieval key information hits, and according to the target operation and maintenance rule, obtain at least one target historical retrieval data in the historical retrieval dataset; Each historical retrieval data entry includes: a historical retrieval question, a historical retrieval analysis statement constructed based on the historical retrieval question, and a set of historical retrieval logs that match the historical retrieval analysis statement; The current key search information is matched with each historical search analysis statement in the target historical search data, and when a third query field is successfully matched, the third query field is added to the current candidate query field set. Return to the previous step and execute the operation of retrieving one key information item at a time as the current key information item, until all key information items have been processed.

4. The method according to claim 3, characterized in that, Provide the set of candidate query fields to the large language model, including: A second prompt word template is constructed based on the business knowledge base, the general knowledge base, and at least one target historical retrieval data obtained from the historical retrieval dataset; The set of alternative query fields corresponding to each key retrieval information, along with the second prompt word template, are provided to the large language model.

5. The method according to claim 4, characterized in that, After filtering the candidate query field set using a large language model to determine the target query field corresponding to each key retrieval information, log retrieval analysis statements for natural language log retrieval problems are formed based on each target query field, including: By using a large language model to filter the set of candidate query fields, the target query fields corresponding to each key piece of retrieval information are determined. Based on the field type of each target query field, a local query statement corresponding to each target query field is constructed using a large language model. By constructing templates based on preset retrieval and analysis statements using a large language model, the various local query statements are assembled to form log retrieval and analysis statements for natural language log retrieval problems.

6. The method according to claim 2, characterized in that, Generate natural language log retrieval questions based on natural language queries, specifically including: Detect whether there is any prior historical dialogue information associated with the natural language question; If so, obtain the result summary information that matches the previous historical dialogue information; wherein, the result summary information includes the previous historical natural language log retrieval question, and the information summary of the historical log set retrieved based on the previous historical natural language log retrieval question; The natural language query is rewritten based on the results summary information to obtain the natural language log retrieval query; If not, directly identify the natural language question as a natural language log retrieval question.

7. The method according to claim 6, characterized in that, After providing user feedback on the retrieved target log set, the process also includes: Information summarization is performed on the natural language log retrieval problem and the target log set to obtain a result summary of the current round of dialogue information for use in the next round of dialogue.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the natural language-based log retrieval method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the natural language-based log retrieval method according to any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the natural language-based log retrieval method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Processing method and device for searching of natural language by remote sensing data

    CN103092979A

  • Log query statement generation method and device, equipment and storage medium

    CN118152341A

  • Database query statement generation method and device, equipment and storage medium

    CN118227655A

  • Information retrieval method, apparatus, device and medium

    WO2020134684A1