Application log intelligent inspection method and system based on large model

Through the large model-based log intelligent inspection method, by establishing a mapping relationship between logs and target tags within a preset time window, generating dynamic vectors and embedding them into the large model, the problem of low intelligence level of existing log inspection methods is solved, and efficient and accurate log analysis is achieved.

CN120763004AActive Publication Date: 2025-10-10STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511271333.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-10
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing log inspection methods mainly rely on log indicator data, supplemented by log text analysis. They suffer from incomplete rule coverage, poor universality, and low intelligence, making them difficult to adapt to log analysis in various scenarios.

Method used

An intelligent inspection method for application logs based on a large model is adopted. By sampling the original logs within a preset time window, a mapping relationship with the target label is established, key logs are identified, and dynamic vectors are generated and embedded into the prompt words of the large model for analysis to obtain log inspection results.

Benefits of technology

It achieves preliminary automatic identification of abnormal logs without the need for frequent manual adjustment of keywords. It has high universality and significantly improves the efficiency and accuracy of log analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763004A_ABST
    Figure CN120763004A_ABST
Patent Text Reader

Abstract

The invention discloses an application log intelligent inspection method and system based on a large model, and relates to the technical field of big data analysis. Original logs in a preset time window are sampled; according to data of a preset target label, establishing a mapping relationship between the original log and the target label, and identifying a key concerned log; generating a dynamic vector according to the focused log and the context thereof, and embedding the dynamic vector into a cue word of a large model to obtain a dynamic large model; and analyzing the focused log and the context thereof by adopting the dynamic large model to obtain a log inspection result. By adopting the embodiment of the invention, the preliminary automatic identification of the abnormal log can be realized, the keyword does not need to be manually adjusted frequently in the subsequent identification process, the self-adaptive updating of the prompt word of the large model can be carried out, and the universality is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data analysis, and in particular to a method and system for intelligent inspection of application logs based on a large model. Background Art

[0002] With the development of IT technology, the number and types of servers, routers, switches and other devices in enterprises and organizations have increased. The amount of log data they generate is huge and continues to grow, which brings challenges to log management and analysis.

[0003] Existing log inspection methods primarily rely on log metric data, supplemented by log text analysis. Abnormal log analysis is often based on empirically configured keyword searches and template wildcard matching. These methods suffer from incomplete rule coverage, poor universality, and low intelligence, making them difficult to adapt to diverse log analysis scenarios. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide an intelligent inspection method and system for application logs based on a large model, which can realize the preliminary automatic identification of abnormal logs, and does not require frequent manual adjustment of keywords in the subsequent identification process. It can adaptively update the prompt words of the large model and has high universality.

[0005] An embodiment of the present invention provides a method for intelligent inspection of application logs based on a large model, including: Sample the original logs within the preset time window; According to the data of the preset target tag, a mapping relationship between the original log and the target tag is established to identify the logs of key concern; Generate a dynamic vector based on the focused log and its context, and embed the dynamic vector into the prompt word of the large model to obtain a dynamic large model; The dynamic large model is used to analyze the key logs and their contexts to obtain log inspection results.

[0006] As an improvement to the above solution, the original logs sampled within the preset time window include: According to the preset time interval, obtain the inspection configuration and start the inspection task; the inspection configuration includes the time range and inspection target; According to the inspection configuration, a list of log data that meets the inspection target within the time range is queried to obtain the original log.

[0007] As an improvement to the above solution, the mapping relationship between the application log and the target tag is established based on the data of the preset target tag, and the key focus log is identified, including: The data of the preset target label is converted into a target vector by using a vectorization service, and the application log is converted into a log vector; The similarity between the log vector and each target vector is calculated, and the target label corresponding to the target vector with the highest similarity to the log vector is taken as the label value of the log vector; According to the label value, a log of focus is identified.

[0008] As an improvement of the above scheme, the label value includes an unexpected label, an error label, and a message label, and the identification of the log of focus according to the label value includes: The original log with the label value of the unexpected label and the error label is marked as the log of focus; If there is no original log with the unexpected label or the error label in a preset time window, the original log at the end of the preset time window is marked as the log of focus.

[0009] As an improvement of the above scheme, the prompt word includes role definition, output target, input-output rule, and dynamic placeholder; wherein the dynamic placeholder is used to embed a dynamic vector to prompt the large model.

[0010] As an improvement of the above scheme, the generation of the dynamic vector according to the log of focus and its context, and the embedding of the dynamic vector into the prompt word of the large model to obtain a dynamic large model includes: The context of the log of focus is obtained, and a log of focus set is obtained in combination with the log of focus; The deployment application type and the associated database are obtained from the log of focus set, and a dynamic vector is generated according to the deployment application type and the associated database; The dynamic vector is embedded into the prompt word of the large model to obtain a dynamic large model.

[0011] As an improvement of the above scheme, the obtaining of the context of the log of focus and the obtaining of the log of focus set in combination with the log of focus includes: According to the time sequence, n upper and lower original logs of the log of focus are obtained as the context; n is a preset context quantity; The log of focus and the context are dynamically aggregated to remove duplicate logs to obtain a log of focus set.

[0012] As an improvement of the above scheme, the analysis of the log of focus and its context by using the dynamic large model to obtain a log inspection result includes: The log of focus and its context are sorted in time sequence to obtain a log sequence; Combined with the historical log inspection results, the dynamic large model is used to perform multiple rounds of reasoning on the log sequence until the confidence level is greater than a preset confidence threshold, thereby obtaining the output of the dynamic large model; The abnormality level, log summary and solution output by the dynamic large model are assembled into a preset structure to obtain the log inspection result.

[0013] As an improvement to the above solution, after analyzing the key logs and their contexts using the dynamic large model to obtain log inspection results, the large model-based application log intelligent inspection method further includes: If the abnormality level in the log inspection result is greater than a preset abnormality level threshold, the log inspection result is stored in a historical result database.

[0014] The embodiment of the present invention further provides an application log intelligent inspection system based on a large model, comprising: Log sampling module, used to sample original logs within a preset time window; A log identification module is used to establish a mapping relationship between the original log and the target tag based on the data of the preset target tag, and identify the logs of key concern; A vector embedding module, configured to generate a dynamic vector based on the focused log and its context, and embed the dynamic vector into the prompt word of the large model to obtain a dynamic large model; The log analysis module is used to analyze the original log using the dynamic large model to obtain log inspection results.

[0015] Compared to the prior art, the present invention discloses a large-scale model-based intelligent application log inspection method and system. The method samples raw logs within a preset time window; establishes a mapping relationship between the raw logs and target tags based on preset target tag data to identify key logs; generates dynamic vectors based on the key logs and their contexts, embeds the dynamic vectors into the prompt words of the large model to obtain a dynamic large model; and uses the dynamic large model to analyze the key logs and their contexts to obtain log inspection results. The embodiments of the present invention can achieve preliminary automatic identification of abnormal logs, eliminate the need for frequent manual adjustment of keywords during subsequent identification, and enable adaptive updates of the large model prompt words, thus having high universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flowchart of a method for intelligent inspection of application logs based on a large model provided by an embodiment of the present invention; Figure 2 This is a structural diagram of an application log intelligent inspection system based on a large model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] In the description of the specification and claims, it should be understood that the terms "first," "second," etc., are used solely for descriptive purposes to distinguish between identical technical features and are not to be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to, nor do they necessarily describe a sequential or chronological order. The terms are interchangeable where appropriate. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one of those features.

[0019] The embodiment of the present invention provides an intelligent inspection method for application logs based on a large model. Figure 1 In this embodiment, the large model-based application log intelligent inspection method is specifically performed through steps S1 to S4: S1. Sample the original logs within the preset time window.

[0020] The original logs in each preset time window are processed as a patrol task. It should be noted that the size of the preset time window can be flexibly adjusted according to the business function corresponding to the original log, and can dynamically adapt to business changes.

[0021] S2. According to the data of the preset target tag, a mapping relationship between the original log and the target tag is established to identify the logs of key concern.

[0022] It should be noted that the data for the preset target tag can be derived from historical data or simulated data that meets the target tag rules. The raw logs within the preset time window typically contain multiple log data entries, each corresponding to a different time or system state. Based on the target tag data, each raw log entry can be associated with the target tag, thereby filtering out anomalous data and focusing on logs.

[0023] S3. Generate a dynamic vector based on the key attention log and its context, and embed the dynamic vector into the prompt word of the large model to obtain a dynamic large model.

[0024] Exemplarily, the large model is an LLM (Large Language Model), which is capable of semantic understanding and reasoning. Existing techniques typically require large models to learn fixed keyword input or wildcard matching to meet reasoning requirements. However, due to the complexity and variability of log traffic, maintaining keywords during log analysis requires significant manpower.

[0025] In a preferred embodiment of the present invention, the required dynamic vectors are quickly embedded in the prompt words in the form of template variables, thereby realizing the automatic generation of intelligent prompt words, thereby achieving the unified encapsulation of log intelligent analysis business and large model service.

[0026] S4. Analyze the key logs and their contexts using the dynamic large model to obtain log inspection results.

[0027] It should be noted that in the embodiment of the present invention, since the key logs have been screened, the dynamic big model only focuses on the key logs and their context, which reduces the data processing volume of the big model and significantly improves the efficiency and accuracy of the big model in log analysis.

[0028] In this approach, sampling within a time window and mapping it to target labels allows for a focus on key logs. Dynamic vectors dynamically inject multi-dimensional information about the logs of interest within the current time window into the large model, enabling adaptive updates of the large model's prompt words. This embodiment is highly universal and significantly improves the efficiency and accuracy of log analysis using the large model compared to existing approaches that fix or simply adjust prompt words.

[0029] As a preferred embodiment, step S1, sampling original logs within a preset time window, includes: According to the preset time interval, obtain the inspection configuration and start the inspection task; the inspection configuration includes the time range and inspection target; According to the inspection configuration, a list of log data that meets the inspection target within the time range is queried to obtain the original log.

[0030] Considering that application logs come from multiple business sources, users need to pre-configure the log configuration information for each business application log. For example, the log configuration information includes the name of the cluster to which the log belongs, the name of the namespace to which the log belongs, the name of the deployment to which the log belongs, whether polling is enabled, the configuration creation time, and the configuration update time.

[0031] Based on whether polling is enabled and the configuration update time, a time range can be obtained as the preset time interval. For example, if polling is enabled for the application log of the first service and the polling period is 10 minutes, the preset time interval is 10 minutes, and the size of the preset time window is also 10 minutes. Furthermore, when the configuration is updated, the preset time window is also updated.

[0032] Preferably, the inspection target includes the target cluster, target space and target deployment name to be monitored. Exemplarily, based on the inspection target, a query request is sent to the ES log storage platform to filter out a list of log data that meets the conditions and obtain the original log.

[0033] Furthermore, in this embodiment of the present invention, if polling is enabled, the Celery scheduled task scheduler is used to trigger inspection tasks at set intervals. Celery is a distributed task queue that is responsible for asynchronous task execution, ensuring that inspection tasks do not block other business logic. The triggered task is added to the task queue and begins executing the specific inspection operation.

[0034] As a preferred embodiment, step S2, establishing a mapping relationship between the application log and the target tag based on the preset target tag data, and identifying the key focus logs, includes: Use vectorization services to convert data with preset target labels into target vectors and application logs into log vectors. Calculating the similarity between the log vector and each of the target vectors, and taking the target label corresponding to the target vector with the highest similarity to the log vector as the label value of the log vector; Identify key logs based on the tag value.

[0035] In some preferred embodiments, a vectorization service is implemented based on the GET universal semantic vector. Specifically, the GTE vectorization service is called to convert data with a preset target tag into a target vector. The GTE vectorization service is then called to calculate each application log into a 768-dimensional floating-point log vector. The similarity principle within the vector space is then utilized to intelligently assign a label to each log using the KNN algorithm, allowing rapid location of all abnormal logs within the log window.

[0036] Specifically, we send batches of 100 logs to the GTE vector computing service to obtain vector embedding expressions for all current logs. Based on the principle of similarity in the vector space, we find the most similar label for each log vector and assign a value, thereby realizing intelligent label recognition capabilities based on the vector space.

[0037] The principle of semantic vector similarity can be simply understood as follows: the cosine similarity between the words "king" and "queen" after being mapped into the vector space is much greater than the similarity between "king" and "benzene propylene." Based on this principle of similarity, we only need to know the content of the target label to achieve fast label matching for common text content, thereby realizing fast, universal, and intelligent identification of abnormal logs.

[0038] Furthermore, in some preferred embodiments, the tag values ​​include an unexpected tag, an error tag, and a message tag, and identifying the key logs according to the tag values ​​includes: Mark the original logs with unexpected and incorrect label values ​​as logs of particular concern; If there are no original logs with unexpected or incorrect labels within the preset time window, the original logs at the end of the preset time window will be marked as key focus logs.

[0039] For example, data with built-in tags such as Exception, Error, and Info can be input into the GTE vector computing service to obtain vector information for each tag. This service searches for all logs within a window with an intelligent tag value of Error or Exception, identifying them as key logs. Furthermore, given the possibility that there may be no abnormal logs within the entire time window, the latest logs at the end of the log queue are prioritized for analysis to ensure the most up-to-date analysis results.

[0040] As a preferred embodiment, the prompt word includes role definition, output target, input and output rules and dynamic placeholders; wherein the dynamic placeholders are used to embed dynamic vectors to prompt the large model.

[0041] It's important to note that each part of the prompt has a specific function and purpose. Role definitions help the large model clarify its core functionality; output targets are specific tasks that the large model needs to complete, such as log parsing, summary generation, and anomaly analysis. The core of input and output rules is to define the operational procedures and output formats that the model must follow. Through these rules, the large model can be constrained to a specific task scope, ensuring standardized and consistent output. Dynamic placeholders are used to dynamically add additional prompt content based on the focus of the log and its context.

[0042] In a specific embodiment, the content of the role definition part in the prompt word is: <text> You are an application log analysis assistant that can help users analyze the original log content entered by users <\TEXT> This allows the large model to clearly define its role as an "application log analysis assistant", whose core function is to process and analyze the original log content provided by users.

[0043] The output target part of the prompt word is: <text> The user input is the original log content. You need to be faithful to the original log content and analyze the log content step by step: Parse each entry in the JSON log Summarize and understand the basic information contained in each log Perform problem and exception analysis on all logs you receive Don't output your thought process, only return the final result Make sure you return the correct JSON format data without adding any explanatory text or special characters <\TEXT> The output goals provide the core tasks that the large model needs to complete from five dimensions: log parsing, summary generation, exception analysis, output requirements, and popularization.

[0044] The content of the input and output rules in the prompt word is: <text> The current system time is {{#Get current time / {x}text#}} The input data is in JSON format, data sample: { "projectName":"xxxx", "app_id":"xxxxx", "timestamp":"xx xx,xxxx@xx:xx:xx.xxx", "cluster_name":"xxxx", "kubernetes.namespace":"xx", "kubernetes.node.name":"xxxxxx", "kubernetes.pod.name":"xxxxxxxxxx", "kubernetes.container.name":"xxxxxxxxx", "container.id":"xxxxxxxxxxxx", "message":"xxxxxxxxxxxxxxxxxxxxx" } Provide a summary of the overall content of the user log, within 500 words Determine whether there is any abnormality in the log and give a conclusion title and a brief description The output data format is JSON format, output sample: { "ai_abstract":"xxxxxxxxx", "ai_level": "Normal / Warning / Severe", "ai_title":"xxxxxxx", "ai_description":"xxxxxxxxxxxxxxxx", "ai_suggestion":"xxxxxxxxxxxxxxxx" } <\TEXT> The input rule requirements ensure the timeliness of log analysis by dynamically obtaining the current system time. It further stipulates that the input must be in JSON format and provides data samples containing fields such as project name (projectName), application ID (app_id), timestamp (timestamp), cluster name (cluster_name), Kubernetes information (kubernetes.namespace), and log content.

[0045] The output rule requirements require a summary of the overall content of the large model log (ai_abstract), a judgment of the log abnormality level (ai_level), a title (ai_title) and a brief description (ai_description), and the final output format must be JSON.

[0046] The content of the dynamic placeholder part in the prompt word is: <text> {{#Start / {x}extra_prompt#}} <\TEXT> In specific calls, model developers can insert additional prompts here based on actual needs. For example, they can further optimize output requirements; add analysis rules customized in the configuration file; and pull system metadata from the system ledger, including the deployed application type (tomcat, nginx, springboot) and the associated database type (Redis, MySQL).

[0047] Furthermore, as a preferred embodiment, step S3, generating a dynamic vector based on the focused log and its context, and embedding the dynamic vector into the prompt word of the large model to obtain a dynamic large model, includes: Obtaining the context of the key attention log, and combining the key attention log to obtain a key attention log set; Acquire a deployment application type and an associated database from the key log set, and generate a dynamic vector according to the deployment application type and the associated database; The dynamic vector is embedded into the prompt word of the large model to obtain a dynamic large model.

[0048] It should be noted that if only the log text of a single exception is analyzed, the exception type may be misjudged. For example, the focus of the log may be on connection timeout, but the manifestations of specific timeouts such as database connection timeout, API call timeout, network timeout, etc. are all different, and different solutions need to be selected. Therefore, by obtaining the context, the scenario where the exception occurred can be clarified, and the complete exception propagation chain can be captured to trace the predecessor exception.

[0049] For example, the application type is determined by the service identifier or container tag in the log; the database connection string is extracted from the log, and the SQL statement features are analyzed to identify the database type and operation object. When generating a dynamic vector, the semantic features of the log collection are focused on based on text extraction; the application type is encoded; database features are constructed based on the database type; temporal features are calculated based on the log timestamp; topological features are calculated based on service dependencies; and an attention mechanism is used to fuse the semantic features, the application type encoding, database features, temporal features, and topological features to generate the dynamic vector.

[0050] Furthermore, preferably, the acquiring of the context of the key focus log and combining the key focus log to obtain a key focus log set includes: Obtain n original logs above and below the key log in chronological order as context; n is the preset number of contexts; Dynamically aggregate the key attention logs and the context, remove duplicate logs, and obtain a key attention log set.

[0051] By dynamically aggregating the focused logs and the context, overlapping logs can be removed. In some embodiments, the number of context logs obtained is updated, and context logs are obtained again to ensure that the preset number of context logs is met.

[0052] As a preferred embodiment, step S4, using the dynamic large model to analyze the key logs and their contexts to obtain log inspection results, includes: Sorting the key logs and their contexts in chronological order to obtain a log sequence; Combined with the historical log inspection results, the dynamic large model is used to perform multiple rounds of reasoning on the log sequence until the confidence level is greater than a preset confidence threshold, thereby obtaining the output of the dynamic large model; The abnormality level, log summary and solution output by the dynamic large model are assembled into a preset structure to obtain the log inspection result.

[0053] In an embodiment of the present invention, the dynamic big model analyzes one or more JSON-formatted log data, each of which contains the project name, application ID, timestamp, cluster name, Kubernetes information (node, Pod, container), and specific log content (message).

[0054] The dynamic big model first parses the input JSON data, extracting key metadata such as project name, application ID, and deployed application category. It then comprehensively analyzes all log content to determine if there are any anomalies or issues. Based on the anomaly level of the log content, the big model generates corresponding conclusions.

[0055] For example, in a specific embodiment, the JSON log input to the dynamic large model is: <json> { "projectName":"MyProject", "app_id":"123456", "timestamp":"10 Oct, 2023 @ 15:30:45.678", "cluster_name":"ClusterA", "kubernetes.namespace":"default", "kubernetes.node.name":"node-01", "kubernetes.pod.name":"pod-001", "kubernetes.container.name":"container-001", "container.id":"abc123def456", "message":"Error: Server is down" } < / json> The dynamic large model performs multiple rounds of reasoning based on prompt words embedded with dynamic vectors, and obtains the following log inspection results: <json> { "ai_abstract": "MyProject's application log shows that a problem occurred in container-001 of pod pod-001 on node node-01 in ClusterA.", "ai_level":"Severe", "ai_title":"Server failure", "ai_description":"The log shows that the server has crashed.", "ai_suggestion":"It is recommended to check the server status immediately and restart the service." } < / json> It's important to note that in this embodiment of the present invention, the dynamic large model not only supports embedding dynamic vectors in prompt words, enabling integration with user-defined rules and system records (current data), but also supports dynamic specification of specific models. For example, the deepseek-r1 model is used by default. However, if the log contains obvious errors, the Qwen2.5-Coder model is used to speed up data processing.

[0056] As a preferred implementation, the application log intelligent inspection method based on a large model further comprises the step S5: if the abnormal level in the log inspection result is greater than a preset abnormal level threshold, storing the log inspection result into a historical result database.

[0057] By storing the analyzed log inspection result into the historical result database, subsequent query, screening and analysis can be facilitated. Exemplarily, the stored content includes the abnormal level, the result title, the log summary and the solution suggestion in the log inspection result, as well as the inspection time, the log window time, the cluster name to which the log belongs, the space to which the log belongs and the deployment name to which the log belongs.

[0058] By using the application log intelligent inspection method based on a large model provided by the embodiment of the present application, the mapping relationship with the target label is established through time window sampling, and the key log can be focused on; the dynamic vector can dynamically inject the multi-dimensional information of the key log under the current time window into the large model, the adaptive update of the large model prompt word can be realized, the method has high universality, and compared with the existing fixed or simple adjustment prompt word mode, the efficiency and accuracy of the log analysis of the large model can be significantly improved.

[0059] The embodiment of the present application provides an application log intelligent inspection system based on a large model. Please refer to Figure 2 The application log intelligent inspection system based on a large model comprises a log sampling module 11, a log identification module 12, a vector embedding module 13 and a log analysis module 14, wherein: The log sampling module 11 is configured to sample the original log in a preset time window. The log identification module 12 is configured to establish a mapping relationship between the original log and the target label according to the data of the preset target label, and identify the key log. The vector embedding module 13 is configured to generate a dynamic vector according to the key log and its context, embed the dynamic vector into the prompt word of the large model, and obtain a dynamic large model. The log analysis module 14 is configured to analyze the original log by using the dynamic large model, and obtain a log inspection result.

[0060] As a preferred implementation, the log sampling module 11 is specifically configured to: According to a preset time interval, obtain an inspection configuration, and start an inspection task; the inspection configuration comprises a time range and an inspection target. According to the inspection configuration, query a log data list that meets the inspection target in the time range, and obtain an original log.

[0061] As a preferred implementation, the log identification module 12 comprises: The vectorization unit is used to convert data with preset target labels into target vectors using the vectorization service, and to convert application logs into log vectors. a similarity calculation unit, configured to calculate the similarity between the log vector and each of the target vectors, and use the target label corresponding to the target vector with the highest similarity to the log vector as the label value of the log vector; The log identification unit is used to identify the key logs according to the tag value.

[0062] Furthermore, preferably, the tag value includes an unexpected tag, an error tag, and a message tag, and the log identification unit is specifically configured to: Mark the original logs with unexpected and incorrect label values ​​as logs of particular concern; If there are no original logs with unexpected or incorrect labels within the preset time window, the original logs at the end of the preset time window will be marked as key focus logs.

[0063] As a preferred embodiment, the prompt word includes role definition, output target, input and output rules and dynamic placeholders; wherein the dynamic placeholders are used to embed dynamic vectors to prompt the large model.

[0064] As a preferred embodiment, the vector embedding module 13 includes: A context acquisition unit, configured to acquire the context of the key attention log and combine the key attention log to obtain a key attention log set; a dynamic vector generating unit, configured to obtain a deployment application type and an associated database from the key log set, and generate a dynamic vector according to the deployment application type and the associated database; The dynamic vector embedding unit is used to embed the dynamic vector into the prompt word of the large model to obtain the dynamic large model.

[0065] Furthermore, preferably, the context acquisition unit is specifically configured to: Obtain n original logs above and below the key log in chronological order as context; n is the preset number of contexts; Dynamically aggregate the key attention logs and the context, remove duplicate logs, and obtain a key attention log set.

[0066] As a preferred implementation, the log analysis module 14 is specifically configured to: Sorting the key logs and their contexts in chronological order to obtain a log sequence; Combined with the historical log inspection results, the dynamic large model is used to perform multiple rounds of reasoning on the log sequence until the confidence level is greater than a preset confidence threshold, thereby obtaining the output of the dynamic large model; The abnormality level, log summary and solution output by the dynamic large model are assembled into a preset structure to obtain the log inspection result.

[0067] As a preferred embodiment, the large model-based application log intelligent inspection system also includes a result storage module, which is used to: if the abnormality level in the log inspection result is greater than the preset abnormality level threshold, the log inspection result is stored in the historical result database.

[0068] An intelligent inspection system for application logs based on a large model provided by an embodiment of the present invention can focus on key logs by sampling in time windows and establishing a mapping relationship with target tags. Dynamic vectors can dynamically inject multi-dimensional information of the logs that are of focus in the current time window into the large model, enabling adaptive updating of the large model's prompt words. This system has high universality and can significantly improve the efficiency and accuracy of log analysis by the large model compared to existing methods of fixing or simply adjusting prompt words.

[0069] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0070] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.< / text> < / text> < / text> < / text>

Claims

1. A method for intelligent inspection of application logs based on a large model, characterized in that: include: Sample the original logs within the preset time window; According to the data of the preset target tag, a mapping relationship between the original log and the target tag is established to identify the logs of key concern; Generate a dynamic vector based on the focused log and its context, and embed the dynamic vector into the prompt word of the large model to obtain a dynamic large model; The dynamic large model is used to analyze the key logs and their contexts to obtain log inspection results.

2. The method for intelligent inspection of application logs based on a large model according to claim 1, characterized in that: The original logs within the preset sampling time window include: According to the preset time interval, obtain the inspection configuration and start the inspection task; the inspection configuration includes the time range and inspection target; According to the inspection configuration, a list of log data that meets the inspection target within the time range is queried to obtain the original log.

3. The method for intelligent inspection of application logs based on a large model according to claim 1, characterized in that: The step of establishing a mapping relationship between the application log and the target tag based on the preset target tag data and identifying the key logs includes: Use vectorization services to convert data with preset target labels into target vectors and application logs into log vectors. Calculating the similarity between the log vector and each of the target vectors, and taking the target label corresponding to the target vector with the highest similarity to the log vector as the label value of the log vector; Identify key logs based on the tag value.

4. The method for intelligent inspection of application logs based on a large model according to claim 3, characterized in that: The tag values ​​include an unexpected tag, an error tag, and a message tag, and identifying the logs to focus on based on the tag values ​​includes: Mark the original logs with unexpected and incorrect label values ​​as logs of particular concern; If there are no original logs with unexpected or incorrect labels within the preset time window, the original logs at the end of the preset time window will be marked as key focus logs.

5. The method for intelligent inspection of application logs based on a large model according to claim 1, characterized in that: The prompt words include role definition, output target, input and output rules and dynamic placeholders; wherein the dynamic placeholders are used to embed dynamic vectors to prompt the large model.

6. The method for intelligent inspection of application logs based on a large model according to claim 1 or 5, characterized in that: Generating a dynamic vector based on the focused log and its context, and embedding the dynamic vector into the prompt word of the large model to obtain a dynamic large model includes: Obtaining the context of the key attention log, and combining the key attention log to obtain a key attention log set; Acquire a deployment application type and an associated database from the key log set, and generate a dynamic vector according to the deployment application type and the associated database; The dynamic vector is embedded into the prompt word of the large model to obtain a dynamic large model.

7. The method for intelligent inspection of application logs based on a large model according to claim 6, characterized in that: The acquiring of the context of the key focus log and combining the key focus log to obtain a key focus log set includes: Obtain n original logs above and below the key log in chronological order as context; n is the preset number of contexts; Dynamically aggregate the key attention logs and the context, remove duplicate logs, and obtain a key attention log set.

8. The method for intelligent inspection of application logs based on a large model according to claim 1, characterized in that: The dynamic large model is used to analyze the key logs and their contexts to obtain log inspection results, including: Sorting the key logs and their contexts in chronological order to obtain a log sequence; Combined with the historical log inspection results, the dynamic large model is used to perform multiple rounds of reasoning on the log sequence until the confidence level is greater than a preset confidence threshold, thereby obtaining the output of the dynamic large model; The abnormality level, log summary and solution output by the dynamic large model are assembled into a preset structure to obtain the log inspection result.

9. The method for intelligent inspection of application logs based on a large model according to claim 1, characterized in that: After analyzing the key logs and their contexts using the dynamic large model to obtain log inspection results, the large model-based application log intelligent inspection method further includes: If the abnormality level in the log inspection result is greater than a preset abnormality level threshold, the log inspection result is stored in a historical result database.

10. An intelligent inspection system for application logs based on a large model, characterized in that: include: Log sampling module, used to sample original logs within a preset time window; A log identification module is used to establish a mapping relationship between the original log and the target tag based on the data of the preset target tag, and identify the logs of key concern; A vector embedding module, configured to generate a dynamic vector based on the focused log and its context, and embed the dynamic vector into the prompt word of the large model to obtain a dynamic large model; The log analysis module is used to analyze the original log using the dynamic large model to obtain log inspection results.

Citation Information

Patent Citations

  • Multi-feature log anomaly detection method and system based on log full semantics

    CN114610515A

  • Semi-supervised log anomaly detection method based on SBERT model

    CN117707813A

  • Log processing method and device

    CN120336110A

  • Log anomaly prediction method based on prompt enhancement

    CN120386686A

  • Abnormal behavior detection method and device based on large model and log big data

    CN120469894A