Log data interactive retrieval visualization and report generation system based on MCP

By adopting a five-module closed-loop architecture based on MCP, the problems of weak semantic understanding and reliance on manual operation for report writing in traditional log analysis solutions are solved. It realizes automated parsing, visualization and report generation of log data, and improves the intelligence and consistency of the system.

CN120874797APending Publication Date: 2025-10-31SHANGHAI NETIS TECH CO LTD

Patent Information

Application Number
CN202510996667.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional log data analysis solutions suffer from weak semantic understanding capabilities, limited presentation of analysis results, reliance on manual reporting which is time-consuming and inconsistent, and fail to meet the needs of multi-role operation and compliance.

Method used

It adopts a five-module closed-loop architecture based on MCP, including a semantic analysis engine, an intelligent visualization engine, a structured report generation engine, an intelligent archiving engine, and an adaptive optimization system. Through context-semantic driven data flow and feedback channels, it automates the entire process from raw log parsing to report generation.

Benefits of technology

It improves the structure, interpretability, and adaptability of log analysis, supports complex scenarios such as fault diagnosis, security auditing, and system situation awareness, and achieves semantically enhanced structured report generation and adaptive optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874797A_ABST
    Figure CN120874797A_ABST
Patent Text Reader

Abstract

The invention discloses an MCP-based log data interactive retrieval visualization and report generation system, and relates to the technical field of log visualization. The system comprises a semantic analysis engine, an intelligent visualization engine, a structured report generation engine, an intelligent filing engine and an adaptive optimization system. Each module constructs a closed loop through a data stream driven by context semantics and a feedback path, and supports the whole process from original log analysis, visual analysis and semantic report generation to archiving management and strategy iteration. The system has a context-enhanced semantic modeling capability, an interactive multi-view linkage mechanism, an automatic report arrangement structure and a feedback-driven optimization capability, improves the structural property, interpretability and adaptability of log analysis compared with a traditional scheme, and is particularly suitable for complex scenes such as troubleshooting, security audit and system situation awareness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of log visualization technology, and in particular relates to an interactive log data retrieval visualization and report generation system based on MCP. Background Technology

[0002] Currently, log data analysis plays a core role in enterprise system operation and maintenance, security auditing, and situational awareness. However, traditional solutions generally face three major problems: First, the semantic understanding capability is weak, relying solely on keyword matching and static rules; second, the presentation of analysis results is monotonous, often presented as lists or fixed charts, lacking interactivity and contextual support; and third, report writing heavily relies on manual operation, which is time-consuming and inconsistent.

[0003] The Model Context Protocol (MCP), as a semantic-driven data context modeling framework, provides application-layer, behavioral-layer, and system-layer context labels for log data, significantly enhancing its structuring and interpretability. This protocol recommends a hierarchical label architecture. Application-layer labels include fields such as application name, module ID, service type, protocol type, and interface version. Behavioral-layer labels cover attributes such as operation type (e.g., login, query, delete), behavior result (success / failure), parameter characteristics, status code, and response time. System-layer labels contain contextual information such as device ID, network protocol, resource usage, session identifier, user role, and permission level. Preferably, the label generation rules are based on a combination of regular expression matching, machine learning classification models (e.g., LSTM, BERT), and a rule engine. A context window mechanism is used to implement the association logic between labels, supporting cross-module semantic data transfer and consistency maintenance. By introducing MCP, the system can conduct multi-dimensional retrieval analysis based on structured semantics and drive visualization and automatic report generation processes, effectively overcoming the shortcomings of traditional solutions in semantic understanding, interactivity, and automation. Furthermore, with the increasing demands for multi-role operation and compliance, the need for capabilities such as visual interaction, context aggregation, and report version control is becoming increasingly prominent. The context fusion mechanism provided by MCP provides a solid foundation for building an intelligent, semantic-driven log analysis system. Summary of the Invention

[0004] This invention provides an interactive log data retrieval, visualization, and report generation system based on MCP (Multi-Channel Programming). Each module constructs a closed loop through context-semantic driven data flow and feedback pathways, supporting the entire process from raw log parsing, visual analysis, semantic report generation to archiving management and policy iteration. The system possesses context-enhanced semantic modeling capabilities, an interactive multi-view linkage mechanism, an automatic report arrangement structure, and feedback-driven optimization capabilities. Compared to traditional solutions, it improves the structure, interpretability, and adaptability of log analysis, making it particularly suitable for complex scenarios such as fault diagnosis, security auditing, and system situational awareness. In summary, it solves the problems mentioned in the background technology.

[0005] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:

[0006] The MCP-based interactive log data retrieval, visualization, and report generation system of the present invention includes a semantic analysis engine (module A), an intelligent visualization engine (module B), a structured report generation engine (module C), an intelligent archiving engine (module D), and an adaptive optimization system (module E);

[0007] Module A serves as the system entry point, receiving raw log data. Its output is structured semantic logs, which are then used by module B.

[0008] Module A.1 is a log structure parser that takes raw logs in various formats as input and outputs structured fields.

[0009] Module A.2 is a semantic feature extractor. It takes structured fields as input and outputs behavioral labels, protocols, and context tags. This module recommends using a feature extraction method that combines rule engines and machine learning models. It extracts key fields (such as timestamps and IP addresses) from logs using regular expressions, extracts behavioral semantic features using word vector models (such as Word2Vec), and generates context tags based on a context window mechanism.

[0010] Preferably, the feature extraction process includes three steps: field parsing, semantic annotation, and context association. The specific process is as follows: 1. Parse the log text into key-value pairs; 2. Perform part-of-speech tagging and semantic classification on behavioral keywords such as "login failure"; 3. Associate with a predefined behavioral tag library (such as "security / login_failure").

[0011] Module A.3 is an information value evaluator that assesses the analytical value of each log segment and outputs weighted indicators. This module recommends calculating the value weight of log segments based on multi-dimensional evaluation indicators, including four dimensions: timeliness (more recent logs have higher weight), scope of influence (logs involving core business have higher weight), repetition frequency (logs appearing for the first time have higher weight), and degree of anomaly. Preferably, the weight calculation method uses a weighted average algorithm, with the weight calculation formula: Weight = α × Timeliness + β × Scope of Influence + γ × Repetition Frequency + δ × Degree of Anomaly, where α + β + γ + δ = 1. The weights of each dimension are adjusted through configuration parameters, the degree of anomaly is determined through statistical deviation analysis, and the scope of influence is calculated through entity correlation.

[0012] Module A.4 is a redundancy pattern detector that removes duplicate logs and outputs a set of semantic logs after redundancy removal.

[0013] Module B is responsible for the visualization and interactive analysis of semantic data. Its output consists of interaction records and semantic aggregation information, which are used by Module C to generate reports.

[0014] Module B.1 is a semantic awareness visualizer that takes a semantic log as input and generates event primitives and context primitives.

[0015] Module B.2 is a visual layout manager that takes a set of primitives and their weights as input and outputs a structured view layout scheme. This module recommends a hybrid layout strategy based on grid layout and force-directed algorithms. View priority is determined by weight values, and responsive layout algorithms automatically adjust view size and position. Preferably, the layout algorithm includes three steps: weight sorting, space allocation, and conflict detection. The layout method (such as force-directed graph or hierarchical layout) is automatically switched based on the complexity of entity relationships. The weight values ​​are calculated using the weight index output by module A.3.

[0016] Module B.3 is an interactive analysis interface that loads the layout structure, receives user clicks, selections, and filtering actions, and outputs interaction logs. This module is recommended to use an event-driven view linkage mechanism, achieving data synchronization between multiple views through a publish-subscribe pattern. When a user selects a time period in the timeline view, the entity relationship diagram and context aggregation panel automatically filter events within that time period.

[0017] Preferably, the linkage mechanism includes three stages: event distribution, data synchronization, and view refresh, supporting cross-view context association and interactive feedback.

[0018] Module B.3.1 is an event timeline view. The input is a timeline, and the output is a draggable and scalable time structure display.

[0019] Module B.3.2 is an entity relationship diagram. It takes entity interaction relationships as input and outputs a visual node diagram of the graph structure.

[0020] Module B.3.3 is a context aggregation panel that aggregates semantic events by protocol, behavior, and session, and outputs the aggregate structure.

[0021] Module B.3.4 is a user interaction-driven analysis interface that integrates user behavior-triggered re-retrieval. The input is user interaction events, and the output is semantic query parameters.

[0022] Module B.4 is a multi-role collaboration supporter that manages shared state and collaborative annotations among different users. Module C combines visual analysis results with user behavior to generate structured reports, which are then stored by Module D.

[0023] Module C.1 is the report template manager. It takes the task type and context label as input and outputs a matching template structure.

[0024] Module C.2 is a content extractor that takes visual interaction logs as input and outputs semantic paragraph materials. This module recommends using a context-based tagging-based content extraction strategy. It identifies key events and behavioral patterns through semantic matching algorithms and uses a template mapping mechanism to convert interaction records into structured paragraphs. Preferably, the extraction rules include three steps: event identification, behavior aggregation, and context association. Key events are determined based on user click counts and filter frequency, or relevant log fragments are extracted through keyword matching, supporting automatic extraction of multi-dimensional semantic materials.

[0025] Module C.3 is a semantic summary generator. It takes extracted paragraphs and context as input and outputs inductive statements and suggestions. This module recommends using a summary algorithm based on a natural language processing model. Through a context-label-driven content generation mechanism, semantic paragraphs are transformed into structured analytical conclusions. Preferably, the summary algorithm includes three stages: content analysis, pattern recognition, and language generation. It supports template-based methods (defining templates such as "[behavior] occurred at [time], affected [entity], and its state is [result]") or deep learning models (such as the GPT series). For example, inputting the context labels "login_failure" and "unauthorized" generates "An unauthorized login failure event occurred on 2025-06-18 at 10:30, involving user ID: user42". It also supports multi-dimensional semantic summaries based on the MCP protocol context.

[0026] Module C.4 is the report version controller, which receives generated reports and outputs structured reports with timestamps and version information.

[0027] Module C.5.1 is the input for semantic analysis results. It takes the semantic summary of module B as input and uses it to start the report generation process.

[0028] Module C.5.2 is for automatic content extraction, extracting contextual events and behaviors as the basis for paragraphs.

[0029] Module C.5.3 is the semantic summary generation module, which summarizes the extracted results to form analytical conclusions. This module recommends using a template-driven semantic summary method, mapping extracted content to structured conclusions through predefined analytical templates and utilizing contextual tags to enhance the accuracy and completeness of the summary. Preferably, the summary process includes three steps: content classification, template matching, and conclusion generation, supporting the automatic generation of multiple types of analytical conclusions.

[0030] Module C.5.4 is the report template management module, which assigns summary content to each template paragraph.

[0031] Module C.5.5 is the version control module, which adds metadata tags and historical numbers to reports.

[0032] Module C.5.6 is the output generation module, which exports structured reports to a document platform or external system.

[0033] Module D indexes the reports and semantic log archives for future tracing and auditing.

[0034] Module D.1 is a hierarchical storage manager that inputs reports and semantic logs and categorizes them into hot, warm, and cold storage areas.

[0035] Module D.2 is the archive strategy optimizer, which adjusts storage strategies based on task priorities.

[0036] Module D.3 is an index builder that constructs multi-dimensional query index structures.

[0037] Module D.4 is a data traceability support unit that integrates index and path information to support behavior chain tracing and cross-period re-retrieval.

[0038] Module E is responsible for evaluating feedback, updating strategies, and applying the optimization results to Modules B and C.

[0039] Module E.1 is the performance evaluator, which collects hit rate, usage frequency, and user preferences.

[0040] Module E.2 is a strategy adaptive adjuster that generates optimization suggestion strategies based on indicators. This module recommends using a weighted scoring algorithm to calculate optimization indicators. It calculates a comprehensive score by allocating weights to three dimensions: hit rate, usage frequency, and user satisfaction, and uses a threshold trigger mechanism to generate strategy adjustment suggestions.

[0041] Preferably, the weight allocation adopts a dynamic adjustment strategy, with a hit rate weight of 0.4, a usage frequency weight of 0.3, and a user satisfaction weight of 0.3. The hit rate calculation formula is: hit rate = (number of correct search results / total number of searches) × 100%, and usage frequency = the number of times a certain function is called within a unit of time. The strategy adjustment rules are determined based on the scoring threshold and historical trend analysis. When the hit rate is lower than the threshold (e.g., 70%), the indexing strategy optimization is automatically triggered.

[0042] Module E.3 is a user feedback integrator that outputs explicit and implicit user behavioral feedback in a structured manner.

[0043] Module E.4 is the knowledge base updater, integrating suggested strategies to update knowledge rules and configuration files, and distributing them to modules B and C. This module recommends using a version-controlled knowledge rule storage structure, achieving real-time updates to the knowledge base through an incremental update mechanism, and distributing optimization strategies to target modules via a configuration distribution interface. Preferably, update triggering conditions include three methods: scoring threshold triggering, time period triggering, and manual triggering. Knowledge rules are stored in JSON format, using an ontology model or rule engine fact base as the storage structure, supporting rule version management and rollback mechanisms. The update process includes user feedback triggering rule priority adjustments, or automatically discovering new semantic association rules through machine learning models.

[0044] The present invention has the following advantages over the prior art:

[0045] (1) A five-module closed-loop log processing architecture based on the MCP protocol is proposed to realize an integrated system for log semantic modeling, visualization, report generation, archiving and policy optimization.

[0046] (2) Design a context-driven multi-view linkage visualization mechanism to support semantic synchronization and interactive feedback of event sequence diagrams, entity relationship diagrams and aggregate views.

[0047] (3) Construct a semantically enhanced structured report generation process, map interactive semantics and context labels to template blocks and automatically generate traceable report versions.

[0048] (4) A pluggable MCP tool registration and calling interface mechanism is proposed to support the decoupling between modules for context-structured query, visual control and report triggering.

[0049] (5) Implement a feedback evaluation and strategy update mechanism to drive the system to adaptively adjust the visualization and report generation strategy based on user behavior and analysis results.

[0050] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a diagram of the overall system architecture of the present invention;

[0053] Figure 2 This is a diagram of the internal structure of module A in this invention;

[0054] Figure 3 This is a flowchart of the workflow of module B in this invention;

[0055] Figure 4 This is a schematic diagram of the report generation mechanism in module C of the present invention;

[0056] Figure 5 This is a schematic diagram of the storage and archiving mechanism of module D in this invention;

[0057] Figure 6 This is a schematic diagram of the adaptive optimization mechanism of module E in this invention;

[0058] Figure 7 This is a schematic diagram of the multi-view linked search interface of the present invention;

[0059] Figure 8 A flowchart for generating the report of this invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] The problems to be solved by this invention are: (1) The original log structure is diverse and lacks semantic uniformity. The system completes semantic standardization through the context modeling mechanism of module A; (2) Static charts and tables cannot express semantic relationships. The system constructs a multi-dimensional semantic linkage visualization view through module B; (3) Reports rely on manual arrangement and lack consistency. The system realizes automatic structured report generation through the template mapping mechanism of module C; (4) Logs and reports are scattered and lack a tracking mechanism. The system realizes unified archiving and indexing support for semantic data and reports through module D; (5) The system lacks self-optimization capability. Module E establishes a behavior feedback closed loop to drive strategy optimization and dynamically update knowledge configuration.

[0062] like Figure 1-8 As shown, the MCP-based interactive log data retrieval, visualization, and report generation system of the present invention includes a semantic analysis engine (module A), an intelligent visualization engine (module B), a structured report generation engine (module C), an intelligent archiving engine (module D), and an adaptive optimization system (module E).

[0063] Module A ( Figure 1 This serves as the system entry point, receiving raw log data. Its output is a structured semantic log, which is used by module B.

[0064] Module A.1 ( Figure 2 This is a log structure parser that takes raw logs in various formats as input and outputs structured fields.

[0065] Module A.2 Figure 2 This module is a semantic feature extractor that takes structured fields as input and outputs behavioral labels, protocols, and context tags. It is recommended to use a feature extraction method combining a rule engine and a machine learning model. This involves extracting key fields (such as timestamps and IP addresses) from logs using regular expressions, extracting behavioral semantic features using word vector models (such as Word2Vec), and generating context tags based on a context window mechanism. Preferably, the feature extraction process includes three steps: field parsing, semantic annotation, and context association. The specific process is as follows: 1. Parse the log text into key-value pairs; 2. Perform part-of-speech tagging and semantic classification on behavioral keywords such as "login failed"; 3. Associate with a predefined behavioral tag library (such as "security / login_failure").

[0066] Module A.3 Figure 2 This module serves as an information value evaluator, assessing the analytical value of each log segment and outputting weighted indicators. It recommends calculating log segment value weights based on multi-dimensional evaluation indicators, including four dimensions: timeliness (more recent logs have higher weight), scope of influence (logs involving core business have higher weight), repetition frequency (logs appearing for the first time have higher weight), and anomaly degree. Preferably, the weight calculation method uses a weighted average algorithm, with the weight calculation formula: Weight = α × Timeliness + β × Scope of Influence + γ × Repetition Frequency + δ × Anomaly Degree, where α + β + γ + δ = 1. The weights of each dimension are adjusted through configuration parameters, the anomaly degree is determined through statistical deviation analysis, and the scope of influence is calculated through entity correlation.

[0067] Module A.4 Figure 2 () is a redundancy pattern recognizer that removes duplicate logs and outputs a set of semantic logs after redundancy removal.

[0068] Module B ( Figure 1This module is responsible for the visualization and interactive analysis of semantic data. Its output consists of interactive records and semantic aggregation information, which are used to generate reports for module C.

[0069] Module B.1 Figure 3 () is a semantic-aware visualizer that takes a semantic log as input and generates event primitives and context primitives.

[0070] Module B.2 Figure 3 This module is a visual layout manager that takes a set of primitives and their weights as input and outputs a structured view layout scheme. It recommends a hybrid layout strategy based on grid layout and force-directed algorithms. View priority is determined by weight values, and responsive layout algorithms automatically adjust view size and position. Preferably, the layout algorithm includes three steps: weight sorting, space allocation, and conflict detection. The layout method (such as force-directed graph or hierarchical layout) is automatically switched based on the complexity of entity relationships. Weight values ​​are calculated using the weight index output by module A.3.

[0071] Module B.3 Figure 3 This module provides an interactive analysis interface, loading the layout structure, receiving user clicks, selections, and filtering actions, and outputting interaction logs. It is recommended to use an event-driven view linkage mechanism, achieving data synchronization between multiple views through a publish-subscribe pattern. When a user selects a time period in the timeline view, the entity relationship diagram and context aggregation panel automatically filter events within that time period. Preferably, the linkage mechanism includes three stages: event distribution, data synchronization, and view refresh, supporting cross-view contextual association and interactive feedback.

[0072] Module B.3.1 Figure 7 This is an event timeline view. The input is a timeline of events, and the output is a draggable and scalable time structure display.

[0073] Module B.3.2 Figure 7 This is an entity relationship diagram. Input the entity interaction relationships and output a visual node diagram of the graph structure.

[0074] Module B.3.3 ( Figure 7 This is a context aggregation panel that aggregates semantic events by protocol, behavior, and session, and outputs the aggregate structure.

[0075] Module B.3.4 Figure 7 This is a user interaction-driven analysis interface that integrates user behavior-triggered re-retrieval. The input is user interaction events, and the output is semantic query parameters.

[0076] Module B.4 Figure 3 It serves as a multi-role collaboration supporter, managing shared state and collaborative annotations among different users.

[0077] Module C ( Figure 1 The visualization analysis results are combined with user behavior to generate a structured report, which is stored in module D.

[0078] Module C.1 Figure 4 This is a report template manager. Input the task type and context label, and output the matching template structure.

[0079] Module C.2 Figure 4 This module acts as a content extractor, taking visual interaction logs as input and outputting semantic paragraph content. It recommends using a context-based tagging strategy for content extraction, identifying key events and behavioral patterns through semantic matching algorithms, and converting interaction records into structured paragraphs using template mapping. Preferably, the extraction rules include three steps: event identification, behavior aggregation, and context association. Key events are determined based on user click counts and filter frequency, or relevant log fragments are extracted through keyword matching, supporting automatic extraction of multi-dimensional semantic content.

[0080] Module C.3 ( Figure 4 This module is a semantic summary generator. It takes extracted paragraphs and context as input and outputs inductive statements and suggestions. It is recommended to use a summary algorithm based on a natural language processing model, which transforms semantic paragraphs into structured analysis conclusions through a context-label-driven content generation mechanism. Preferably, the summary algorithm includes three stages: content analysis, pattern recognition, and language generation. It supports template-based methods (defining templates such as "[behavior] occurred at [time], affected [entity], and its state is [result]") or deep learning models (such as the GPT series). For example, inputting the context labels "login_failure" and "unauthorized" generates "An unauthorized login failure event occurred on 2025-06-18 at 10:30, involving user ID: user42". It also supports multi-dimensional semantic summaries based on the MCP protocol context.

[0081] Module C.4 Figure 4 It acts as a report version controller, receiving generated reports and outputting structured reports with timestamps and version information.

[0082] Module C.5.1 Figure 8 The semantic analysis result is input, which is the semantic summary of input module B, used to start the report generation.

[0083] Module C.5.2 Figure 8 The content is automatically extracted, and contextual events and behaviors are used as the basis for paragraphs.

[0084] Module C.5.3 Figure 8This module is a semantic summary generation module that summarizes the extracted results to form analytical conclusions. It is recommended to use a template-driven semantic summary method, which maps extracted content to structured conclusions through predefined analytical templates and utilizes contextual tags to enhance the accuracy and completeness of the summary. Preferably, the summary process includes three steps: content classification, template matching, and conclusion generation, supporting the automatic generation of multiple types of analytical conclusions.

[0085] Module C.5.4 Figure 8 This is the report template management module, which assigns summary content to each template paragraph.

[0086] Module C.5.5 Figure 8 This is the version control module, which adds metadata tags and historical numbers to reports.

[0087] Module C.5.6 Figure 8 This is the output generation module, which exports structured reports to a document platform or external system.

[0088] Module D ( Figure 1 The report and semantic log archive index will be used for future attribution and auditing. Module D.1 Figure 5 It is a hierarchical storage manager that categorizes input reports and semantic logs into hot, warm, and cold storage areas.

[0089] Module D.2 Figure 5 () is an archive strategy optimizer that adjusts storage strategies based on task priorities.

[0090] Module D.3 Figure 5 () is an index builder that constructs multi-dimensional query index structures.

[0091] Module D.4 Figure 5 It serves as a data traceability support tool, integrating index and path information to support behavioral chain tracing and cross-period re-retrieval.

[0092] Module E ( Figure 1 This is responsible for evaluating feedback, updating strategies, and applying the optimization results to modules B and C.

[0093] Module E.1 ( Figure 6 It serves as an effectiveness evaluator, collecting hit rate, usage frequency, and user preferences.

[0094] Module E.2 Figure 6This module acts as a strategy adaptive adjuster, generating optimization suggestions based on metrics. It recommends using a weighted scoring algorithm to calculate optimization metrics, assigning weights to three dimensions: hit rate, usage frequency, and user satisfaction. A threshold-triggered mechanism is then used to generate strategy adjustment suggestions. Preferably, a dynamic adjustment strategy is employed, with hit rate weighted at 0.4, usage frequency at 0.3, and user satisfaction at 0.3. The hit rate formula is: Hit Rate = (Number of Correct Search Results / Total Searches) × 100%. Usage frequency = Number of times a function is called per unit time. Strategy adjustment rules are determined based on scoring thresholds and historical trend analysis. When the hit rate falls below a threshold (e.g., 70%), indexing strategy optimization is automatically triggered.

[0095] Module E.3 ( Figure 6 It is a user feedback integrator that structures and outputs explicit and implicit user behavioral feedback.

[0096] Module E.4 Figure 6 This module acts as a knowledge base updater, integrating suggested strategies to update knowledge rules and configuration files, and distributing the updates to modules B and C. This module is recommended to use a version-controlled knowledge rule storage structure, achieving real-time updates of the knowledge base through an incremental update mechanism, and distributing optimization strategies to target modules via a configuration distribution interface. Preferably, update triggering conditions include three methods: scoring threshold triggering, time period triggering, and manual triggering. Knowledge rules are stored in JSON format, using an ontology model or rule engine fact base as the storage structure, supporting rule version management and rollback mechanisms. The update process includes user feedback triggering rule priority adjustments, or automatically discovering new semantic association rules through machine learning models.

[0097] The typical usage and key technical implementation details of each module in this invention include structural examples and the MCP tool registration interface.

[0098] Module A receives raw logs from the network probe, such as PCAP or Syslog, and generates the following semantic context after structuring:

[0099]

[0100] This module provides standard semantic retrieval capabilities through the MCP tool interface search_contextual_logs, the interface structure of which is as follows:

[0101]

[0102]

[0103]

[0104] After module B loads the semantic logs mentioned above, the user configures the following aggregation strategy in the "Context Aggregation Panel":

[0105]

[0106] Module B provides the explore_visual_context interface for external platforms to use in conjunction with it.

[0107]

[0108]

[0109]

[0110] Module C automatically matches report templates to generate analysis reports. A template example is shown below:

[0111] Module C exposes the generate_contextual_report interface for external calls:

[0112]

[0113]

[0114] Module D receives the report archive and builds a semantic index, such as:

[0115]

[0116] Module E summarizes the feedback and calls the optimization interface submit_feedback_and_optimize, which is defined as follows:

[0117]

[0118]

[0119] Ultimately, the modules collaborated to complete a closed-loop process from semantic structure construction, visual linkage analysis, automatic report generation to policy evaluation and feedback, fully validating the practicality and intelligence of this system in semantically enhanced log processing.

[0120] Competitive technology analysis

[0121] This chapter aims to analyze publicly available technologies or products that are comparable to the target functions and key mechanisms of this invention, focusing on the differences in system capabilities and structural mechanisms in areas such as semantic log processing, log visualization, and automatic report generation.

[0122] Splunk US Patent US20130297653A1 Semantic Log Visualization System

[0123] Technical features: This patent proposes a method for converting logs into semantic graphs for analysis, focusing on the extraction and display of semantic connections between log entities.

[0124] Limitations: This technology does not form a complete report generation and policy feedback mechanism, and its structure cannot be extended to a closed-loop log processing system.

[0125] Advantages of this invention: This patent constructs a hierarchical context tagging system and semantic association mechanism based on the MCP protocol. Through the synergistic effect of application layer, behavior layer, and system layer tags, it achieves a five-module closed-loop processing capability from log parsing to archiving, effectively solving the technical bottlenecks of scattered semantic understanding and lack of context association in traditional solutions.

[0126] ELK Stack + Kibana visualization solution

[0127] Technical features: It indexes log data through Elasticsearch, uses Kibana for chart visualization, and supports some search and statistical functions.

[0128] Limitations: The charts are static and lack a semantic tag-driven linkage mechanism; reports need to rely on external plugins or be completed manually.

[0129] Advantages of this invention: This system supports semantic linkage and contextual backtracking, automatically generates structured reports, and can be associated with interaction records and version control.

[0130] US20190011635A1 Auto-Generating Incident Reports from Logs

[0131] Technical features: A method is proposed to automatically extract keywords and time series data from logs to generate event summary reports.

[0132] Limitations: Limited to keyword and time series logic, lacking the ability to integrate contextual structure and user behavior.

[0133] Advantages of this invention: This invention supports multi-layer context model-driven content extraction and multi-template generation mechanisms, enhancing the structural integrity and interpretability of reports.

[0134] Graylog Visual Analysis Platform

[0135] Technical features: Provides search functionality for user-defined fields and event alerts, and integrates a basic chart analysis interface.

[0136] Limitations: Semantic processing relies on regular expressions and field extraction, and there is no unified semantic view modeling capability.

[0137] Advantages of this invention: This system constructs an MCP semantic engine and a contextual structure view to achieve the fusion of semantic-driven search and graphics rendering.

[0138] While the aforementioned technologies each possess capabilities in log analysis, display, and event extraction, they generally lack a system-level closed-loop design. This invention, through a modular MCP protocol system architecture, constructs a complete process from semantic analysis to report output, possessing capabilities in semantic unification, view interaction, report automation, and strategy optimization. It exhibits significant advantages in system integration, scalability, and intelligence.

[0139] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A log data interactive retrieval, visualization, and report generation system based on MCP, characterized in that, include: Module A: MCP-driven log semantic analysis engine; As the system entry point, it receives raw log data; Its output is a structured semantic log, which is used by module B; module A includes a log structure parser, a semantic feature extractor, an information value evaluator, and a redundancy pattern recognizer; the log structure parser is used to input raw logs in various formats and output structured fields; the semantic feature extractor is used to input structured fields and output behavior labels, protocols, and context markers; The information value evaluator is used to evaluate the analytical value of each log segment and output a weight index; the redundancy pattern recognizer is used to remove duplicate logs and output a set of semantic logs after redundancy removal. Module B, the MCP intelligent visualization engine, is responsible for the visualization and interactive analysis of semantic data. Its output is interactive records and semantic aggregation information, which are used for report generation by Module C. Module B includes a semantic awareness visualizer, a visualization layout manager, an interactive analysis interface, and a multi-role collaboration supporter. The semantic awareness visualizer is used to input semantic logs and generate event primitives and context primitives. The visual layout manager is used to input the set of graphic elements and their weights, and output a structured view layout scheme; the interactive analysis interface is used to load the layout structure, receive user clicks, selections, and filtering behaviors, and output interaction logs. The multi-role collaboration supporter is used to manage shared state and collaborative annotations among different users; Module C, the MCP intelligent report generation engine, combines visual analysis results with user behavior to generate structured reports for storage by module D. Module C includes a report template manager, a content extractor, a semantic summary generator, a report version controller, and semantic input and structured report export. The report template manager is used to input task types and context labels and output matching template structures. The content extractor is used to input visual interactive logs and output semantic paragraph materials. The semantic summary generator is used to extract paragraphs and context from the input and output inductive statements and suggestions. The report version controller is used to receive generated reports and output structured reports with timestamps and version information; The semantic input and structured report export are used to export a structured report; Module D, MCP intelligent storage and archiving engine; Archive and index reports and semantic logs for future source tracing and auditing; Module E, the adaptive learning and optimization system, is responsible for evaluating feedback, updating strategies, and applying the optimization results to Modules B and C. Module E includes an effect evaluator, a strategy adaptive adjuster, a user feedback integrator, and a knowledge base updater. The effect evaluator is used to collect hit rate, usage frequency, and user preferences. The strategy adaptive adjuster is used to generate optimization suggestion strategies based on the indicators. The user feedback integrator is used to output structured user explicit and implicit behavioral feedback. The knowledge base updater is used to integrate suggestion strategies, update knowledge rules and configuration files, and distribute them to modules B and C.

2. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The semantic feature extractor adopts a feature extraction method based on a combination of rule engine and machine learning model. It extracts key fields including timestamps and IP addresses from logs through regular expressions, extracts behavioral semantic features using word vector model, and generates context tags based on context window mechanism.

3. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 2, characterized in that, The extraction process of the behavioral semantic features includes three steps: field parsing, semantic annotation, and context association. The specific process is as follows: (1) Parse the log text into key-value pairs; (2) Perform part-of-speech tagging and semantic classification on behavioral keywords such as "login failed"; (3) Associate with a predefined behavior tag library.

4. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The information value evaluator calculates the value weight of log fragments based on multi-dimensional evaluation indicators, including four dimensions: timeliness, scope of influence, number of repetitions, and degree of anomaly. Timeliness refers to the weight of recently occurring logs; scope of influence refers to the weight of logs involving core business; and number of repetitions refers to the weight of logs appearing for the first time. The weight calculation method adopts a weighted average algorithm, and the weight calculation formula is: weight = α × timeliness + β × scope of influence + γ × number of repetitions + δ × degree of anomaly, where α + β + γ + δ = 1. The weights of each dimension are adjusted through configuration parameters, the degree of anomaly is determined through statistical deviation analysis, and the scope of influence is calculated through entity correlation.

5. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The visual layout manager adopts a hybrid layout strategy based on grid layout and force-directed algorithm. It determines the view priority through weight values ​​and automatically adjusts the view size and position using responsive layout algorithm. The layout algorithm includes three steps: weight sorting, space allocation, and conflict detection. It automatically switches the layout mode according to the complexity of entity relationships. The weight values ​​are calculated through the weight index output by the information value evaluator.

6. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The interactive analysis interface adopts an event-driven view linkage mechanism, and realizes data synchronization between multiple views through a publish-subscribe model. When the user selects a time period in the timeline view, the entity relationship diagram and the context aggregation panel automatically filter the events within that time period. The linkage mechanism includes three stages: event distribution, data synchronization, and view refresh, and supports cross-view context association and interactive feedback.

7. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 6, characterized in that, The interactive analysis interface includes an event timeline view, an entity relationship diagram, a context aggregation panel, a user interaction-driven analysis interface, and a multi-role collaboration supporter; the event timeline view is used to input the behavior timeline and output a draggable and scalable time structure display. The entity relationship graph is used to input entity interaction relationships and output a visual node graph of the graph structure; the context aggregation panel is used to aggregate semantic events by protocol, behavior, and session and output an aggregation structure. The user interaction-driven analysis interface is used to integrate user behavior-triggered re-retrieval, with user interaction events as input and semantic query parameters as output.

8. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The semantic input and structured report export process includes the following steps: Module C.5.1 is the input for semantic analysis results. It takes the semantic summary of module B as input and uses it to start report generation. Module C.5.2 is for automatic content extraction, extracting contextual events and behaviors as the basis for paragraphs; Module C.5.3 is a semantic summary generation module that summarizes the extracted results and forms analytical conclusions. This module adopts a template-driven semantic summary method, which maps the extracted content to structured conclusions through predefined analytical templates and uses context labels to enhance the accuracy and completeness of the summary. The summary process includes three steps: content classification, template matching, and conclusion generation, and supports the automatic generation of multiple types of analytical conclusions. Module C.5.4 is the report template management module, which assigns summary content to each template paragraph; Module C.5.5 is the version control module, which adds metadata tags and historical numbers to reports; Module C.5.6 is the output generation module, which exports structured reports to a document platform or external system.

9. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The module includes a hierarchical storage manager, an archive strategy optimizer, an index builder, and a data traceability supporter. The hierarchical storage manager is used to input reports and semantic logs. The archive strategy optimizer is used to adjust storage strategies based on task priorities. The index builder is used to construct multi-dimensional query index structures. The data traceability supporter is used to integrate index and path information, supporting behavior chain tracing and cross-period re-retrieval.

10. The MCP-based interactive log data retrieval, visualization, and report generation system according to claim 1, characterized in that, The knowledge base updater adopts a version-controlled knowledge rule storage structure, realizes real-time updates of the knowledge base through an incremental update mechanism, and distributes optimization strategies to target modules using a configuration distribution interface. The update triggering conditions include three methods: scoring threshold triggering, time period triggering, and manual triggering. The knowledge rules are stored in JSON format, using an ontology model or rule engine fact base as the storage structure, and support rule version management and rollback mechanisms. The update process includes user feedback triggering rule priority adjustment, or automatically discovering new semantic association rules through machine learning models.

Citation Information

Patent Citations

  • Metadata storage management offloading for enterprise applications

    US20130297653A1

Cited By

  • Business-driven log intelligent hierarchical storage method and system

    CN121277906A

  • Standard document duplicate checking result visual display method and system

    CN122019768A