An industrial big data and AI intelligent agent-based network information system operation and maintenance management system
Patent Information
- Application Number
- CN202610973073.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本发明要解决的技术问题是:现有技术中存在真实故障异常,从而导致日志检索范围扩大、故障定位偏移和运维报告可信度降低的问题的缺点,为此我们提出一种基于工业大数据与AI智能体的网络信息系统运维管理系统
[0016]本发明的技术效果和优点:本发明中,通过多源异构数据感知网关、动态配置同步中枢和多模态索引模块的配合,将工业网络中的遥测指标、原始日志、采集侧物理链路状态以及配置侧端口物理状态统一映射到节点标识、端口标识和端口状态版本号下,使拓扑关系邻接图索引、时序时间窗索引和全文本倒排索引能够在同一端口状态基础上进行关联检索,避免现有运维系统仅依据告警时间或日志关键词进行孤立分析而导致的误关联;同时,系统在拓扑遍历过程中区分有效可达节点集合和物理反证节点集合,将物理上处于阻断、休眠、未协商成功、维护隔离或不可达状态的邻接节点作为反证对象保留,结合弹性时间跨度对有效可达节点生成正向时序突变数据、对物理反证节点生成反证时序越限数据,并进一步根据越限时间是否落入预设过渡时间范围以及多采集来源是否存在矛盾越限,将节点分流为真实异常节点、物理状态冲突节点和待复核节点,由此能够减少采集缓存上报、传感器漂移、状态同步滞后或维护隔离期间残留数据造成的伪异常误判;此外,系统分别生成正向联合主键、反证联合主键和复核联合主键,对故障日志、采集异常或漂移日志、端口刷新或配置变更日志进行分流检索,使输出的结构化运维报告不仅能够定位真实故障,还能够给出伪异常来源和状态版本复核依据;进一步地,系统将物理状态冲突标记、状态版本待复核标记、异常类型、采集来源标识和冲突发生时间回写至多模态索引模块,更新冲突可信度记录,并据此调整后续同一端口状态版本下相似时序越限的检索路径,从而降低重复伪异常进入真实故障检索路径的概率,提高工业网络运维检索的准确性、可解释性和持续处理效率。
Smart Images

Figure CN122802357A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network information system operation and maintenance management technology, and in particular to a network information system operation and maintenance management system based on industrial big data and AI intelligent agents. Background Technology
[0002] Industrial network information systems typically deploy devices such as industrial switches, industrial servers, edge gateways, PLC control nodes, data acquisition terminals, production management servers, and host computer monitoring platforms. To ensure the stable operation of industrial production processes, the operation and maintenance system usually needs to continuously collect equipment telemetry indicators, physical link status, raw logs, configuration change records, and maintenance operation records, and locate the fault location, cause of the fault, and related logs when abnormal alarms occur.
[0003] Current operation and maintenance management methods typically employ multi-source data fusion, inputting time-series indicators, network topology, log text, and alarm events into an analysis platform. This data is then jointly analyzed based on topological relationships, time-series anomalies, and log keywords. For example, when a node experiences a sudden traffic surge, increased packet loss rate, abnormal port error rate, or excessive device resource utilization, the system combines the topological relationships of adjacent nodes with log information from the same time period to perform fault retrieval and root cause determination. While this approach can improve fault location efficiency to some extent, it still carries the risk of misjudgment in industrial network scenarios.
[0004] Therefore, it is necessary to provide a new network information system operation and maintenance management system that enables topology status, timing indicators, and log records to be associated under the same port status version, and to handle the actual fault, physical status conflict, and status version pending verification separately when physical status and timing exceed the limit, thereby improving the accuracy of fault log retrieval and operation and maintenance report output. Summary of the Invention
[0005] The technical problem this invention aims to solve is that existing technologies suffer from real-world faults and anomalies, leading to an expanded scope of log retrieval, fault location deviation, and reduced credibility of maintenance reports. To address this, we propose a network information system maintenance management system based on industrial big data and AI intelligent agents.
[0006] To achieve the above objectives, this application adopts the following technical solution: a network information system operation and maintenance management system based on industrial big data and AI intelligent agents, including a sensing gateway, a synchronization hub, a multimodal index module, an intent intelligent agent, and a retrieval intelligent agent; the sensing gateway collects telemetry indicators, logs, and link status and calibrates timestamps; the synchronization hub updates the port status version number when the physical status of the configuration-side port changes and maintains a preset transition time range; the multimodal index module constructs a topology, time sequence, and inverted index that associates node identifiers, port identifiers, and port status version numbers, and maintains conflict confidence records; when responding to abnormal alarms or retrieval requests, the retrieval intelligent agent reads the topology index starting from a determined alarm source node, and takes the nodes reachable by the connection edges that are enabled, successfully negotiated, not dormant, not blocked, not isolated, and whose status update time exceeds the preset transition time range as the valid reachable set, and takes the adjacent nodes that are blocked, dormant, unsuccessfully negotiated, or underwent maintenance as the valid reachable set. Nodes excluded due to isolation or unreachability are designated as a set of physical counter-evidence. The intent agent generates an elastic time span based on the complexity of the effective reachable set, the boundary density of the physical counter-evidence set, and the conflict credibility record. The retrieval agent reads the time-series index within the elastic time span, generating positive time-series mutation data and counter-evidence time-series limit-crossing data. The system defines nodes with valid reachable mutation values as real abnormal nodes, physical counter-evidence nodes not falling within the preset transition time range as conflict nodes, and physical counter-evidence nodes whose limit-crossing time falls within this range or whose limit-crossing is caused by multiple sources within the same version as nodes awaiting review. The system generates a joint primary key for positive, counter-evidence, and review based on the node's identifier, version number, and time parameters to retrieve fault, collection anomaly, or configuration change logs. The generated node classification tags, anomaly types, collection sources, and conflict times are written back to the multimodal index module to update the conflict credibility record and adjust the retrieval path for similar limit-crossing events of the same version.
[0007] Preferably, when the synchronization center changes the port start / stop state, link negotiation state, sleep state, blocking state, or maintenance isolation state, or when the state refresh result is inconsistent with the physical state combination corresponding to the previous port state version number, it updates the port state version number and distinguishes the index data before and after the update according to the port state version number.
[0008] Preferably, the preset transition time range is determined based on at least one of the collection cycle, configuration synchronization delay, collection agent refresh cycle, and log entry delay, and is used to determine whether the timing violation that occurs after the port status version number is updated belongs to the violation situation corresponding to the node to be reviewed during the status switch period.
[0009] Preferably, when generating the physical counter-evidence set, the retrieval agent only includes nodes that have a direct topological adjacency with nodes in the valid reachable set and are excluded because their port physical states do not meet the valid reachability conditions in the physical counter-evidence set, and records the corresponding counter-evidence connection edge, port state version number, and physical state type.
[0010] Preferably, the flexible time span expands as the complexity of the effective reachable set increases, and tightens as the boundary density of the physical proof set or the conflict credibility record increases; when there is a port status version update, a change in the collection source, or a multi-source contradiction exceeding the limit, the tightening effect of the conflict credibility record on the flexible time span is reduced.
[0011] Preferably, the node classification marker includes a physical state conflict marker and a state version pending review marker; the system generates the physical state conflict marker for the conflicting node and prevents the conflicting node from entering the real abnormal node based on the corresponding counter-evidence timing limit violation data; the system generates the state version pending review marker for the node pending review and allows it to enter the log retrieval path corresponding to the review composite primary key.
[0012] Preferably, the multi-source conflict exceeding the limit includes: under the same port status version number, different collection sources produce conflicting results of exceeding the limit and not exceeding the limit for the same indicator, results of exceeding the limit in opposite directions, or results of exceeding the limit with a time difference exceeding a preset synchronization tolerance.
[0013] Preferably, the forward composite primary key includes node identifier, port identifier, port status version number, anomaly occurrence time, and temporal mutation type; the counter-evidence composite primary key includes node identifier, port identifier, port status version number, conflict time, data collection source, and the physical status conflict flag; the verification composite primary key includes node identifier, port identifier, port status version number, version update time, the preset transition time range, and the status version pending verification flag.
[0014] Preferably, the conflict confidence record includes node identifier, port identifier, port status version number, anomaly type, number of conflict occurrences, most recent conflict time, data collection source, and confidence value; when a data collection anomaly, sensor drift, or cached reporting log is retrieved, the corresponding confidence value is increased; when a real fault log is retrieved, the corresponding confidence value is decreased.
[0015] Preferably, when the same or similar type of counter-evidence time-series out-of-limit data appears again under the same port status version number, the system increases the retrieval priority of the counter-evidence composite primary key or the review composite primary key according to the conflict credibility record, and adjusts the elastic time span to reduce the retrieval path of similar out-of-limit data repeatedly entering the real abnormal node.
[0016] The technical effects and advantages of this invention are as follows: This invention, through the cooperation of a multi-source heterogeneous data sensing gateway, a dynamic configuration synchronization hub, and a multimodal index module, maps telemetry indicators, raw logs, physical link status on the acquisition side, and physical port status on the configuration side of the industrial network to node identifiers, port identifiers, and port status version numbers. This enables the topology adjacency graph index, time-series time window index, and full-text inverted index to perform associated retrieval based on the same port status, avoiding erroneous associations caused by existing operation and maintenance systems that rely solely on isolated analysis based on alarm time or log keywords. Simultaneously, during topology traversal, the system distinguishes between the set of validly reachable nodes and the set of physically contradictory nodes. Adjacent nodes physically in a blocked, dormant, unnegotiated, maintenance-isolated, or unreachable state are retained as contradictory objects. Combined with a flexible time span, positive time-series mutation data is generated for validly reachable nodes, and contradictory time-series exceeding-limit data is generated for physically contradictory nodes. Furthermore, the system determines whether the exceeding-limit time falls within a preset transition time range and other parameters. If there are any contradictions or exceeding limits in the data collection sources, nodes are categorized into genuine abnormal nodes, physical state conflict nodes, and nodes awaiting verification. This reduces false anomaly misjudgments caused by data collection cache reporting, sensor drift, state synchronization lag, or residual data during maintenance isolation. Furthermore, the system generates positive composite primary keys, negative proof composite primary keys, and verification composite primary keys to perform traffic-based retrieval of fault logs, data collection anomaly or drift logs, and port refresh or configuration change logs. This ensures that the output structured maintenance report not only locates genuine faults but also provides evidence for verifying the source of false anomalies and state versions. Further, the system writes back physical state conflict markers, state version awaiting verification markers, anomaly types, data collection source identifiers, and conflict occurrence times to the multimodal index module, updates the conflict credibility record, and adjusts the retrieval path for similar time-series exceeding limits under the same port state version accordingly. This reduces the probability of duplicate false anomalies entering the genuine fault retrieval path, improving the accuracy, interpretability, and continuous processing efficiency of industrial network maintenance retrieval. Attached Figure Description
[0017] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts:
[0018] Figure 1 This is a block diagram of the overall system structure of the present invention; Figure 2 This is a schematic diagram illustrating the data acquisition and configuration synchronization relationship of the present invention; Figure 3 This is a schematic diagram illustrating the relationships within the present invention; Figure 4 This is a schematic diagram of the process of the present invention; Figure 5 This is a schematic diagram of the flow determination process of the present invention; Figure 6 This is a flowchart illustrating the retrieval path of the present invention. Detailed Implementation
[0019] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.
[0020] Reference Figure 1-6 As shown, this invention provides a technical solution: a network information system operation and maintenance management system based on industrial big data and AI intelligent agents, applied to an industrial network environment. The industrial network environment may include industrial switches, industrial servers, edge gateways, PLC control nodes, data acquisition terminals, log servers, production management servers, and a host computer monitoring platform. The system includes a multi-source heterogeneous data sensing gateway, a dynamic configuration synchronization hub, a multimodal index module, an intent parsing intelligent agent, and a collaborative retrieval intelligent agent. The multi-source heterogeneous data sensing gateway communicates with network devices, computing nodes, acquisition agents, and log servers in the industrial network. The dynamic configuration synchronization hub communicates with network device configuration management interfaces, port status reporting interfaces, or the operation and maintenance management platform. The multimodal index module is connected to both the multi-source heterogeneous data sensing gateway and the dynamic configuration synchronization hub. The intent parsing intelligent agent and the collaborative retrieval intelligent agent communicate with the multimodal index module.
[0021] Multi-source heterogeneous data sensing gateways are used to collect telemetry metrics, raw logs, and physical link status on the acquisition side in industrial networks. Telemetry metrics include one or more of the following: port traffic, bit error rate, packet loss rate, communication latency, CPU utilization, memory utilization, device temperature, acquisition agent status, and service communication count. Raw logs include one or more of the following: device system logs, port event logs, acquisition agent logs, service anomaly logs, configuration change logs, maintenance operation logs, and status synchronization logs. Physical link status on the acquisition side includes one or more of the following reported by acquisition agents, probes, or network devices: port start / stop status, link negotiation status, dormant status, blocked status, maintenance isolation status, and unreachable status.
[0022] When collecting data, the multi-source heterogeneous data sensing gateway records the original timestamp of the data source, the gateway's receiving timestamp, the collection agent identifier, and the estimated latency. Based on the latency estimate, it performs collection timestamp calibration and latency normalization on data from different sources. The processed telemetry metrics, raw logs, and physical link status on the collection side are mapped to a unified state time base. Through this processing, the system can determine whether different time-series metrics, logs, and physical statuses are within the same port state version, reducing erroneous associations caused by asynchronous collection times.
[0023] The dynamic configuration synchronization hub is used to monitor changes in the configuration-side physical state of network nodes and their ports in real time. The configuration-side physical state is provided by the network device configuration management interface, port status reporting interface, or operation and maintenance management platform, and is used to characterize the port physical state confirmed by the configuration management side. The physical link state collected by the multi-source heterogeneous data perception gateway and the configuration-side physical state maintained by the dynamic configuration synchronization hub are compared under a unified state time base to identify inconsistencies between the collected-side state and the configuration-side state.
[0024] When the start / stop status, link negotiation status, sleep status, blocking status, or maintenance isolation status of any port changes, or when the status refresh result is inconsistent with the physical state combination corresponding to the previous port status version, the dynamic configuration synchronization center performs an auto-increment update of the port status version number corresponding to that port and records the version update time. The port status version number is used to represent the effective version of the port under a certain physical state combination. For example, when a port switches from the enabled state to the maintenance isolation state, the port status version number is updated from version one to version two; after the port is released from maintenance isolation and successfully renegotiation, the port status version number is updated from version two to version three. Through the port status version number, the system can distinguish the topology, timing indicators, and log records of the same port under different physical states.
[0025] The dynamic configuration synchronization hub also maintains a preset transition time range for each port status version. This preset transition time range represents the brief transition period that may occur after a port status version update, including configuration synchronization, data acquisition agent refresh, status index update, and log storage. The preset transition time range can be set based on port type, data acquisition agent refresh cycle, network device category, status synchronization latency, and historical synchronization time. For example, the preset transition time range for industrial switch ports can be set based on link negotiation time and data acquisition cycle; the preset transition time range for edge gateway ports can be set based on status reporting cycle and configuration refresh cycle.
[0026] The multimodal indexing module receives data from the multi-source heterogeneous data sensing gateway and the dynamically configured synchronization center, and builds and maintains a hybrid index library. The hybrid index library includes a topology adjacency graph index, a time-series time window index, and a full-text inverted index. The multimodal indexing module establishes a unified node identifier, port identifier, port status version number, data collection timestamp alignment relationship, and conflict confidence record for each network node and physical port.
[0027] Each edge in the topology adjacency graph index represents the physical connection between two ports or two nodes. Each edge is associated with at least the following records: start node identifier, end node identifier, start port identifier, end port identifier, current physical state, port status version number, and status update time. The time-series window index is organized by node identifier, port identifier, port status version number, metric type, data collection source identifier, and time slice, enabling rapid retrieval of telemetry metrics under the same port status version. The full-text inverted index is built using keywords, log type, node identifier, port identifier, port status version number, and log timestamp, allowing fault logs, data collection anomaly logs, status synchronization logs, configuration change logs, and maintenance operation logs to be retrieved based on a composite primary key.
[0028] Conflict credibility records are used to represent the credibility of records judged as having physical state conflicts or pending state version verification in historical searches for a specific node, port state version, and anomaly type. Conflict credibility records include node identifier, port identifier, port state version number, anomaly type, number of conflicts, most recent conflict time, data collection source identifier, conflict marker type, and credibility value. Credibility values can be updated based on the number of physical state conflicts, the number of pending state version verifications, the hit rate of counter-evidence logs, and the confirmation status of verification logs, for example, using exponential smoothing algorithms or cumulative decay models for quantitative calculation. Specifically, when counter-evidence logs hit data collection anomaly logs, sensor drift logs, or data collection agent cached reporting logs, the conflict credibility of the corresponding anomaly type is increased; when verification logs show that port state refresh, configuration changes, or maintenance operations caused a short-term limit violation, the corresponding record is marked as pending verification and credible; when fault logs confirm a real fault in the node, the corresponding conflict credibility is decreased. When the port status version number changes, the system will isolate the conflict credibility of the old version from the new version; when only non-critical fields change in the combination of the old and new physical statuses, the conflict credibility of the old version can also be inherited to the new version according to the preset attenuation coefficient.
[0029] When the system receives an abnormal alarm or maintenance retrieval request, the collaborative retrieval agent first determines the alarm source node, alarm port, alarm metric type, and alarm occurrence time. If the abnormal alarm or maintenance retrieval request does not directly specify the alarm source node, the intent parsing agent parses the request text, alarm fields, device name, port name, or log keywords to obtain candidate alarm source nodes, and sends the candidate alarm source nodes to the collaborative retrieval agent.
[0030] The collaborative retrieval agent reads the topological adjacency graph index and traverses the adjacency relationships starting from the alarm source node. During the traversal, the collaborative retrieval agent reads the port status version number and port physical status of each connection edge and determines whether the connection edge is in a truly active state. In this embodiment, a truly active state must at least meet the following conditions: the port is enabled, the link negotiation is successful, the port is not in a dormant, blocked, or maintenance isolated state, and the status update time of the corresponding physical status is earlier than the alarm occurrence time or the retrieval reference time corresponding to the maintenance retrieval request, and the preset transition time range set by the dynamic configuration synchronization center has been exceeded since the status update time.
[0031] Edges that are truly active are considered valid edges, and nodes reachable from the alarm source node along these valid edges are included in the set of valid reachable nodes. For nodes that are directly adjacent to nodes in the set of valid reachable nodes but are not included because their corresponding ports are in a blocked, dormant, unsuccessful negotiation, maintenance isolation, or unreachable state, the collaborative retrieval agent includes them in the set of physical counter-evidence nodes. The set of physical counter-evidence nodes also records the corresponding counter-evidence edge, port state version number, physical state type, and state update time. Through this method, physically unreachable boundary nodes are not directly discarded but are retained as counter-evidence objects for subsequent identification of false anomalies.
[0032] After the collaborative retrieval agent completes the traversal, it generates the first result. The first result includes the set of valid reachable nodes, the set of physically disproving nodes, the set of valid connections, the set of disproving connections, the corresponding port status version number, and the node physical status label. Among them, the range of the set of physically disproving nodes is limited to the vicinity of the physical blocking boundary of the valid reachable area, so as to avoid including unreachable nodes that are far away from the alarm source and have no direct relationship with the current alarm propagation path in the subsequent calculation range.
[0033] After receiving the first result, the intent-parsing agent calculates the local retrieval constraint value. This local retrieval constraint value determines the elastic time span of the subsequent reading time window index. The local retrieval constraint value can be calculated based on the structural complexity of the set of effectively reachable nodes, the boundary density of the set of physically contradictory nodes, and the conflict confidence record under the corresponding port state version number. The structural complexity of the set of effectively reachable nodes can be calculated based on at least two of the following: the number of nodes, node degree distribution, number of reachable levels, number of effective connecting edges, and node type distribution. The boundary density of the set of physically contradictory nodes can be calculated based on at least two of the following: the number of physically contradictory nodes, the number of contradictory connecting edges, the number of effectively reachable nodes, and the total number of traversed adjacent edges. For example, the boundary density can be set as the ratio of the number of contradictory connecting edges to the total number of traversed adjacent edges, or as the ratio of the number of physically contradictory nodes to the number of effectively reachable nodes, or a weighted sum of both.
[0034] Flexible time span Based on the basic time span The structural complexity of the set of effectively reachable nodes Boundary density of the set of nodes for physical proof and the credibility of the conflict The calculation yielded, where All values are normalized to between 0 and 1. Elastic time span. satisfy: ;in The preset weighting coefficients are used to calculate the elastic time span. Further restricted to the minimum time span and maximum time span Between. When the review condition is triggered, the system degrades. The value of is chosen to weaken the tightening effect of conflict credibility on the elastic time span.
[0035] When the structural complexity of the set of reachable nodes is high, it indicates that the current alarm may have propagated through many devices or levels, and the elastic time span tends to expand to reduce missed detections in complex local networks. When the boundary density of the set of physical counter-evidence nodes is high, it indicates that there are many physical blocking or isolation boundaries around the alarm source, and the elastic time span tends to tighten to reduce residual collected values outside the blocking boundaries from entering the scope of real fault retrieval. When the conflict confidence under the corresponding port status version number is high, it indicates that the timing limit exceeding the physical state has occurred multiple times under the same port status version, and the elastic time span tends to tighten, increasing the priority of subsequent counter-evidence retrieval and verification retrieval. Through the above weighted mapping function, the system transforms the set status of the first result into a determined elastic time span parameter.
[0036] To avoid missing genuine faults due to excessively high conflict confidence, a review trigger condition is set. When a port status version update, data collection source change, the duration of the violation exceeds the review threshold, multiple data collection sources produce contradictory violation results under the same port status version, or multiple independent data collection sources simultaneously violate the limit under the same port status version and the violation duration exceeds the review threshold, the intent parsing agent reduces the tightening weight of conflict confidence on the elastic time span and allows the collaborative retrieval agent to trigger a joint primary key retrieval for review. Therefore, the system can reduce invalid searches when there are many duplicate pseudo-anomalies, and retain review opportunities when state changes, multi-source contradictions, or multi-source persistent anomalies occur.
[0037] The collaborative retrieval agent reads the time-series time window index within the flexible time span and processes the set of effectively reachable nodes and the set of physically contradictory nodes separately. For nodes in the set of effectively reachable nodes, the collaborative retrieval agent extracts time-series indicator data related to the alarm indicator type and calculates the time-series mutation correlation value to form positive time-series mutation data. The time-series mutation correlation value can be jointly determined by the indicator mutation amount, mutation time proximity, out-of-limit duration, and reachability level weight, for example, using a linear weighted summation algorithm: ;in Values related to time-series abrupt changes. This represents the normalized index mutation rate. This represents the absolute time difference between the time of the mutation and the time of the alarm. For the duration of exceeding the limit, This is the number of reachable topology hops between this node and the alarm source node. These are preset weighting coefficients. This mathematical model achieves quantitative convergence of multi-dimensional features. The higher the magnitude of the indicator mutation, the closer the mutation time is to the alarm occurrence time, the longer the duration of the exceeded limit, and the closer the reachable level is to the alarm source node, the higher the correlation value of the time-series mutation. For example, for port traffic anomalies, we can calculate whether the traffic mutation direction is consistent between the alarm source node and the effectively reachable node within the elastic time span, whether the mutation times are adjacent, and whether the mutation magnitude exceeds the threshold. For CPU utilization or packet loss rate anomalies, we can calculate the mutation magnitude, duration, and correlation with the anomaly times of adjacent nodes within the elastic time span. Positive time-series mutation data is used to identify true anomaly candidate nodes.
[0038] For nodes in the physical counter-evidence node set, the collaborative retrieval agent does not include them in the actual fault ranking. Instead, it extracts their time-series index values, data source identifiers, time of occurrence of the violation, duration of the violation, and direction of the violation. A time-series violation marker is generated when the time-series index value reaches a second preset threshold, forming counter-evidence time-series violation data. The collaborative retrieval agent combines the positive time-series mutation data with the counter-evidence time-series violation data to generate a second result. The time-series violation marker indicates that although the node violates the telemetry index, it belongs to the physical counter-evidence node set in the first result, suggesting that the violation may be physically unsupported. The data source identifier is used to distinguish whether the violation data comes from the device itself, the data acquisition agent, the edge gateway, the mirror acquisition end, or the log parsing result. By retaining the data source identifier, the system can further determine whether the violation originates from a single data acquisition source's cached report, an abnormal data acquisition agent, or sensor drift.
[0039] The system then performs a consistency check on the first and second results. If a node belongs to the set of valid reachable nodes, and the temporal mutation correlation value of that node reaches a first preset threshold, and the temporal mutation state corresponding to that node meets the preset minimum duration condition, then the system identifies that node as a real anomaly candidate node. Real anomaly candidate nodes can first be sorted according to their temporal mutation correlation value, reachability level with the alarm source node, and alarm indicator type; after completing the fault log retrieval based on the positive composite primary key, the system performs a secondary sorting of real anomaly candidate nodes based on the log matching degree (e.g., the hit rate of fault feature keywords in the log text or the similarity score calculated by TF-IDF weight).
[0040] If a node belongs to the set of physical counter-evidence nodes, and its timing index value reaches the second preset threshold and generates a timing limit violation flag, and the time of this violation does not fall within the preset transition time range after the port status version update defined by the dynamic configuration synchronization hub, then the system generates a physical state conflict flag and prohibits the node from entering the set of real abnormal candidate nodes based on this timing limit violation flag. This physical state conflict flag indicates that the node is physically in a blocked, dormant, unsuccessful negotiation, maintenance isolated, or unreachable state, but its telemetry index still exceeds the limit. Therefore, this violation is more likely due to residual acquired values, acquisition agent cache reporting, sensor noise floor drift, or state synchronization anomalies.
[0041] If a node belongs to the set of physical counter-evidence nodes, and its timing indicator value reaches the second preset threshold and generates a timing violation flag, but the violation occurs within the preset transition time range after the port status version update, the system generates a status version pending review flag. If multiple data acquisition sources show contradictory violation results under the same port status version, the system also generates a status version pending review flag. Contradictory violation results include the same indicator exceeding the limit in one data acquisition source while not exceeding the limit in another, the violation directions of different data acquisition sources being opposite, the time difference of the violation occurrence of different data acquisition sources exceeding the preset synchronization tolerance, or the node's physical status indicating unreachability but a certain telemetry source continuously reporting high-frequency violations.
[0042] After completing the consistency check, different composite primary keys are generated based on different markers. For genuine anomaly candidate nodes, the system generates a positive composite primary key based on the node identifier, port identifier, port status version number, elastic time span, time-series mutation type, alarm indicator type, and anomaly occurrence time, and retrieves fault logs from the full-text inverted index based on the positive composite primary key. Fault logs include one or more of the following: device fault logs, service interruption logs, port error logs, resource exhaustion logs, and communication anomaly logs.
[0043] For nodes with physical state conflict markers, the system generates a counter-evidence composite primary key based on the node identifier, port identifier, port status version number, conflict occurrence time, data acquisition source identifier, anomaly type, and physical state conflict marker. Based on this counter-evidence composite primary key, the system retrieves one or more of the following from the full-text inverted index: data acquisition anomaly logs, sensor drift logs, data acquisition agent cache reporting logs, status synchronization anomaly logs, port status change logs, and maintenance isolation operation logs. The counter-evidence composite primary key is used to trace why the timing violation did not enter the actual fault candidate path, and to determine which data acquisition or status synchronization anomaly it might originate from.
[0044] For nodes marked with a status version pending review, the system generates a composite primary key for review based on the node identifier, port identifier, port status version number, version update time, preset transition time range, data collection source identifier, and status version pending review mark. Then, based on this composite primary key, the system retrieves one or more of the following logs from the full-text inverted index: port status refresh log, configuration change log, maintenance operation log, status synchronization log, and data collection agent reconnection log. The composite primary key is used to confirm whether the current violation is occurring during a port status switch, configuration refresh, or data collection source resynchronization process.
[0045] The search results for positive composite primary keys, negative proof composite primary keys, and verification composite primary keys generate a structured operation and maintenance report. The structured operation and maintenance report includes the alarm source node, real anomaly candidate nodes, physical state conflict nodes, nodes whose status versions need verification, the port status version number corresponding to each node, the elastic time span, fault log matching results, negative proof log matching results, verification log matching results, conflict filtering reasons, and the final judgment conclusion. For real anomaly candidate nodes, the report provides their time-series mutation-related values, corresponding fault logs, and suggested handling directions. For physical state conflict nodes, the report provides their physical unreachability reasons, out-of-limit collection sources, conflict occurrence time, and pseudo-anomaly reasons. For nodes whose status versions need verification, the report provides the basis for their falling within the preset transition time range, the contradiction in collection sources, and the configuration or status logs that need verification.
[0046] After outputting a structured operation and maintenance report, the system writes back the physical status conflict marker, status version pending verification marker, anomaly type, data collection source identifier, and conflict occurrence time to the multimodal index module to update the conflict credibility record under the corresponding node and port status version number. By continuously accumulating this record, the system dynamically downgrades the weight of nodes with repeated false alarms. When the same or similar anomaly type is subsequently received again under the same port status version, the multimodal index module provides the corresponding conflict credibility record to the intent parsing agent and the collaborative retrieval agent. The intent parsing agent tightens the elastic time span or increases the priority of counter-evidence retrieval based on the conflict credibility record. Before the verification trigger condition is met, the collaborative retrieval agent prioritizes performing counter-evidence joint primary key retrieval or verification joint primary key retrieval, reducing the repeated entry of duplicate false anomalies under the same port status version into the real fault retrieval path. This forms a historical conflict record-driven retrieval and update mechanism that includes retrieval, judgment, feedback write-back, and subsequent path adjustment.
[0047] In an exemplary scenario, port P1 of industrial switch A enters a new port state version V5 due to maintenance isolation. The dynamic configuration synchronization center records the state of port P1 as maintenance isolation and records the version update time and a preset transition time range. Subsequently, the multi-source heterogeneous data sensing gateway still receives traffic violation data corresponding to port P1 from a certain acquisition agent. When the collaborative retrieval agent traverses the topology adjacency graph index centered on the alarm source node, because port P1 is in maintenance isolation, the adjacent nodes connected to port P1 are included in the physical counter-evidence node set. If the traffic violation occurred after the preset transition time range, the system generates a physical state conflict marker and retrieves the acquisition agent's cached reporting logs or sensor drift logs through the counter-evidence federated primary key, not considering this node as a real fault candidate node. If the traffic violation occurred within the preset transition time range immediately after the maintenance isolation state update, the system generates a state version pending verification marker and retrieves the port state refresh logs and configuration change logs through the verification federated primary key. The above marking results are then written back to the multimodal index module to update the conflict confidence record of port P1 under port state version V5. If a similar violation occurs again on port P1 under the same port status version V5, the system will prioritize performing a counter-evidence search or a verification search to reduce duplicate false positives.
[0048] Through the above implementation methods, this embodiment can unify the processing of topology physical status, timing indicators, and log records under the port status version number during industrial network operation and maintenance. It also uses a set of effectively reachable nodes and a set of physically contradictory nodes to triage and judge timing anomalies. For physically reachable and timing-related nodes, the system can locate the actual fault and retrieve fault logs. For physically unreachable nodes that still exhibit timing exceedances, the system can identify physical status conflicts and trace collection anomalies or status synchronization anomalies. For nodes during status transitions, those with conflicting multiple collection sources, or those with persistent anomalies from multiple collection sources, the system can trigger status version verification. Furthermore, the system writes the conflict and verification results back to the multimodal index module, updates the conflict credibility record, thereby constraining subsequent similar exceedance retrieval paths under the same port status version, reducing the probability of false retrievals caused by duplicate pseudo-anomalies, and improving the accuracy and interpretability of structured operation and maintenance reports.
[0049] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. A network information system operation and maintenance management system based on industrial big data and AI intelligent agents, characterized in that, The system includes a perception gateway, a synchronization hub, a multimodal indexing module, an intent agent, and a retrieval agent. The perception gateway collects telemetry metrics, logs, and link status, and calibrates timestamps. The synchronization hub updates the port status version number when the physical status of the configured port changes and maintains a preset transition time range. The multimodal indexing module constructs a topology, time-series, and inverted index that associates node identifiers, port identifiers, and port status version numbers, and maintains conflict confidence records. When responding to an anomaly alarm or retrieval request, the retrieval agent reads the topology index starting from the identified alarm source node. The system defines a valid reachable set as nodes reachable by connected edges that are enabled, successfully negotiated, not dormant, not blocked, not isolated, and whose state update time exceeds the preset transition time range. It defines a physical counter-evidence set as nodes adjacent to these nodes but excluded due to blocking, dormancy, unsuccessful negotiation, maintenance isolation, or unreachability. The intent agent generates an elastic time span based on the complexity of the valid reachable set, the boundary density of the physical counter-evidence set, and the conflict credibility record. The retrieval agent reads the temporal index within the elastic time span, generating positive temporal mutation data and counter-evidence temporal limit violation data. The system defines nodes with valid reachable mutation values as real abnormal nodes, physical counter-evidence nodes that do not fall within the preset transition time range as conflict nodes, and physical counter-evidence nodes whose time exceeds the limit falls within the range or whose time exceeds the limit due to multiple sources within the same version as nodes to be reviewed. The system generates joint primary keys for positive, counter-evidence, and review based on the identifier, version number, and time parameters of the above nodes, respectively, to retrieve fault, collection anomaly, or configuration change logs. The system also writes back the generated node classification tags, anomaly types, collection sources, and conflict times to the multimodal index module to update the conflict credibility records and adjust the retrieval paths for similar time limits exceeding the limit within the same version.
2. The network information system operation and maintenance management system according to claim 1, characterized in that, When the port start / stop state, link negotiation state, sleep state, blocking state, or maintenance isolation state changes, or when the state refresh result is inconsistent with the physical state combination corresponding to the previous port state version number, the synchronization center updates the port state version number and distinguishes the index data before and after the update according to the port state version number.
3. The network information system operation and maintenance management system according to claim 1, characterized in that, The preset transition time range is determined based on at least one of the following: collection cycle, configuration synchronization delay, collection agent refresh cycle, and log entry delay. It is used to determine whether the timing violation that occurs after the port status version number is updated belongs to the violation situation corresponding to the node to be reviewed during the status transition period.
4. The network information system operation and maintenance management system according to claim 1, characterized in that, When generating the physical counter-proof set, the retrieval agent only includes nodes that have a direct topological adjacency with nodes in the valid reachable set and are excluded because their port physical states do not meet the valid reachability conditions. The agent also records the corresponding counter-proof connection edge, port state version number, and physical state type.
5. The network information system operation and maintenance management system according to claim 1, characterized in that, The elastic time span expands as the complexity of the effective reachable set increases, and tightens as the boundary density of the physical proof set or the conflict credibility record increases; when there is a port status version update, a change in the collection source, or a multi-source contradiction exceeding the limit, the tightening effect of the conflict credibility record on the elastic time span is reduced.
6. The network information system operation and maintenance management system according to claim 1, characterized in that, The node classification markers include physical state conflict markers and state version pending verification markers; the system generates the physical state conflict markers for the conflicting nodes and prevents the conflicting node from entering the real abnormal node based on the corresponding counter-evidence timing limit violation data; The status version pending review marker is generated for the node to be reviewed, and it is allowed to enter the log retrieval path corresponding to the review composite primary key.
7. The network information system operation and maintenance management system according to claim 1, characterized in that, The multi-source conflict exceeding the limit includes: under the same port status version number, different collection sources produce conflicting results of exceeding the limit and not exceeding the limit for the same indicator, results with opposite directions of exceeding the limit, or results where the time difference of the exceeding the limit exceeds the preset synchronization tolerance.
8. The network information system operation and maintenance management system according to claim 1, characterized in that, The forward composite primary key includes node identifier, port identifier, port status version number, anomaly occurrence time, and temporal mutation type; the counter-evidence composite primary key includes node identifier, port identifier, port status version number, conflict time, data collection source, and the physical status conflict flag; the verification composite primary key includes node identifier, port identifier, port status version number, version update time, the preset transition time range, and the status version pending verification flag.
9. The network information system operation and maintenance management system according to claim 1, characterized in that, The conflict credibility record includes node identifier, port identifier, port status version number, anomaly type, number of conflicts, most recent conflict time, data collection source, and credibility value. When a data collection anomaly, sensor drift, or cached reporting log is retrieved, the corresponding credibility value is increased; when a real fault log is retrieved, the corresponding credibility value is decreased.
10. The network information system operation and maintenance management system according to claim 1, characterized in that, When the same or similar type of counter-evidence time-series out-of-limit data appears again under the same port status version number, the system increases the retrieval priority of the counter-evidence composite primary key or the review composite primary key according to the conflict credibility record, and adjusts the elastic time span to reduce the retrieval path of similar out-of-limit data repeatedly entering the real abnormal node.