Multi-mode operation and maintenance data processing method and electronic equipment
By using multimodal data processing methods, operation and maintenance image and text data are mapped to the same semantic space and causal reasoning is performed, which solves the problem of insufficient single-modal data, improves the accuracy and efficiency of root cause determination of operation and maintenance events, and reduces operation and maintenance costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-14
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies rely on single-modal data for fault diagnosis in network intelligent operation and maintenance, resulting in limited information on equipment operating status and difficulty in fully presenting the associated scenarios and characteristics of operation and maintenance events. When multimodal data is fused, the semantic expression systems differ greatly, making it difficult to form an effective correlation mapping, leading to insufficient determination of the root cause of the fault, long diagnosis cycle and high cost.
By acquiring multimodal operation and maintenance data, including operation and maintenance images and text, feature extraction methods are used to form image feature sets and text feature sets. These are then mapped to the same semantic space through a shared projection layer. Based on feature similarity, the same operation and maintenance event is associated. A chain causal reasoning layer is used to determine the root cause and output the processing action.
It has achieved effective information integration of multimodal operation and maintenance data, improved the accuracy and efficiency of root cause determination of operation and maintenance events, reduced the frequency of manual intervention, shortened the fault handling cycle, and reduced network operation and maintenance costs.
Smart Images

Figure CN121723408A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of information technology, and more specifically, to a multimodal operation and maintenance data processing method and electronic device. Background Technology
[0002] Currently, in the field of intelligent network operations and maintenance (O&M), many technologies rely on single-modal data (such as performance indicator curves or device status logs) for fault diagnosis and maintenance decisions, achieving preliminary judgment of O&M events through the analysis of a single type of data. Meanwhile, some solutions attempt to introduce multimodal data fusion approaches, such as improving O&M task performance through joint analysis of signal and text data, or achieving multimodal detection of network traffic based on rule extraction, to address O&M needs in complex network environments.
[0003] However, existing technologies still face many challenges in practical applications: the equipment operating status information that single-modal data can carry is limited, making it difficult to fully present the associated scenarios and characteristics of maintenance events, resulting in insufficient determination of the root cause of failures; when multimodal data is fused, the semantic expression systems of different types of data (such as visual information and text descriptions) are different, making it difficult for various types of data to form an effective association mapping, affecting the logical coherence of maintenance decisions; in addition, the adaptability of existing solutions is poor in the face of massive amounts of incompletely labeled maintenance data, and it is difficult to quickly adjust the adaptation strategy in dynamic scenarios such as network topology changes and the emergence of new faults, resulting in a long fault diagnosis cycle, a high frequency of manual intervention, and significant room for improvement in overall maintenance efficiency and cost control.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this disclosure is to provide a multimodal operation and maintenance data processing method and electronic device to shorten the operation and maintenance fault handling cycle and reduce network operation and maintenance costs.
[0006] According to a first aspect of the present disclosure, a multimodal operation and maintenance data processing method is provided, comprising: acquiring multimodal operation and maintenance data, wherein the multimodal operation and maintenance data includes operation and maintenance images and operation and maintenance text, wherein the operation and maintenance images include actual photos of the data center, equipment connection topology diagrams, and curves of preset performance indicators, and the operation and maintenance text includes speech-recognized text and written report text; extracting image features of the operation and maintenance images using a feature extraction method corresponding to the types of the operation and maintenance images to form an image feature set, extracting contextual semantic features of the operation and maintenance text to form a text feature set, and connecting the image feature set with the text feature set through a shared projection layer. The feature set is projected onto the same semantic space of a preset dimension; the feature similarity between the image feature set and the text feature set is determined in the same semantic space; when the feature similarity is greater than a preset threshold, the image feature set and the text feature set are determined to be associated with the same operation and maintenance event, forming an associated feature set, which includes the image feature set, the text feature set, and the operation and maintenance event; the associated feature set is input into the inference layer, and the inference layer performs chain causal inference on the operation and maintenance event to determine the root cause of the operation and maintenance event, and outputs the processing action corresponding to the root cause of the operation and maintenance event.
[0007] According to a second aspect of this disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method as described in any one of the preceding methods based on instructions stored in the memory.
[0008] According to a third aspect of this disclosure, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the multimodal operation and maintenance data processing method as described in any of the preceding claims.
[0009] According to a fourth aspect of this disclosure, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in any of the preceding claims.
[0010] This embodiment of the disclosure acquires multiple types of operation and maintenance images and operation and maintenance text, and uses feature extraction methods adapted to image types and contextual semantic feature extraction methods to obtain two types of feature sets respectively. Through a shared projection layer, cross-modal features are mapped to the same semantic space. Based on feature similarity, the same operation and maintenance event is associated and a set of associated features is formed. Then, through chain causal reasoning in the inference layer, the root cause of the operation and maintenance event is determined and the corresponding processing action is output. This can fully integrate the effective information of multimodal operation and maintenance data, improve the degree and efficiency of root cause determination of operation and maintenance events, reduce the frequency of manual intervention, shorten the fault handling cycle, reduce network operation and maintenance costs, and adapt to the diverse operation and maintenance data and real-time processing needs in complex network environments.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0013] Figure 1 This is a flowchart of a multimodal operation and maintenance data processing method in an exemplary embodiment of this disclosure.
[0014] Figure 2 This is a schematic diagram of a pre-trained model provided in an exemplary embodiment of this disclosure.
[0015] Figure 3 This is a sub-flowchart of step S3 in an exemplary embodiment of this disclosure.
[0016] Figure 4 This is a sub-flowchart of step S4 in an exemplary embodiment of this disclosure.
[0017] Figure 5 This is a sub-flowchart of method 100 in an exemplary embodiment of this disclosure.
[0018] Figure 6 This is a diagram illustrating the training process of the inference layer in an exemplary embodiment of this disclosure.
[0019] Figure 7 This is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0020] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0021] Furthermore, the accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0022] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0023] Figure 1 This is a flowchart of a multimodal operation and maintenance data processing method in an exemplary embodiment of this disclosure.
[0024] refer to Figure 1 The multimodal operation and maintenance data processing method 100 may include: Step S1: Obtain multimodal operation and maintenance data, which includes operation and maintenance images and operation and maintenance text. The operation and maintenance images include real-time photos of the data center, equipment connection topology diagrams, and curves of preset performance indicators. The operation and maintenance text includes voice-recognized text and written report text. Step S2: Extract image features of the operation and maintenance image using a feature extraction method corresponding to the type of the operation and maintenance image to form an image feature set; extract contextual semantic features of the operation and maintenance text to form a text feature set; and project the image feature set and the text feature set to the same semantic space of a preset dimension using a shared projection layer. Step S3: Determine the feature similarity between the image feature set and the text feature set in the same semantic space. When the feature similarity is greater than a preset threshold, determine that the image feature set and the text feature set are associated with the same operation and maintenance event, forming an associated feature set. The associated feature set includes the image feature set, the text feature set, and the operation and maintenance event. Step S4: Input the associated feature set into the inference layer, and use the inference layer to perform chain causal reasoning on the operation and maintenance event to determine the root cause of the operation and maintenance event, and output the processing action corresponding to the root cause of the operation and maintenance event.
[0025] This embodiment of the disclosure acquires multiple types of operation and maintenance images and operation and maintenance text, and uses feature extraction methods adapted to image types and contextual semantic feature extraction methods to obtain two types of feature sets respectively. Through a shared projection layer, cross-modal features are mapped to the same semantic space. Based on feature similarity, the same operation and maintenance event is associated and a set of associated features is formed. Then, through chain causal reasoning in the inference layer, the root cause of the operation and maintenance event is determined and the corresponding processing action is output. This can fully integrate the effective information of multimodal operation and maintenance data, improve the degree and efficiency of root cause determination of operation and maintenance events, reduce the frequency of manual intervention, shorten the fault handling cycle, reduce network operation and maintenance costs, and adapt to the diverse operation and maintenance data and real-time processing needs in complex network environments.
[0026] The following is a detailed explanation of each step in the multimodal operation and maintenance data processing method 100.
[0027] In step S1, multimodal operation and maintenance data is acquired. The multimodal operation and maintenance data includes operation and maintenance images and operation and maintenance text. The operation and maintenance images include actual photos of the data center, equipment connection topology diagrams, and curves of preset performance indicators. The operation and maintenance text includes speech recognition text and written report text.
[0028] In this embodiment of the disclosure, the operation and maintenance data used for root cause analysis is multimodal data generated during network operation and maintenance that can reflect the operating status of equipment, fault phenomena and operation and maintenance scenario information. This provides a comprehensive and multidimensional basis for the determination and handling of the root causes of operation and maintenance events, avoiding the one-sidedness of information caused by a single data type.
[0029] Multimodal operation and maintenance data includes, but is not limited to, operation and maintenance images and operation and maintenance text.
[0030] The operation and maintenance images visually present the physical status, connection relationships, and dynamic operating trends of devices and networks, and can include various types of images. Any image that reflects the operation and maintenance status can be used as an operation and maintenance image. In an exemplary embodiment, the operation and maintenance images include real-life photos of the data center, device connection topology diagrams, and graphs of preset performance indicators.
[0031] Data center photos refer to images of the equipment and its surrounding environment taken on-site in the data center. These photos can include details of the status of key components, such as the color of switch port indicator lights (flashing red, solid green, etc.), port labels (e.g., G0 / 1 port number), error codes displayed on the equipment casing, and the plugging and unplugging status of cables (e.g., loose fiber optic connectors, tangled and worn cables). These photos directly reflect the physical operating status of the equipment.
[0032] A device connection topology diagram is a visual graphic that reflects the connection relationship between network devices. Nodes represent network devices (such as switches, routers, and servers), and edges represent communication links between devices. It can also label attributes such as link bandwidth, real-time connection status, and traffic distribution. For example, it can show the connection link between aggregation layer switches and routers, the relationship between branch nodes and aggregation nodes, etc. It can display the network topology and the dependencies between devices.
[0033] The preset performance index curves are time-series visualization charts showing how device performance parameters change over time. Preset performance indexes include, but are not limited to, bandwidth utilization, port packet loss rate, link latency, CPU utilization, memory usage, etc. For example, a curve showing how the packet loss rate of a certain port gradually increases from 0.1% to 15% in the past 24 hours, or a trend chart showing that the bandwidth utilization rate remains above 90% during peak hours. These can dynamically reflect the operating load and health status of the device and the link.
[0034] Operations and maintenance (O&M) texts are verbal descriptions of O&M scenarios by O&M personnel, including spoken feedback and written records, containing the O&M personnel's experiential judgments and descriptions of fault phenomena. O&M texts can include speech recognition text and written report text.
[0035] Voice-recognized text is the text that is recorded by maintenance personnel using voice input tools during inspections and troubleshooting, and then recognized and converted. Examples include "The switch with number x keeps dropping BGP neighbors, and it disconnects every 10 minutes", "The router indicator light in rack 3 of the data center keeps flashing red", and "The aggregation layer link has been very unstable recently, and the service is occasionally stuck". Because voice information is direct and highly contextual, voice-recognized text can quickly convey the fault phenomena that are discovered immediately.
[0036] Written reports are standardized written records created by operations and maintenance personnel after routine maintenance and fault debriefing. Examples include "On [Date], the aggregation layer link experienced abnormal stability with multiple short-term interruptions, affecting branch nodes" and "Intermittent packet loss occurred on the G0 / 2 port of the business access device S7706, requiring investigation of physical link or interface faults." These reports are relatively rigorous and include information such as key times, device models, and the scope of the fault's impact, which can be used to pinpoint the scenario.
[0037] In some embodiments, the operation and maintenance data also includes device operation audio data. Device operation audio data refers to sound signals collected during device operation, such as abnormal noises from server fans (e.g., scraping sounds, high-frequency whistling), current noises caused by poor port contact, abnormal sounds from hard drive read / write operations, etc. It can supplement potential clues of device failure from an acoustic perspective, further enrich the information dimension of multimodal data, and improve the comprehensiveness of operation and maintenance event judgment.
[0038] In step S2, image features of the operation and maintenance image are extracted using a feature extraction method corresponding to the type of the operation and maintenance image to form an image feature set. Contextual semantic features of the operation and maintenance text are extracted to form a text feature set. The image feature set and the text feature set are then projected to the same semantic space of a preset dimension using a shared projection layer.
[0039] In step S2, after feature extraction of multimodal operation and maintenance data, dimensional unification and semantic alignment are used to lay the foundation for subsequent operation and maintenance event correlation and root cause reasoning, thereby comprehensively obtaining accurate results from multimodal data and overcoming the problem of decreased analysis accuracy caused by different data types.
[0040] In some embodiments, after acquiring the operation and maintenance image, the operation and maintenance image is first preprocessed by pixel normalization and key area cropping, and then image features are extracted. The key areas are determined based on the fault-related areas of the preset operation and maintenance scenario. The key areas include the device interface area, indicator light area, and topology graph node connection area.
[0041] For example, firstly, all types of operation and maintenance images are uniformly normalized (e.g., adjusted to 224×224 pixel specifications) to ensure the consistency of input for feature extraction; then, based on the preset fault-related areas of the operation and maintenance scenario, key areas of the image are cropped, such as high-frequency fault-related areas like device interface areas, indicator light areas, and topology graph node connection areas, and redundant information such as background and irrelevant equipment are removed, which not only improves the efficiency of feature extraction but also enhances the recognizability of fault-related features.
[0042] After preprocessing, features of various operation and maintenance images are extracted according to the principle of image type-adaptive extraction method to form an image feature set.
[0043] In an exemplary embodiment, extracting image features may include: using a convolutional neural network or a visual transformer to extract local and global scene features from a real-world image of a computer room; using a graph convolutional network to extract node and link association features from a device connection topology diagram; and using a temporal convolutional network to extract trend and fluctuation features from a curve of a preset performance index.
[0044] This disclosure employs different methods for feature extraction on different types of images to improve information accuracy and optimize the accuracy of subsequent analysis results. It is understood that since the types of maintenance images used each time may differ, any of the following extraction methods can be selected for implementation. Alternatively, based on the specific type of maintenance image and its similarity to the three image types mentioned above, the appropriate method for image feature extraction can be determined in real-time.
[0045] Specifically, for real-world images of the server room, feature extraction can be performed using convolutional neural networks (such as ResNet-50) or visual transformers (such as ViT-Base). In an exemplary embodiment, detailed information such as indicator light colors, port label characters, and error codes can be captured through local feature extraction. Simultaneously, overall environmental information such as device layout and cable connection status can be obtained through global scene feature extraction, achieving comprehensive feature extraction covering both details and the overall scene. For example, after the input image is preprocessed to 224×224 pixels, local features (such as indicator light colors and port label characters) and global scene features (such as device layout and cable connection status) are extracted through convolutional layers or Transformer layers. Local features can be obtained by locating key areas (such as indicator light areas) through an attention mechanism and then using ROI pooling. Global features can be obtained by outputting 2048-dimensional (ResNet-50) or 768-dimensional (ViT-Base) vectors through an average pooling layer.
[0046] For the device connection topology graph, Graph Convolutional Networks (GCNs) can be used for feature extraction. Nodes in the topology graph represent network devices, and edges represent links between devices. Network layer operations capture the correlation between node attributes (such as device type and real-time CPU utilization) and link attributes (such as link bandwidth and latency), outputting graph structure features that reflect the network topology and connection status. For example, an adjacency matrix can be constructed based on the topology connections, and after two layers of GCN convolution, the output graph structure features with a node dimension of 256 can be obtained.
[0047] For the curves of preset performance indicators, a Temporal Convolutional Network (TCN) can be used for feature extraction. By employing multiple layers of dilated convolutions, the long-range dependencies of the time-series data can be captured, extracting trend features (such as continuously increasing bandwidth utilization and stepwise growth in packet loss rate) and volatility features (such as sudden fluctuations in latency and periodic fluctuations in indicators) to reflect the dynamic changes in the equipment's operating status. For example, given time-series data with a time step of 24 (such as bandwidth utilization in the previous 24 hours), three dilated convolutional layers (with dilation factors of 1, 2, and 4) can capture long-range dependencies, outputting time-series features with a dimension of 128.
[0048] While extracting image features, contextual semantic features of the operation and maintenance text can also be extracted using a large language model to form a text feature set. In an exemplary embodiment, a pre-trained large language model (such as BERT-base or LLaMA-7B) can be used as the base model. The operation and maintenance text (including speech recognition text and written report text) is first preprocessed by word segmentation and the addition of semantic tags (such as [CLS] tags). Then, contextual semantic modeling is performed through the model's multi-layer Transformer structure. For example, the large language model takes the output at the [CLS] position as the contextual semantic feature, with a dimension of 768 (BERT-base) or 4096 (LLaMA-7B), and performs domain fine-tuning on operation and maintenance corpora (such as historical fault logs) to enhance the understanding of technical terms.
[0049] Therefore, by leveraging the contextual understanding capabilities of the large language model, it can not only capture the literal semantics of fault descriptions such as frequent port interruptions and BGP neighbor loss, but also understand the contextual relevance of colloquial expressions (such as the semantic equivalence between constant interruptions and intermittent interruptions) and the logical relevance of written texts (such as the causal relationship between link anomalies and business stagnation), ultimately outputting high-dimensional vector features that can represent the semantics of the text.
[0050] The above-mentioned feature extraction process can be operated through the feature extraction layer of a pre-trained model, which can be used to implement method 100.
[0051] Figure 2 This is a schematic diagram of a pre-trained model provided in an exemplary embodiment of this disclosure.
[0052] refer to Figure 2 In this embodiment of the disclosure, the pre-trained model 200 includes at least an input layer 21, a feature extraction layer 22, a shared projection layer 23, and an inference layer 24.
[0053] The input layer 21 is used to preprocess the input multimodal operation and maintenance data. In step S1, the multimodal operation and maintenance data can be directly input into the input layer 21.
[0054] The feature extraction layer 22 may include various neural network layers or processing modules as mentioned in the above embodiment for image feature extraction, and may also include a large language model that is built-in or called externally.
[0055] The shared projection layer 23 is used for subsequent processing in step S2.
[0056] The inference layer 24 is used to output root cause analysis and processing actions in step S4.
[0057] In step S2, after forming the image feature set and the text feature set, the dimensionality unification and semantic mapping of cross-modal features are achieved through the shared projection layer 23. For example, the above-mentioned image feature set (including various features of the real-world images of the computer room, topology maps, and performance curves) and text feature set can be input into the shared projection layer 23 of the pre-trained model 200. Through linear transformation, the high-dimensional features of heterogeneous modalities are uniformly mapped to the same semantic space of a preset dimension (for example, the input dimensions are the maximum dimension of visual features, 2048 dimensions, and the text features, 768 dimensions, respectively, and the output is uniformly 768 dimensions).
[0058] The shared projection layer 23 can eliminate the dimensional differences and semantic barriers of different modal features, making the visual features of the indicator light flashing red and the text features of poor port contact comparable in the same semantic space, providing a premise for subsequent similarity-based operation and maintenance event association determination.
[0059] In some embodiments, when the provided multimodal operation and maintenance data includes audio data, the processing flow of step S2 further includes an audio feature extraction step.
[0060] For example, the audio feature extraction step includes first preprocessing the collected device operating audio data: converting the audio data into mono audio segments with a preset sampling rate (e.g., 16kHz), reducing environmental noise interference and improving the recognition of effective signals through operations such as pre-emphasis and frame-by-frame windowing (e.g., using Hamming windows); then extracting features from the preprocessed audio data to form an audio feature set. For example, first converting the audio signal into a Mel spectrogram (e.g., 320×128 pixel specification) through short-time Fourier transform, and then using a convolutional neural network (CNN, 3 layers of convolution + max pooling) to extract time-frequency local features (e.g., frequency distribution and intensity changes of abnormal noise) from the spectrogram. At the same time, using a long short-term memory network (LSTM, e.g., 128-dimensional hidden layer) or a gated recurrent unit (GRU) to capture the time-dependent features of the audio signal (e.g., duration of fan noise and interval patterns of current noise), and outputting audio features (e.g., 256 dimensions) to capture fault-related signals in the device operating audio.
[0061] By introducing audio modalities to collect abnormal sounds during device operation and extract audio features, the multimodal alignment capability of Method 100 can be extended, allowing for the perception of the device's operating status from more dimensions. For example, abnormal sounds during device operation may be early signals of faults, and introducing audio modalities can help detect potential faults earlier.
[0062] The remaining processing steps are the same as described above: After pixel normalization and key area cropping preprocessing of the operation and maintenance images, image features are extracted according to image type using convolutional neural networks / visual transformers, graph convolutional networks, and temporal convolutional networks to form image feature sets; for operation and maintenance text, a pre-trained large language model is used to extract contextual semantic features to form text feature sets; finally, the image feature sets, text feature sets, and newly added audio feature sets are input into the shared projection layer 23, and the three types of heterogeneous modal features are uniformly mapped to the same semantic space of a preset dimension (such as 768 dimensions) through linear transformation, eliminating the modal differences and semantic barriers between audio, image, and text data, so that audio features such as current noise, visual features such as red indicator light flashing, and text features such as poor port contact are comparable in the same semantic space, providing more comprehensive feature support for subsequent multimodal feature association of the same operation and maintenance event.
[0063] In step S3, the feature similarity between the image feature set and the text feature set is determined in the same semantic space. When the feature similarity is greater than a preset threshold, the image feature set and the text feature set are associated with the same operation and maintenance event, forming an associated feature set. The associated feature set includes the image feature set, the text feature set, and the operation and maintenance event.
[0064] After mapping cross-modal features to the same semantic space in step S2, all image feature sets and text feature sets are now in the same semantic coordinate system, providing a basis for direct similarity comparison. Therefore, in step S3, the correlation determination of cross-modal features is completed, and multimodal feature combinations pointing to the same operation and maintenance event are selected by quantifying feature similarity, providing joint feature input for subsequent root cause inference.
[0065] In step S3, the feature similarity between the image feature set and the text feature set is first quantified and calculated in the same semantic space based on a preset similarity calculation rule. For example, a cosine similarity algorithm can be used as the calculation method. By measuring the angle between two types of feature vectors, the semantic association is determined: the smaller the vector angle, the closer the similarity value is to 1, indicating that the operation and maintenance information carried by the two types of features is more related; the larger the vector angle, the closer the similarity value is to 0, indicating that the operation and maintenance information corresponding to the two types of features is less related.
[0066] To ensure the accuracy of correlation determination, a local key feature weighting strategy can be adopted during the similarity calculation process. For example, higher calculation weights can be assigned to local key features such as device indicator status, port labels, and error codes in the image feature set, and semantic key features such as fault description words (e.g., interruption, packet loss, anomalies) in the text feature set. This avoids redundant information in the global features from interfering with the correlation determination results, making the similarity calculation more focused on the characteristics of the operation and maintenance event. See the exemplary embodiment of similarity calculation for details. Figure 3 describe.
[0067] After obtaining the similarity calculation results, the calculated feature similarity is compared with a preset threshold (the preset threshold can be determined based on training with historical associated samples in the operation and maintenance field, and its value range adapts to the needs of the operation and maintenance scenario). If the feature similarity is greater than the preset threshold, it indicates that the current image feature set and text feature set reflect the same operation and maintenance event (for example, if the image feature of a flashing red indicator light and the text feature of frequent port interruptions have a similarity higher than the preset threshold, it can be determined that both point to the operation and maintenance event of a physical port connection failure); if the feature similarity is less than or equal to the preset threshold, it is determined that the two types of features belong to different operation and maintenance events, and no association integration is performed.
[0068] Finally, the image feature sets and text feature sets determined to be related to the same maintenance event are integrated to form a related feature set. This related feature set includes not only the original image feature set and text feature set, but also the association identifier of the same maintenance event, clarifying the attribution relationship between the two types of feature sets. This allows the subsequent inference layer to directly conduct root cause analysis of maintenance events based on the integrated joint features, avoiding the breakdown of inference logic caused by the scattered processing of multimodal features.
[0069] Figure 3 This is a sub-flowchart of step S3 in an exemplary embodiment of this disclosure.
[0070] refer to Figure 3 In an exemplary embodiment, the process of calculating similarity in step S3 may include: Step S31: Calculate the similarity of local key features between the image feature set and the text feature set. Local key features include device indicator status, port labels, error codes in the maintenance image, and fault description words in the maintenance text. Step S32: Calculate the global scene feature similarity between the image feature set and the text feature set; Step S33: Perform a weighted calculation on the local key feature similarity and the global scene feature similarity to determine the feature similarity between the image feature set and the text feature set.
[0071] In step S31, local key features refer to features that are directly related to the essence of the operation and maintenance event and can quickly pinpoint the fault. The types of local key features corresponding to different modalities can be determined manually in advance, and then local key features in the feature set of the corresponding modality can be extracted according to the preset types of local key features. For example, the types of local key features corresponding to the image modality may include, but are not limited to, device indicator status (such as flashing red, solid green), port labels (such as G0 / 1 port number), error codes (such as ERR-001 displayed on the device casing), and other local features directly related to the fault; the types of local key features corresponding to the text modality may include, but are not limited to, semantic features such as fault description words (such as interruption, packet loss, abnormality, lag, neighbor loss, etc.).
[0072] After extracting local key features from images and text, the matching degree between these local key features and text local key features is quantified using similarity algorithms such as cosine similarity, resulting in local key feature similarity. This local key feature similarity reflects the correlation strength between features of two modalities at the fault information level. For example, the local key feature similarity between the image local features of a flashing red indicator light and the text fault description words of frequent port interruptions will be significantly higher.
[0073] In step S32, global scene features refer to the overall information supporting the context of operation and maintenance events. They are used to supplement scene correlations not covered by local key features, avoiding biases in correlation judgments due to insufficient clues. Similarly, the types of global scene features corresponding to features of different modalities can be determined manually in advance, thereby extracting global scene features from the feature set of the corresponding modality based on the preset types of global scene features.
[0074] In specific processing, the types of global scene features corresponding to the image modality may include, but are not limited to, the layout of equipment in the data center, the overall status of cable connections (image type), the overall association of network topology (topology graph type), and the overall trend of performance indicator curves (curve graph type); the types of global scene features corresponding to the text modality may include, but are not limited to, the occurrence scenario of the operation and maintenance event (such as the data center aggregation layer link), the scope of impact (such as the full user access of branch node services), time characteristics (such as peak periods lasting 2 hours), and other contextual information.
[0075] After extracting global scene features, similarity algorithms such as cosine similarity can be used to calculate the degree of matching between the two types of features in the global scene dimension, thus obtaining the global scene feature similarity. For example, the global scene features of an image from a data center switch and the global scene features of text from a lost BGP neighbor of the switch will maintain a high level of global scene feature similarity, providing scene support for local feature association.
[0076] In step S33, to balance the accuracy of local clues with the comprehensiveness of the global scene, the local key feature similarity obtained in step S31 and the global scene feature similarity obtained in step S32 are weighted and fused to finally determine the comprehensive feature similarity of cross-modal features. The weighting coefficients can be optimized based on the actual needs of the operation and maintenance scenario. For example, the weight of local key feature similarity can be set higher (e.g., a weight ratio of 0.7) to ensure the dominant role of fault clues; the weight of global scene feature similarity can be relatively lower (e.g., a weight ratio of 0.3), mainly used to correct the judgment results when local features are insufficient.
[0077] Ultimately, the feature similarity calculated by using the formula Feature Similarity = Local Key Feature Similarity × Local Weight + Global Scene Feature Similarity × Global Weight ensures that matching clues leads to association, while also preventing the omission of isolated local features that are highly relevant to the scene. This significantly improves the reliability of cross-modal feature association determination.
[0078] In an exemplary scenario, when the local features of an image showing a flashing yellow indicator light on a port match the text description of unstable link signals in the aggregation layer, the local key features have a high similarity. Even if there are slight differences in the global scene features, a high overall similarity will still be obtained, and the two will be determined to be related to the same maintenance event. However, if there is no obvious match in the local key features, but the global scene features are highly consistent (such as the global features of an image of abnormal equipment in the data center and the global features of a text description of a fault in a certain equipment in the data center), the addition of global weights can also avoid misjudging it as an irrelevant event, providing a more comprehensive basis for association for the subsequent inference layer.
[0079] In some embodiments, if the operation and maintenance data includes an audio feature set, the logic for calculating feature similarity is the same as... Figure 3 As shown, the feature similarity of the audio feature set, image feature set and text feature set is determined in the same semantic space. When the feature similarity is greater than a preset threshold, the audio feature set, image feature set and text feature set are associated with the same operation and maintenance event to form an associated feature set. The associated feature set includes the audio feature set, image feature set, text feature set and operation and maintenance event.
[0080] For example, the similarity between the audio feature set and the image feature set and the text feature set can be calculated separately. If the similarity between the audio feature set and any other modal feature set is greater than a preset threshold, the audio feature set is included in the associated feature set to form an associated feature set that fuses the image, text and audio three modalities. This further enriches the feature dimensions of the operation and maintenance event and improves the comprehensiveness of the subsequent root cause determination.
[0081] In step S4, the associated feature set is input into the inference layer, and the root cause of the operation and maintenance event is determined by chain causal reasoning through the inference layer. The processing action corresponding to the root cause of the operation and maintenance event is then output.
[0082] Step S4 can transform the associated feature set into the root cause determination result of the operation and maintenance event and the executable processing actions through the chain-like causal reasoning of the inference layer 24, realizing the key transformation from multimodal feature fusion to operation and maintenance decision output. For example, the inference layer 24 can adopt a Transformer decoder architecture, with the input being a multimodal feature sequence (visual features + text features + acoustic features), which is processed by 6 decoder layers and outputs the inference chain, the root cause determination result, and the corresponding processing actions.
[0083] For example, the set of associated features formed in step S3 (which can be image-text bimodal or image-text-audio trimodal fusion features, with the same operation and maintenance event association identifier) is completely input into the inference layer 24. The inference layer 24 is based on the preset operation and maintenance event-root cause-handling action association logic, and has a built-in causal reasoning knowledge base built based on historical failure cases and expert experience, which can quickly locate the contradictions of operation and maintenance events based on multimodal fusion features.
[0084] Figure 4 This is a sub-flowchart of step S4 in an exemplary embodiment of this disclosure.
[0085] refer to Figure 4 In an exemplary embodiment, step S4 may include: Step S41: Determine the fault phenomenon in the matching of associated feature sets; Step S42: Determine at least one hypothetical root cause corresponding to the fault phenomenon; Step S43: Call the historical operation and maintenance event causal reasoning path library, and determine the target hypothetical root cause with the highest degree of matching with the associated feature set as the root cause of the operation and maintenance event by matching the similarity between the associated feature set and the causal reasoning path corresponding to each hypothetical root cause. Step S44: Determine the root cause of the maintenance event and output the corresponding processing action.
[0086] Figure 4 The embodiment shown achieves the logical chain deduction of the final root cause through phenomenon matching, hypothesis root cause generation, reasoning path comparison, and processing action output, and improves the accuracy and efficiency of root cause determination by leveraging historical operation and maintenance experience.
[0087] In step S41, the inference layer 24 first integrates and parses the input associated feature set using multimodal information, extracting key performance dimensions of the operation and maintenance event from image, text (and audio, if any) features, and extracting concrete fault phenomena. For example, combining the image feature of the flashing red indicator light on the G0 / 1 port of the aggregation layer switch, the text feature of frequent port interruptions and repeated down BGP neighbor status, and the audio feature of slight current noise in the port area, the inference layer 24 integrates the following to obtain a clear fault phenomenon: intermittent signal transmission interruption on the G0 / 1 port of the aggregation layer switch, accompanied by abnormal physical connection characteristics, providing a clear starting point for subsequent root cause inference.
[0088] In step S42, based on the visualized fault phenomenon in step S41, the inference layer 24 can call the built-in historical operation and maintenance experience library (covering common fault types and root cause association rules) to generate at least one hypothetical root cause highly related to the fault phenomenon. This process must ensure the comprehensiveness of the hypothetical root causes and avoid omitting key possibilities. For example, for a fault phenomenon characterized by intermittent signal interruption and physical anomalies at the port, the generated hypothetical root causes may include fiber optic interface oxidation, port physical looseness, poor contact of the interface board, link transmission medium loss, abnormal port configuration parameters, etc., with each hypothetical root cause corresponding to a certain manifestation of the fault phenomenon.
[0089] In some embodiments, the inference layer 24 can also eliminate contradictory root causes through multimodal feature cross-validation, making the root causes derived in step S42 more accurate. For example, it can combine the feature of a surge in packet loss rate but normal bandwidth utilization in the performance index curve to eliminate the link bandwidth overload hypothesis; or, it can combine the feature of only the target port link being abnormal while other ports are normal in the topology diagram to eliminate the board damage hypothesis, and finally filter out the root causes that match all modal features (such as fiber optic interface oxidation).
[0090] In step S43, in this embodiment of the disclosure, to improve the reliability of the final target hypothesis root cause determination, a historical operation and maintenance event causal reasoning path library is introduced (this library stores verified causal reasoning paths of fault phenomena-related features-root causes, which can be formed through historical operation and maintenance cases and expert experience). The optimal root cause is selected by matching the path similarity. For example, for each hypothetical root cause generated in step S42, the causal reasoning path corresponding to the hypothetical root cause is retrieved from the path library (for example, if the hypothetical root cause is fiber optic interface oxidation, its corresponding historical path is "indicator red flashing + port interrupt text + current noise → fiber optic interface oxidation").
[0091] Next, the similarity between the current set of associated features and the associated feature parts in each causal inference path is calculated (using algorithms such as cosine similarity to calculate the matching degree between local key features and global scene features). The hypothetical root cause corresponding to the causal inference path with the highest similarity is selected as the final root cause of the operation and maintenance event (i.e., the target hypothetical root cause). For example, if the similarity between the current set of associated features and the causal inference path corresponding to fiber optic interface oxidation reaches 92%, which is significantly higher than the similarity between the causal inference paths of other hypothetical root causes, then the root cause of the operation and maintenance event is determined to be fiber optic interface oxidation.
[0092] By reusing verified historical reasoning experience, misjudgments caused by isolated deductions can be avoided, while significantly shortening the root cause determination time and ensuring the practicality and accuracy of the reasoning results.
[0093] In step S44, the inference layer 24 matches the corresponding processing action from a preset operation and maintenance action library based on the built-in root cause-processing action mapping relationship. This operation and maintenance action library can include standard operations for various operation and maintenance scenarios such as hardware repair, configuration adjustment, and environment optimization. After determining the final root cause in step S43, the corresponding processing action is directly matched from the operation and maintenance action library. For example, for the root cause of fiber optic interface oxidation, the matching actions are: cleaning the fiber optic interface with an alcohol swab, detecting interface loss with an optical power meter, and re-plugging and securing the interface; for the root cause of BGP neighbor loss (caused by configuration conflicts), the matching actions are: checking routing policy configuration, modifying conflicting route entries, and restarting the BGP process.
[0094] Finally, the inference layer 24 outputs the root cause of the operation and maintenance event and the corresponding handling actions in a preset format. The output content may include a root cause description (such as oxidation of the fiber optic interface of the G0 / 1 port of the aggregation layer switch, resulting in signal transmission interruption), the confidence level of the root cause, the handling action steps (arranged in execution order), and the action execution priority (such as emergency handling and routine handling). In some embodiments, risk warnings corresponding to the handling actions may also be attached (such as closing the corresponding port when cleaning the interface to avoid service interruption), ensuring that operation and maintenance personnel can directly and quickly perform operations based on the output results and efficiently resolve operation and maintenance events.
[0095] In the exemplary embodiment, when outputting the processing action corresponding to the root cause of the operation and maintenance event, one or more of the following information may also be output: the confidence level of the root cause of the operation and maintenance event, the execution priority of the processing action, the operation step guidance corresponding to the processing action, and the risk warning.
[0096] The pre-trained model 200 of this embodiment can achieve online incremental adaptive updates.
[0097] Figure 5 This is a sub-flowchart of method 100 in an exemplary embodiment of this disclosure.
[0098] refer to Figure 5 In an exemplary embodiment, method 100 further includes: Step S51: Obtain user feedback on the accuracy of the root cause of the operation and maintenance events output by the inference layer. Step S52: When it is determined from multiple accuracy feedbacks that the accuracy of the inference layer output has decreased by more than a preset ratio, freeze the feature extraction parameters, adjust the parameters of the shared projection layer and the inference layer, and perform online incremental adaptive updates of the model.
[0099] Figure 5 In the illustrated embodiment, by receiving user feedback and triggering online incremental adaptive updates, it is possible to ensure that the pre-trained model 200 can continuously adapt to the dynamic changes in the network environment (such as the access of new devices or the occurrence of new faults) and maintain a high accuracy rate in root cause determination over a long period of time.
[0100] In step S51, a feedback interaction entry point can be provided for users performing maintenance operations (such as maintenance engineers) to submit feedback on the accuracy of root cause determination based on the actual processing results. Feedback can take the form of standardized options (e.g., completely accurate, partially accurate, completely inaccurate) or a confidence-based rating (e.g., 0-10 points) to verify whether the root cause output by the inference layer 24 matches the actual fault root cause. For example: if the fault is completely resolved after the user operates according to the root cause of fiber optic interface oxidation and the corresponding processing action, the feedback can be completely accurate; if the fault is not resolved after operating according to the output root cause, and the actual root cause is found to be a damaged port board, the feedback can be completely inaccurate; if the output root cause is an interface abnormality (not clearly oxidized), and the actual root cause is a loose interface, the feedback can be partially accurate. Multiple user feedback data are continuously accumulated to form a feedback dataset.
[0101] In step S52, the feedback dataset is statistically analyzed periodically (e.g., daily / weekly), the actual accuracy of the inference layer output results (i.e., the proportion of feedbacks that are completely accurate to the total number of feedbacks) is calculated, and compared with a preset accuracy threshold (e.g., 90%). When the judgment accuracy drops below a preset proportion (e.g., 5%), the model is triggered to perform online incremental adaptive updates.
[0102] In this embodiment of the disclosure, the online incremental adaptive update process of the pre-trained model 200 includes, but is not limited to, parameter freezing, parameter adjustment, and incremental training.
[0103] To avoid compromising the proven and effective basic feature extraction capabilities, the feature extraction-related parameters involved in step S2 (including image preprocessing and the underlying parameters of various modal feature extraction models) can be frozen, retaining only the adjustable parameters of the shared projection layer 23 and the inference layer 24. Then, based on the accumulated feedback dataset and recently added operational scenario data (such as multimodal features of new faults and operational data of new equipment), the upper-layer parameters of the shared projection layer 23 and the inference layer 24 are fine-tuned. For example, the semantic mapping weights of the shared projection layer 23 are adjusted to make the association judgment of cross-modal features more adaptable to new scenarios; the association rules of fault phenomena and hypothetical root causes and the weight coefficients of causal inference paths in the inference layer 24 are optimized to correct the inference logic biases that previously led to a decrease in accuracy. During parameter adjustment, historical effective data can be mixed into the training at a preset ratio (e.g., 20%) to avoid overfitting to new data and ensure that the model adapts to new scenarios without losing its ability to handle common faults.
[0104] By using online incremental adaptive updates, the real-time service of the pre-trained model 200 can be optimized while in use without long-term interruption. This solves the problem that traditional models are difficult to adapt to dynamic changes after training. Furthermore, by freezing parameters and incremental training, the model can reduce the risk of degradation while ensuring update efficiency, thus ensuring that the model 200 can meet the operation and maintenance needs of complex network environments in the long term.
[0105] In some embodiments, a federated learning framework can be constructed to address the data privacy requirements of multiple operators. Each operator deploys its model locally, retaining only the shared projection layer 23 and inference layer 24. In each iteration, local operational data is used to fine-tune the model, and parameters are uploaded to a central server for updates. The server then aggregates the parameters and distributes them. This approach improves the model's generalization ability across cross-operator scenarios while protecting privacy. Since data from different operators varies, federated learning can leverage data from multiple operators to improve model performance without disclosing the original data, making it suitable for a wider range of network operation and maintenance scenarios.
[0106] During the training of the pre-trained model 200, different training methods can be used for different layers.
[0107] For example, for feature extraction layer 22, text encoders, visual encoders, and audio encoders can be set separately to achieve feature extraction for the corresponding modalities. The audio encoder shares projection layer 23 with the visual and text encoders (input 256+2400+768=3424 dimensions, output 768 dimensions). During training, an audio-text comparison learning task is added (e.g., the similarity target between the audio of "current noise" and the text of "poor port contact" is ≥0.75) to expand the multimodal alignment capability.
[0108] For the training of the shared projection layer 23, the parameters can be initialized using the Xavier method. During training, the contrastive learning loss (InfoNCE) is used to optimize the process, maximizing the cosine similarity (target value ≥ 0.8) between the visual feature of "device indicator light flashing abnormally" and the text description of "port may have poor contact" in the semantic space, and minimizing the similarity of negative sample pairs (target value ≤ 0.3), thereby ensuring the semantic consistency of cross-modal representation.
[0109] The training of inference layer 24 can construct "image-text-processing action" triples and form a chain-like causal reasoning path based on fault samples with root cause annotations and expert feedback. In the specific implementation, fault samples include multimodal inputs (such as an image of a switch port light flashing, a text description of "frequent interruptions of G0 / 1 port"), root cause labels (such as "fiber optic interface oxidation"), and historical maintenance actions (such as "normal operation restored after cleaning the interface"); expert feedback supplements the causal relationships that are not explicitly labeled (such as "high temperature environment easily aggravates interface oxidation").
[0110] Figure 6 This is a diagram illustrating the training process of the inference layer in an exemplary embodiment of this disclosure.
[0111] refer to Figure 6 In this embodiment of the disclosure, the training process of the inference layer 24 includes: Step S61: Based on the operation and maintenance failure samples with root cause annotation and expert feedback, construct multiple sets of related features and the causal reasoning paths corresponding to the related feature sets. The related feature sets include image feature sets, text feature sets, and operation and maintenance events corresponding to both image feature sets and text feature sets. Step S62: A reinforcement learning strategy is adopted, using the inference path annotated by experts as a reward signal, to iteratively optimize the accuracy and interpretability of the root causes of operational events output by the inference layer.
[0112] In this embodiment of the disclosure, the inference layer 24 is an important component of the pre-trained model 200. In order to improve the accuracy and interpretability of the inference results output by the inference layer 24, expert knowledge is used to train the inference layer 24.
[0113] In step S61, a high-quality training dataset is constructed. Based on labeled real-world operation and maintenance data and expert experience, training samples with one-to-one correspondence between related feature sets and causal reasoning paths are formed. A large number of operation and maintenance fault samples with root cause annotations in network operation and maintenance scenarios can be selected, while integrating fault troubleshooting experience feedback from operation and maintenance experts to ensure that the data source is both authentic and professional. Among them, the operation and maintenance fault samples with root cause annotations include multimodal data (corresponding to operation and maintenance images, operation and maintenance text, and some samples include audio data), verified fault root causes and handling experience, and expert feedback supplements the troubleshooting logic of rare and complex faults, improving the sample coverage.
[0114] For each set of fault samples with multimodal data, according to the processing logic of steps S2 and S3 (feature extraction adapted to image type, cross-modal feature projection, similarity association determination), a corresponding associated feature set (such as image feature set, text feature set, and operation and maintenance event identifiers corresponding to both types of features) is generated, which is consistent with the associated feature set format in the real-time processing process to ensure consistency between training and inference.
[0115] Based on root cause annotation of fault samples and expert feedback, a corresponding causal inference path is matched for each set of associated features. This path includes, for example, visualized fault phenomena, generation of hypothetical root causes, verification and screening of root causes, and determination of the final root cause. Each path node corresponds to a multimodal feature in the associated feature set. For example, a set of associated features corresponds to the text "indicator red flashing + port interruption". Its causal inference path includes: visualized fault phenomena (intermittent port disconnection) → generation of hypothetical root causes (interface oxidation / port looseness / configuration anomaly) → feature cross-validation (excluding configuration anomalies) → determination of the final root cause (interface oxidation), ensuring that the inference path is interpretable and traceable. Finally, multiple sets of associated feature sets - standard causal inference path training samples are formed, providing learning information for the iterative optimization of inference layer 24.
[0116] In step S62, a reinforcement learning strategy is used to train the reasoning layer. The standard causal reasoning path annotated by experts is used as a reward signal to guide the reasoning layer to gradually learn the derivation logic from related features to root causes, while strengthening the interpretability of the reasoning process.
[0117] The inference layer can be configured as a reinforcement learning agent, with each set of associated features as environmental input, the causal reasoning path and root cause output by the inference layer as action output, and the standard causal reasoning path annotated by experts as reward signals, thus constructing a complete reinforcement learning training loop. The reward signal's scoring rules correspond to two dimensions: root cause accuracy (higher score if the output root cause matches the standard root cause) and inference path interpretability (higher score if the node matching degree between the output inference path and the standard path is higher), avoiding the problem of the inference layer only pursuing correct root causes while resulting in chaotic reasoning logic. For example, the reward function is designed as follows: 1 point for each step of reasoning matching the expert path, 0.5 points for a partial match, and -0.5 points for a non-match, with the total reward being the sum of the scores from all four steps.
[0118] Next, the training samples constructed in step S61 are input into the inference layer in batches. The inference layer outputs the corresponding causal inference path and root cause based on the initial parameters. The output of the inference layer is compared with the standard path and standard root cause labeled by experts, and a reward score is calculated—the highest reward is given if the root cause is completely accurate and the node matching degree of the inference path is ≥90%; a medium reward is given if the root cause is accurate but the inference path has logical jumps; and a negative reward is given if the root cause is incorrect. In addition, the parameters of the inference layer can be optimized by backpropagation based on the reward score, prioritizing the adjustment of parameters such as the correlation weight between the fault phenomenon and the hypothetical root cause, and the matching coefficient between the hypothetical root cause and the historical path, so that the inference layer can gradually learn the expert's troubleshooting logic. Finally, through multiple rounds of iterative training (the number of iterations is adapted to the sample size adjustment), the accuracy of the root cause output by the inference layer is continuously improved, while the logic and interpretability of the inference path are optimized simultaneously, ensuring that the inference layer can not only correctly identify the root cause, but also clarify the logic.
[0119] During the overall training process of the pre-trained model 200, the optimization algorithm can adopt Proximal Policy Optimization (PPO), with the policy network being a Transformer decoder and the value network being a fully connected network. Through interactive training for 100 rounds (each round containing 50 fault samples), the accuracy and interpretability of the inference sequence generated by the model are optimized, ensuring that it can explicitly output logically coherent operation and maintenance decisions when faced with new faults.
[0120] The pre-trained model 200 supports weakly supervised learning and online incremental adaptation.
[0121] Weakly supervised learning utilizes a large amount of incompletely labeled operational data (such as logs that only describe fault phenomena but lack clear root causes) to mine potential patterns from low-confidence samples through bootstrapping. In the initial stage, a baseline model (a unified cross-modal representation module + 24 inference layers) is trained using 5% labeled data. This model is then used to predict the root causes of unlabeled data. Samples with a confidence level ≥ 0.7 are added to the training set (approximately 30% of the unlabeled data). Subsequently, the model is updated by training with a mixture of new and old data, iterating for 3 rounds (each round adding 20% high-confidence samples). Inconsistent samples (such as samples where the same phenomenon predicts conflicting root causes) are eliminated through cross-validation, reducing the dependence on fully labeled data.
[0122] Online incremental adaptation triggers incremental model training by receiving new operation and maintenance scenario data in real time (such as topology changes after new equipment is connected, new fault cases): when the amount of new data reaches the threshold of 1,000 samples per week or the accuracy of the model on the test set drops by more than 5%, the underlying parameters of the visual encoder and language encoder (the first 3 convolutional layers or Transformer layers) are frozen, and only the upper layer parameters of the shared projection layer and inference layer 24 are fine-tuned (the learning rate is set to 1e-5). The old data is mixed in at a ratio of 20% to participate in the training, and the semantic space mapping parameters and causal inference rules are dynamically updated to ensure that the model continues to adapt to the dynamic evolution of the network environment.
[0123] In summary, in Model 200, Input Layer 21 is responsible for aggregating various operational data, covering four main categories: device data, environmental data, business data, and terminal data. Device data includes, for example, device status data (indicator status, port labels), performance data (CPU utilization, bandwidth usage), and log data (error codes, process status); environmental data includes, for example, data center temperature and humidity, device operating audio (e.g., fan noise), and physical connection status (cable plugging / unplugging); business data includes, for example, business access latency, service availability, and user fault feedback (voice recognition text, written reports); and terminal data includes, for example, terminal device access status and traffic characteristics.
[0124] The feature extraction layer 22 uses an appropriate feature extraction method based on the data type of the input multi-source data, and is divided into a first part (device / environment data processing) and a second part (business / terminal data processing).
[0125] Part 1 (Device / Environment Data): For device status data (such as indicator lights and ports), CNN / ViT is used to extract local and global features; for topology data, GCN is used to extract node-link association features; for performance curve data, TCN is used to extract time-series trend / fluctuation features; for environmental audio data, CNN+LSTM is used to extract time-frequency and time-series dependency features.
[0126] Each data point is processed by its corresponding feature extractor, which outputs a standardized feature vector.
[0127] Part Two (Business / Terminal Data): For business text data (fault feedback), a pre-trained large language model is used to extract contextual semantic features; for terminal access data, structured features such as terminal identifier and traffic characteristics are extracted.
[0128] This ultimately forms a multimodal feature set, preparing for cross-modal association.
[0129] In the shared projection layer 23, the feature sets of all modalities are input into the shared projection layer 23 and uniformly mapped to the same semantic space to eliminate the differences in modal dimensions. Based on the weighted similarity calculation of local key features (such as indicator lights and fault description words) and global scene features (such as equipment layout and business scenarios), it is determined whether different modal features are associated with the same operation and maintenance event. The feature sets that are successfully associated are integrated into a multimodal associated feature set, with an operation and maintenance event identifier attached, and passed to the subsequent inference module.
[0130] Inference layer 24 receives the set of associated features, calls the historical causal reasoning path library, and outputs the root cause of the operation and maintenance event and the handling action through chain causal reasoning. In addition, the pre-trained model 200 also includes a learning feedback module (not shown) and an application interface (not shown).
[0131] The learning feedback module is used to collect user feedback on root cause accuracy. When the accuracy drops below a preset percentage, the feature extraction parameters are frozen, and only the parameters of the shared projection layer and inference layer are fine-tuned to achieve online incremental adaptive updates.
[0132] The application interface outputs the root cause and handling actions to the operation and maintenance management platform in a structured form, and supports functions such as root cause tracing and action execution recording to help operation and maintenance personnel handle faults efficiently.
[0133] Model 200 can cover a variety of operation and maintenance scenarios in practical applications.
[0134] For example, in fault handling scenarios, visual cues of “device abnormal indicator lights” (such as local features of red flashing indicator lights) can be quickly associated with textual descriptions of “BGP neighbor interruption” (such as semantic embedding of “neighbor status frequently down”) through cross-modal alignment. Combined with causal inference chains, hardware faults (such as board damage), connection problems (such as loose fiber optic cables) or configuration errors (such as routing policy conflicts) can be located. In routine maintenance scenarios, based on historical causal reasoning path analysis, potential risks of change operations (such as software upgrades and configuration modifications) are analyzed (e.g., the historical root cause probability of port compatibility issues after upgrades), and security change suggestions are provided (e.g., it is recommended to verify in a test environment first). By mining periodic fault patterns in multimodal data (e.g., link congestion that occurs frequently during high-temperature periods (above 35℃), combined with the volatility analysis of time-series characteristics (e.g., the probability of congestion increases by 15% for every 5℃ increase in temperature), preventive inspection plans are generated (e.g., key inspections of equipment in high-temperature areas are conducted daily from 14:00 to 16:00).
[0135] The above capabilities effectively shorten fault recovery time (from an average of 45 minutes to 12 minutes), reduce repetitive manual troubleshooting (reducing the number of troubleshooting sessions per day by 60%), and significantly reduce network operation and maintenance costs (annual operation and maintenance costs decrease by 35%).
[0136] For example, in a fault handling scenario in a carrier's core data center, the system input includes: 1) Equipment - Data center image (the indicator light on port G0 / 1 of a certain aggregation layer switch is flashing red), after being normalized to 224×224 pixels, is input into ResNet-50, the indicator light area is located through the attention mechanism (local features), and features such as color (red) and flashing frequency (2Hz) are extracted through ROI pooling. The global features are output as a 2048-dimensional vector through average pooling. 2) Topology visualization graph (the link latency between the switch and the adjacent router suddenly increased from 5ms to 80ms), and the 256-dimensional graph structure features were output by processing the node (switcheroo type S7706, CPU utilization 85%) and edge (link bandwidth 10Gbps, current latency 80ms) features through GCN. 3) The performance index curve (the packet loss rate of this port increased from 0.1% to 15% in the past 24 hours) is used to extract the temporal fluctuation features through the 3-layer dilated convolution of TCN (dilation factors 1, 2, and 4), and output a 128-dimensional vector. 4) The operation and maintenance text “G0 / 1 port frequently interrupts, neighbor status Down” is segmented by BERT-base and marked with [CLS]. The 768-dimensional semantic features are output from the [CLS] position.
[0137] In the cross-modal representation unification stage, visual features (2048+256+128=2400 dimensions) and text features (768 dimensions) are unified into 768 dimensions through a linear projection layer. During training, InfoNCE contrastive learning is used. The cosine similarity of positive samples (red flashing feature and "port interrupted" text) reaches 0.85, and the similarity of negative samples (normal green indicator feature and the same text) is 0.22, which meets the semantic consistency requirements.
[0138] In the causal chain reinforcement and fine-tuning stage, the model input is a multimodal embedding (768 dimensions), which is processed by a 6-layer Transformer decoder into a triple of "visual (red flashing) - text (frequent interruptions) - action (historical cleaning record)" to generate an inference chain: "fault phenomenon (G0 / 1 port interruption) → root cause hypothesis (fiber optic interface oxidation) → verification step (using an optical power meter to detect interface loss) → repair instruction (cleaning the interface with alcohol swabs)". PPO optimization is used during training. After 100 rounds (50 samples per round), the inference chain matches the expert path 92%, and the fault recovery time is reduced from 45 minutes to 10 minutes.
[0139] In summary, the multimodal operation and maintenance data processing method and corresponding pre-trained model 200 provided in this disclosure, with a multimodal large model as the core, integrates visual information and operation and maintenance text information to construct cross-modal semantic alignment and causal reasoning capabilities to achieve intelligent network operation and maintenance. A two-stage alignment strategy is adopted to achieve deep fusion of multimodal information.
[0140] The first stage unifies cross-modal representations by designing a shared semantic space for the visual encoder and language encoder, mapping heterogeneous modal data to the same embedding domain. The visual encoder uses differentiated feature extraction methods for different visual input types, while the language encoder performs contextual semantic modeling of the operation and maintenance text based on a pre-trained large language model. The two types of encoders map visual and text features to the same dimensional semantic space through a shared projection layer.
[0141] The second stage involves fine-tuning the causal chain. Based on fault samples with root cause annotations and expert feedback, a "visual-text-action" triplet is constructed to form a chain-like causal reasoning path for training model 200. Model 200 connects the triplets through reasoning layer 24 to form a reasoning chain. The training adopts a reinforcement learning strategy to optimize the accuracy and interpretability of the reasoning sequence generated by the model.
[0142] Furthermore, Model 200 supports weakly supervised learning and online incremental adaptation. Weakly supervised learning utilizes a large amount of incompletely labeled operational data to discover potential patterns through bootstrapping; online incremental adaptation receives new operational scenario data in real time, triggering incremental training of the model and dynamically updating parameters and rules.
[0143] Therefore, Method 100 and Model 200 can cover multiple operation and maintenance scenarios, locate various problems in fault handling scenarios, and provide security change suggestions and preventive inspection plans in routine maintenance scenarios. They can solve the technical problems in existing network operation and maintenance, such as difficulty in fault location, reliance on large amounts of labeled data, difficulty in adapting to dynamic network changes, and high manual operation and maintenance costs. They achieve the technical effects of quickly locating faults, providing security change suggestions and preventive inspection plans, shortening fault recovery time, reducing repetitive manual labor, and significantly reducing network operation and maintenance costs.
[0144] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0145] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0146] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as circuits, modules, or systems.
[0147] The following reference Figure 7 An electronic device 700 according to this embodiment of the present invention is described. Figure 7 The electronic device 700 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0148] like Figure 7 As shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one processor 710, at least one memory 720, and a bus 730 connecting different system components (including memory 720 and processor 710).
[0149] The memory stores program code that can be executed by the processor 710, causing the processor 710 to perform the steps described in the exemplary method section of this specification according to various exemplary embodiments of the present invention. For example, the processor 710 can perform the methods shown in the embodiments of this disclosure.
[0150] The memory 720 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 7201 and / or cache 7202, and may further include read-only memory (ROM) 7203.
[0151] The memory 720 may also include a program / utility 7204 having a set (at least one) of program modules 7205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0152] Bus 730 can represent one or more of several types of bus structures, including a memory bus or memory controller, peripheral bus, graphics acceleration port, processor, or a local bus using any of the various bus structures.
[0153] Electronic device 700 can also communicate with one or more external devices 800 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 700, and / or with any device that enables electronic device 700 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 750. Furthermore, electronic device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 760. As shown, network adapter 760 communicates with other modules of electronic device 700 via bus 730. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 700, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0154] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0155] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the invention may also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the invention described in the exemplary method section of this specification.
[0156] According to embodiments of the present invention, a program product for implementing the above-described method may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium that includes or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0157] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0158] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0159] The program code included on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0160] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0161] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0162] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and concept of this disclosure are indicated by the claims.
Claims
1. A multimodal operation and maintenance data processing method, characterized in that, include: Acquire multimodal operation and maintenance data, which includes operation and maintenance images and operation and maintenance text. The operation and maintenance images include actual photos of the data center, equipment connection topology diagrams, and curves of preset performance indicators. The operation and maintenance text includes speech recognition text and written report text. Image features of the operation and maintenance images are extracted using a feature extraction method corresponding to the type of the operation and maintenance images to form an image feature set. Contextual semantic features of the operation and maintenance text are extracted to form a text feature set. The image feature set and the text feature set are then projected to the same semantic space of a preset dimension using a shared projection layer. In the same semantic space, the feature similarity between the image feature set and the text feature set is determined. When the feature similarity is greater than a preset threshold, the image feature set and the text feature set are determined to be associated with the same operation and maintenance event, forming an associated feature set. The associated feature set includes the image feature set, the text feature set, and the operation and maintenance event. The associated feature set is input into the inference layer, and the inference layer performs chain causal reasoning on the operation and maintenance event to determine the root cause of the operation and maintenance event, and outputs the processing action corresponding to the root cause of the operation and maintenance event.
2. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, Image features of the maintenance images are extracted using a feature extraction method corresponding to the type of the maintenance image to form an image feature set, including: Convolutional neural networks or visual transformers are used to extract local and global scene features from real-life images of the computer room. Graph convolutional networks are used to extract the node-link association features from the device connection topology graph; A temporal convolutional network is used to extract trend and volatility features from the curves of preset performance indicators.
3. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, Image features of the maintenance images are extracted using a feature extraction method corresponding to the type of the maintenance image to form an image feature set, including: After acquiring the operation and maintenance image, the image is first preprocessed by pixel normalization and key area cropping, and then the image features are extracted. The key areas are determined based on the fault association areas of the preset operation and maintenance scenario. The key areas include the device interface area, indicator light area, and topology graph node connection area.
4. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, Determining the feature similarity between the image feature set and the text feature set in the same semantic space includes: Calculate the similarity of local key features between the image feature set and the text feature set. The local key features include the device indicator status, port labels, and error codes in the maintenance image, and the fault description words in the maintenance text. Calculate the global scene feature similarity between the image feature set and the text feature set; The similarity between local key features and the similarity between global scene features are weighted and calculated to determine the feature similarity between the image feature set and the text feature set.
5. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, The step of determining the root cause of the operation and maintenance event through chain causal reasoning in the inference layer and outputting the corresponding processing action for the root cause of the operation and maintenance event includes: Determine the fault phenomenon matching the associated feature set; Identify at least one hypothetical root cause corresponding to the fault phenomenon; The historical operation and maintenance event causal reasoning path library is invoked. By matching the similarity between the associated feature set and the causal reasoning path corresponding to each hypothetical root cause, the target hypothetical root cause with the highest degree of matching with the associated feature set is determined as the root cause of the operation and maintenance event. Determine the root cause of the maintenance event and output the corresponding processing action.
6. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, Also includes: When outputting the processing action corresponding to the root cause of the operation and maintenance event, the system also outputs one or more of the following: the confidence level of the root cause of the operation and maintenance event, the execution priority of the processing action, the operation step guidance corresponding to the processing action, and the risk warning.
7. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, The multimodal mode also includes device operating audio data, and the method further includes: Extract the time-frequency features and time-series dependency features of the audio data to form an audio feature set; The shared projection layer projects the audio feature set, the image feature set, and the text feature set to the same semantic space of a preset dimension. In the same semantic space, the feature similarity of the audio feature set, the image feature set, and the text feature set is determined. When the feature similarity is greater than a preset threshold, the audio feature set, the image feature set, and the text feature set are associated with the same operation and maintenance event, forming an associated feature set. The associated feature set includes the audio feature set, the image feature set, the text feature set, and the operation and maintenance event.
8. The multimodal operation and maintenance data processing method as described in claim 1, characterized in that, Also includes: Obtain user feedback on the accuracy of the root cause of the operation and maintenance event output by the inference layer; When it is determined from multiple accuracy feedbacks that the accuracy of the inference layer output decreases by more than a preset proportion, the feature extraction parameters are frozen, the parameters of the shared projection layer and the inference layer are adjusted, and the model is updated incrementally and adaptively online.
9. The multimodal operation and maintenance data processing method as described in claim 1 or 8, characterized in that, The training process for the inference layer includes: Based on the operation and maintenance failure samples with root cause annotation and expert feedback, multiple sets of associated feature sets and the causal reasoning paths corresponding to the associated feature sets are constructed. The associated feature sets include image feature sets, text feature sets, and operation and maintenance events that are jointly corresponding to the image feature sets and the text feature sets. A reinforcement learning strategy is adopted, using expert-annotated inference paths as reward signals, to iteratively optimize the accuracy and interpretability of the root causes of the operational events output by the inference layer.
10. An electronic device, characterized in that, include: Memory; as well as A processor coupled to the memory, the processor being configured to perform the method as described in any one of claims 1-9 based on instructions stored in the memory.
Citation Information
Patent Citations
Traffic event detection system based on multi-modal data
CN119293551A
Vectorization fault repair method and system based on large model
CN120780516A
Cross-modal semantic alignment driven power grid equipment fault diagnosis method and system
CN121071667A
Visual operation and maintenance method for electric power communication network based on digital twinning
CN121077918A
Unsupervised multi-modal causal structure learning for root cause analysis
US20250062951A1
Cited By
Government inspection method and device, storage medium and electronic equipment
CN122046174A
Government affairs inspection method and device, storage medium and electronic equipment
CN122046174B