A multi-modal fusion pipe network data digitization archive intelligent management system
Patent Information
- Application Number
- CN202610770461.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]然而,当前技术在处理这些管网数据时面临显著挑战,首先,数据融合度低,各类单模态数据如视觉、文本、音频彼此孤立,缺乏有效的跨模态对齐与融合技术,难以形成描述同一管网事件的统一、高维语义特征表达,导致现场全景信息被割裂,其次,自动化分类与语义标注能力不足,传统方法依赖人工判读,不仅效率低下,而且难以从海量多模态特征中自动发现数据的内在结构关联并生成精准的主题标签,再者,数据存储安全与操作追溯机制薄弱,现有系统多采用集中式存储,存在单点故障与数据被恶意篡改的风险,同时缺乏对数据全生命周期操作行为的可靠、不可篡改的审计追踪手段,最后,数据检索与分析方式单一,未能构建起管网数据间的深层次语义与时空逻辑关联,导致跨数据的关联挖掘与历史追溯过程繁复且效率低下,为了解决这一技术问题,于是我们提供了一种多模态融合的管网数据数字化档案智能管理系统
本发明通过并行专用模型提取多模态特征,并利用跨模态对齐技术与多头注意力机制进行动态加权融合,能够生成高度聚合且语义统一的多模态特征嵌入,实现了对管网现场复杂事件的数字描述,通过结合密度聚类与主题生成模型,能够自动挖掘数据内在结构并提炼精确的语义化主题标签,配合专门训练的双分支神经网络分类器,提升了新数据类别判定与多主题标注的自动化水平与准确性,并将分布式多副本存储与区块链校验机制深度融合,为管网数据生成了存储档案,在保障数据高可用的同时,建立起覆盖数据全生命周期的不可篡改操作日志链,极大增强了数据的安全性与可审计性,最后,通过构建以管网数据为实体的知识图谱,并集成多跳遍历关联检索功能,用户可以挖掘和追溯数据间深层的时空逻辑与语义关联,提升了数据检索、调阅与分析的深度和效率。
Smart Images

Figure CN122594565A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pipeline network data digital management technology, and more specifically, to a multimodal integrated intelligent management system for pipeline network data digital archives. Background Technology
[0002] As urban pipeline networks become increasingly large and complex, their operation and maintenance management generates massive amounts of heterogeneous data from multiple sources, including inspection images, equipment nameplate text photos, and on-site audio recordings.
[0003] However, current technologies face significant challenges in processing this pipeline network data. First, the data integration is low; various single-modal data, such as visual, text, and audio, are isolated from each other, lacking effective cross-modal alignment and fusion technologies. This makes it difficult to form a unified, high-dimensional semantic feature expression describing the same pipeline network event, resulting in fragmented panoramic information. Second, the automated classification and semantic annotation capabilities are insufficient. Traditional methods rely on manual interpretation, which is not only inefficient but also makes it difficult to automatically discover the inherent structural relationships of data from massive multimodal features and generate accurate topic tags. Third, data storage security and operation traceability mechanisms are weak. Existing systems mostly use centralized storage, which poses risks of single points of failure and malicious data tampering. At the same time, there is a lack of reliable and tamper-proof auditing and tracing methods for the entire lifecycle of data operations. Finally, the data retrieval and analysis methods are simplistic, failing to build deep semantic and spatiotemporal logical relationships between pipeline network data. This results in complex and inefficient cross-data association mining and historical tracing processes. To solve this technical problem, we provide a multimodal fusion-based intelligent management system for digital archives of pipeline network data. Summary of the Invention
[0004] The purpose of this invention is to provide a multimodal integrated intelligent management system for digital archives of pipeline network data, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, a multimodal integrated intelligent management system for digital archives of pipeline network data is provided, comprising: The multimodal data recognition unit receives multimodal raw data from the official website and preprocesses the multimodal raw data to obtain image, text and audio data. It extracts features from the image, text and audio data in parallel using a dedicated deep model to generate corresponding high-dimensional feature vectors. It maps each high-dimensional feature vector to the semantic space through cross-modal alignment technology and uses an attention mechanism to assign weights for weighted fusion to output multimodal feature embedding. The intelligent data classification unit groups multimodal feature embeddings using a clustering algorithm to obtain the internal structure. At the same time, it combines a topic generation model to refine the internal structure groups and generate topic labels. The multimodal feature embeddings and topic labels are used together as input features to train the classification model. The classification model is used to determine the category and label the topic of the newly input pipeline network data. The distributed secure storage unit is used to store pipeline data that has been categorized and labeled with a topic. The storage process introduces a blockchain verification mechanism to store pipeline data operation logs on the blockchain and obtain storage files. The interactive control unit provides a visual user interface and integrates knowledge graph-based association retrieval functions to respond to user commands for retrieving, accessing, and analyzing stored files.
[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention extracts multimodal features through a parallel dedicated model and uses cross-modal alignment technology and multi-head attention mechanism for dynamic weighted fusion, generating highly aggregated and semantically unified multimodal feature embeddings. This enables digital description of complex events in pipeline networks. By combining density clustering and topic generation models, it can automatically mine the inherent structure of data and extract precise semantic topic tags. With a specially trained dual-branch neural network classifier, it improves the automation level and accuracy of new data category determination and multi-topic labeling. Furthermore, it deeply integrates distributed multi-replica storage with blockchain verification mechanisms to generate storage archives for pipeline network data. While ensuring high data availability, it establishes an immutable operation log chain covering the entire data lifecycle, greatly enhancing data security and auditability. Finally, by constructing a knowledge graph with pipeline network data as entities and integrating multi-hop traversal association retrieval functions, users can mine and trace deep spatiotemporal logic and semantic relationships between data, improving the depth and efficiency of data retrieval, access, and analysis. Attached Figure Description
[0007] Figure 1 This is an overall block diagram of the present invention.
[0008] The meanings of the labels in the diagram are as follows: 1. Multimodal data recognition unit; 2. Intelligent data classification unit; 3. Distributed secure storage unit; 4. Interactive control unit. Detailed Implementation
[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0010] This invention provides a multimodal fusion intelligent management system for digital archives of pipeline network data. Please refer to [link / reference]. Figure 1 As shown, it includes: The multimodal data recognition unit receives multimodal raw data from the official website and preprocesses the multimodal raw data to obtain image, text and audio data. It extracts features from the image, text and audio data in parallel using a dedicated deep model to generate corresponding high-dimensional feature vectors. It maps each high-dimensional feature vector to the semantic space through cross-modal alignment technology and uses an attention mechanism to assign weights for weighted fusion to output multimodal feature embedding. The intelligent data classification unit groups multimodal feature embeddings using a clustering algorithm to obtain the internal structure. At the same time, it combines a topic generation model to refine the internal structure groups and generate topic labels. The multimodal feature embeddings and topic labels are used together as input features to train the classification model. The classification model is used to determine the category and label the topic of the newly input pipeline network data. The distributed secure storage unit is used to store pipeline data that has been categorized and labeled with a theme. The storage process introduces a blockchain verification mechanism to record the pipeline data operation logs on the blockchain and obtain the storage archive. The interactive control unit provides a visual user interface and integrates knowledge graph-based association retrieval functions to respond to user commands for retrieving, accessing, and analyzing stored files.
[0011] The multimodal raw data includes on-site inspection images of the pipeline network, photos of equipment nameplate text, and on-site audio recordings. The parallel dedicated deep learning models include convolutional neural networks, optical character recognition models, and audio feature extraction networks. Further explanation is needed regarding the expression for calculating the visual feature vector: In the formula Represents visual feature vectors. This represents the feature extraction function of a convolutional neural network. This represents the preprocessed inspection image matrix.
[0012] The optical character recognition (OCR) model receives a pre-processed image of the equipment nameplate text. Pre-processing includes image grayscale conversion, binarization, and tilt correction. The OCR model accurately identifies the text content on the nameplate through a three-stage process: text region detection, character segmentation, and character recognition, outputting the equipment parameter text. The text embedding model receives the equipment parameter text, first performing word segmentation and word vector mapping, and then converting the discrete word vector sequence into a fixed-dimensional text feature vector through multiple fully connected layers. The expression for calculating the text feature vector is as follows: In the formula Represents the text feature vector. This represents the text embedding model transformation function. This represents a text sequence of device parameters. The audio feature extraction network receives preprocessed on-site recordings. Preprocessing includes audio noise reduction, framing, and windowing. The audio feature extraction network converts the time-domain audio signal into frequency-domain features using a short-time Fourier transform, and then extracts key audio features such as Mel-frequency spectrum and power spectrum through convolutional or recurrent layers, outputting a fixed-dimensional audio feature vector. The expression for calculating the audio feature vector is... In the formula Represents the audio feature vector. This represents the audio feature extraction network function. This indicates the pre-processed on-site audio recording signal.
[0013] Cross-modal alignment technology utilizes a semantic mapping layer. Visual, text, and audio feature vectors are input into this layer, which is a three-layer fully connected network. Through linear transformations and non-linear activation functions, it projects single-modal feature vectors with different dimensions and distributions onto the same high-dimensional semantic space, eliminating inter-modal semantic offsets and outputting visual, text, and audio semantic vectors with completely consistent dimensions. The semantic mapping calculation expression is as follows: In the formula This represents the semantic vector of the k-th modality, where k can be img, txt, or aud. This represents the weight matrix for the corresponding mode. This represents the original feature vector of the corresponding mode. This represents the bias vector for the corresponding modality. The attention mechanism receives visual semantic vectors, text semantic vectors, and audio semantic vectors, and calculates the relevance weights between each semantic vector. When using a multi-head attention mechanism, each semantic vector is first mapped to a query vector, key vector, and value vector. A relevance score is calculated using a dot product operation, and then normalized using the Softmax function to obtain the relevance weights of each modality's semantic vectors. The weight calculation expression is as follows: In the formula This represents the relevance weight of the k-th mode. Represents the query vector. This represents the transpose of the k-th modality key vector. m iterates through the three modalities (img, txt, aud) to obtain relevance weights. Then, a weighted summation operation is performed on the visual semantic vector, text semantic vector, and audio semantic vector according to these weights, outputting the final multimodal feature embedding. The weighted summation expression is: In the formula This represents multimodal feature embedding. , , These are the relevance weights for visual, text, and audio modalities, respectively. , , These are visual, text, and audio semantic vectors, respectively. The fusion process dynamically assigns higher weights to highly relevant modalities, effectively integrates multimodal information, generates multimodal feature embeddings, and completes the entire process from single-modal feature extraction to cross-modal fusion.
[0014] After completing the initial projection of single-modal features into the semantic space and generating visual semantic vectors, text semantic vectors, and audio semantic vectors, the semantic mapping layer needs to enter the training and optimization stage. By comparing and learning the loss function, the network parameters are iteratively updated to ensure that the cross-modal semantic alignment effect reaches the preset standard. At the same time, the attention mechanism is switched to multi-head attention mode, and multiple rounds of cross-attention calculation are performed to capture the deep semantic associations between multimodalities, ultimately generating multimodal feature embeddings with stronger representation capabilities.
[0015] The semantic mapping layer is a three-layer fully connected network. The network parameters include the weight matrix and bias vector corresponding to each modality. During the training phase, two types of sample pairs need to be constructed for loss calculation: positive sample pairs, composed of visual, text, and audio feature vectors corresponding to the same pipeline network event; and negative sample pairs, randomly combined from visual, text, and audio feature vectors corresponding to different pipeline network events. The core objective of the contrastive learning loss function is to reduce the vector distance between positive sample pairs in the semantic space while increasing the vector distance between negative sample pairs, thereby constraining the semantic mapping layer to learn common semantic representations among modalities. This loss function adopts a triplet contrastive loss form, and its mathematical expression is: In the formula This represents the contrastive learning loss value. This indicates the number of samples in the training batch. This represents the threshold for the distance between positive and negative samples, used to control semantic discriminative power. The function represents the Euclidean distance calculation for two vectors in the semantic space. , , Let these represent the visual semantic vector, text semantic vector, and audio semantic vector corresponding to the i-th positive sample pair, respectively. , Let represent the audio semantic vector and text semantic vector corresponding to the j-th negative sample pair, respectively.
[0016] During training, gradient descent is used to perform backpropagation. In each round of forward propagation, the contrastive learning loss value of the current batch of samples is calculated. Then, the gradient of the weight matrix and the bias vector is calculated in reverse based on the loss value. The network parameters are updated along the gradient descent direction. This process is iterated until the contrastive learning loss value converges to the preset threshold. At this time, the semantic mapping layer is trained, which can achieve close alignment of semantic vectors of different modalities of the same event and effective differentiation of semantic vectors of different events.
[0017] After completing the training and optimization of the semantic mapping layer, the attention mechanism enables a multi-head attention structure. The multi-head attention mechanism sets up multiple independent attention computing heads in parallel to capture semantic associations of different dimensions and granularities between multimodal semantic vectors, thus avoiding the limitation of single-head attention focusing only on local associations.
[0018] During multi-head attention computation, the visual semantic vector, text semantic vector, and audio semantic vector output from the semantic mapping layer are first concatenated dimensionally to obtain the initial fusion vector. Then, using three independent linear projection matrices, the initial fusion vector is mapped to a query vector, a key vector, and a value vector, respectively. The mathematical expression for linear projection is as follows: , , In the formula Represents the query vector. Represents the key vector. Represents a value vector. , , These represent the query projection matrix, key projection matrix, and value projection matrix, respectively. This represents a concatenated vector of visual semantic vectors, text semantic vectors, and audio semantic vectors.
[0019] Subsequently, the query vector, key vector, and value vector are split into equal-dimensional segments based on the number of attention heads. Each attention head is assigned a set of independent sub-query vectors, sub-key vectors, and sub-value vectors. Each attention head independently performs cross-attention calculation. The calculation process first obtains an initial relevance score by performing a dot product operation between the query sub-vector and the key sub-vector. Then, it is scaled by dividing by the square root of the key vector's dimension to avoid excessively high scores due to high dimensionality. After normalization using the Softmax function, the relevance weights between the semantic vectors of each modality are obtained. Finally, the relevance weights are weighted and summed with the value sub-vectors to generate the intermediate fusion vector corresponding to the current attention head. The mathematical expression for single-head cross-attention calculation is as follows: In the formula This represents the intermediate fusion vector output by the h-th attention head. , , These represent the subquery vector, subkey vector, and subvalue vector corresponding to the h-th attention head, respectively. This represents the total dimension of the semantic vector. This indicates the number of attention heads.
[0020] After all attention heads have completed cross-attention calculations, all intermediate fusion vectors are concatenated in dimensional order to obtain a concatenated fusion vector. The dimension of the concatenated fusion vector is the product of the dimension of a single intermediate fusion vector and the number of attention heads. Finally, a linear transformation network is used to compress the dimension and integrate the features of the concatenated fusion vector. The mathematical expression for the linear transformation is as follows: In the formula This represents the final generated multimodal feature embedding. Represents the linear transformation weight matrix. This represents the concatenated fusion vector obtained by splicing all intermediate fusion vectors. After linear transformation, the output multimodal feature embedding dimension is unified and the semantic representation is comprehensive. It can accurately reflect the multimodal comprehensive features of pipeline network field events and provide reliable input for subsequent clustering analysis of intelligent data classification units.
[0021] After multimodal feature embedding, the intelligent data classification unit immediately starts the clustering algorithm to perform feature space grouping operation, completes similar feature aggregation based on feature distance threshold judgment rules, generates grouping labels that represent the internal structure of pipeline network data, and then calls the topic generation model to generate semantic topic labels based on grouping features, providing structured label data for subsequent classification model training.
[0022] The clustering algorithm employs a density-based clustering method. First, it maps all multimodal feature embeddings to a high-dimensional feature space, where each multimodal feature embedding corresponds to a high-dimensional coordinate point. The feature space distance, referring to the Euclidean distance between the corresponding coordinate points of two feature embeddings, is used to quantify the similarity between features. A preset threshold, a pre-defined distance threshold, serves as the criterion for feature aggregation. The clustering algorithm traverses all multimodal feature embedding coordinate points in the feature space, calculating the feature space distance between the current coordinate point and all other coordinate points point by point. The distance calculation expression is: In the formula This represents the feature space distance between the i-th feature embedding and the j-th feature embedding. This represents the dimension of the multimodal feature embedding. This represents the value of the i-th feature embedded in the k-th dimension. This represents the value of the j-th feature embedded in the k-th dimension. When the feature space distance between two coordinate points is less than a preset threshold, they are judged as similar features and grouped into the same group. During the traversal, similar feature points are continuously merged to form dense feature clusters. At the same time, discrete feature points whose feature space distance is greater than the preset threshold are excluded to avoid abnormal data interfering with the grouping results. After traversing all feature points, each dense feature cluster corresponds to an independent group. A unique number is assigned to each group as a group label. The group label is used to identify the inherent structural category of the pipeline data.
[0023] After completing clustering and generating group labels, the topic generation model receives all multimodal feature embeddings within the same group. First, it calculates the mean vector of all feature embeddings within that group. The expression for calculating the mean vector is as follows: In the formula This represents the feature mean vector of the g-th group. This represents the number of multimodal feature embeddings within the g-th group. Let represent the embedding of the i-th multimodal feature within the g-th group. The mean vector is used as the initial hidden state of the neural network decoder. The neural network decoder employs a recurrent neural network structure to map high-dimensional feature vectors into natural language text sequences. The decoder performs an autoregressive generation operation, calculating the generation probability of each word in the vocabulary at each step based on the current hidden state. The probability calculation expression is: In the formula This indicates that the word is generated in step t. The probability, This represents the decoder's hidden state at step t-1. This represents the output layer weight matrix. The output layer bias vector is used to select the word with the highest probability as the current generated word. The word vector corresponding to this word is input into the decoder to update the hidden state. This process is repeated until a preset end symbol is generated. The generated word sequence is then concatenated into a text phrase, which is the topic label describing the common attributes of the group. The topic generation model simultaneously calculates the confidence score of this topic label. The confidence score calculation expression is: In the formula This represents the confidence score of the topic label for the g-th group. This represents the number of words in the generated text phrases. The generated topic labels are associated and bound with the corresponding group labels and multimodal feature embeddings, serving as the basic data for subsequent classification model training.
[0024] The intelligent data classification unit performs standardized construction of training samples. The training samples consist of three corresponding data categories: multimodal feature embedding, predefined pipeline data category labels, and topic labels. The predefined pipeline data category labels are pre-set discrete category identifiers, covering all preset pipeline data types such as pipeline inspection, equipment failure, normal operation, environmental anomalies, and pipeline leaks. The category labels are converted into numerical vectors using one-hot encoding. One-hot encoding maps a single category to a vector with a dimension equal to the total number of categories, where only the corresponding category position has a value of 1 and the rest have values of 0, ensuring that discrete categories can be processed by the neural network. The topic labels are text phrases output by the topic generation model, which are converted into fixed-dimensional topic feature vectors through a text embedding model. The processed multimodal feature embeddings, category label vectors, and topic feature vectors are sorted in batches. The batch size is set between 8 and 64 based on the hardware computing power. All samples within a batch are randomly shuffled to avoid order deviations during training, forming training batch data that can be directly input into a multilayer neural network classifier.
[0025] The multi-layer neural network classifier employs a multi-branch fully connected network structure, consisting of an input layer, two or more hidden layers, and a dual-branch output layer. The input layer dimension is identical to the multimodal feature embedding dimension, used to receive standardized multimodal feature embedding data. The hidden layers are fully connected, with the number of neurons in each layer set between 256 and 1024. The neuron activation function is the ReLU function, with the function expression as follows: The hidden layer uses weight matrices and bias vectors to transfer data between layers, achieving layer-by-layer abstraction and non-linear expression of features. The output layer is divided into two independent functional branches: the class output branch has the same number of neurons as the total number of predefined network data categories, and uses the Softmax function to output the probability distribution of each predefined category; the topic output branch has the same number of neurons as the dimension of the topic feature vector, and uses the Sigmoid function to output the relevance probability of each topic feature dimension. Supervised training uses a batch gradient descent iterative mechanism. A single training iteration includes four consecutive steps: forward propagation, joint loss calculation, backpropagation gradient solution, and network parameter update. The first step, forward propagation, embeds the multimodal features of each sample in the training batch into the input layer of the input classifier. Data is transferred from the input layer to the hidden layer. The output calculation expression of the l-th hidden layer is: In the formula This represents the output vector of the l-th hidden layer. This represents the weight matrix of the l-th hidden layer. This represents the output vector of the (l-1)th hidden layer. This represents the bias vector of the l-th hidden layer. Data is processed through multiple hidden layers before being passed to the output layer. The class output branch calculates the class probability distribution using the Softmax function, with the following expression: In the formula Represents the category probability distribution vector, This represents the weight matrix of the category output layer. This represents the output vector of the last hidden layer. This represents the bias vector of the category output layer. The topic output branch calculates the topic relevance probability using the Sigmoid function, with the following expression: In the formula This represents a probability vector indicating topic relevance. This represents the weight matrix of the topic output layer. This represents the bias vector of the topic output layer. The second step involves calculating the joint loss using a weighted sum of the class cross-entropy loss and the topic binary cross-entropy loss. The expression for the joint loss is as follows: In the formula This represents the total loss value for a single iteration. This represents the weighting coefficient, with a value ranging from 0 to 1. This represents the cross-entropy loss value. This represents the topic binary cross-entropy loss value. The category cross-entropy loss measures the deviation between the predicted category probability and the true category label, and its calculation expression is: In the formula This indicates the number of samples in a single training batch. This indicates the total number of predefined pipeline network data categories. This represents the one-hot label value of the i-th sample in class c. This represents the predicted probability value of the i-th sample in class c. The topic binary cross-entropy loss is used to measure the deviation between the predicted probability of topic relevance and the true topic label, and its calculation expression is: In the formula The dimension of the topic feature vector. This represents the topic label value of the i-th sample in the t-th dimension. Let represent the predicted probability value of the i-th sample in dimension t. The third step performs backpropagation gradient calculation. Based on the joint loss value, the gradients of all weight matrices and bias vectors in the network are calculated layer by layer from the output layer to the input layer using the chain rule to avoid the gradient vanishing problem. The gradient calculation expression for the hidden layer weight matrix is: [expression here], and the gradient calculation expression for the hidden layer bias vector is: [expression here]. In the formula, each parameter corresponds to a variable in the forward propagation process. After the gradient is solved, the gradient values of all trainable parameters are obtained. The fourth step is to update the network parameters using the Adam optimization algorithm, combining the gradient with historical gradient momentum to update the weight matrix and bias vector. The update expression for the weight matrix is as follows: The update expression for the bias vector is: In the formula This represents the learning rate, with a value ranging from 0.0001 to 0.001. Describing first-order momentum, Indicates second-order momentum. The constant is 10^{-8} to prevent division by zero. The above four-step iterative training is repeated until the joint loss value of the validation set no longer decreases for 20 consecutive rounds. The model is then considered to have converged. All weight matrices, bias vectors and network structure parameters of the classifier are saved to complete the classification model training.
[0026] After the training and deployment of the classification model, it receives the multimodal feature embeddings corresponding to the new input pipeline network data, performs forward propagation calculations in the inference phase, and the inference process is consistent with the forward propagation process in the training phase. The new multimodal feature embeddings are input into the input layer of the classifier, and after feature processing through multiple hidden layers, they are passed to the output layer. The category output branch outputs the probability distribution of each predefined category, and selects the predefined category corresponding to the maximum probability as the category determination result of the new pipeline network data. The topic output branch outputs the relevance probability of each topic feature dimension, and selects topics with a probability greater than a preset threshold of 0.5 as the topic labeling result of the new pipeline network data. The inference process only performs forward propagation calculations and does not update the model parameters. Finally, it outputs structured category determination results and topic labeling results, completing the automated classification and semantic annotation of the new pipeline network data.
[0027] The intelligent data classification unit then initiates a continuous execution process of density-based clustering, sequence generation model topic decoding, and dual-branch prediction of the output layer of a multi-layer neural network classifier. Relying on the unsupervised grouping capability of density clustering, the semantic decoding capability of the sequence generation model, and the multi-task prediction capability of the dual-branch output layer, it achieves accurate grouping of pipeline network data features, semantic generation of topic tags, and automated labeling of categories and multiple topics, providing structured labeled data support for subsequent data storage and retrieval.
[0028] Density-based clustering methods essentially divide feature clusters by identifying differences in the distribution density of feature points in a high-dimensional feature space, rather than relying on a single distance threshold. This effectively adapts to the high-dimensional distribution characteristics of multimodal feature embeddings and automatically filters out abnormal feature points. During clustering, the feature space is first initialized by mapping all multimodal feature embeddings to a d-dimensional continuous feature space. Each multimodal feature embedding corresponds to a d-dimensional coordinate point in this space, with each coordinate component corresponding one-to-one with the numerical values of each dimension of the feature embedding. Two core clustering parameters are then set: the neighborhood radius eps and the minimum number of neighborhood samples minPts. The neighborhood radius eps refers to the radius of the spherical neighborhood space centered on a single feature point, used to define the local density statistical range. The minimum number of neighborhood samples minPts refers to the minimum number of feature points in the neighborhood required to determine a dense region, used to distinguish between dense clusters and sparse regions. After setting the parameters, all unvisited feature points in the feature space are traversed. Using the currently traversed point as the center, the Euclidean distance between the current traversed point and all its neighboring feature points is calculated. The distance calculation expression is: In the formula This represents the spatial distance between the s-th feature point and the t-th feature point. This represents the k-th dimension coordinate value of the s-th feature point. Let represent the k-th dimension coordinate value of the t-th feature point. After traversal, the total number of feature points in the neighborhood is counted. If the total number is greater than or equal to minPts, the current feature point is marked as the core point. The core point is the starting reference point of the dense cluster. Starting from the core point, all reachable feature points in its neighborhood are recursively traversed. All reachable feature points are merged into a dense cluster. Each dense cluster is assigned a unique number as a grouping label. If the total number of feature points in the neighborhood is less than minPts, the current feature point is marked as a boundary point or a discrete point. Boundary points are feature points that are close to the dense cluster but do not meet the density requirements. Discrete points are abnormal feature points that are completely separated from the dense region and are sparsely distributed. After traversing all feature points and completing the dense cluster division, all discrete multimodal feature embedding points are directly removed. Only the feature embeddings and grouping labels corresponding to the dense clusters are retained to avoid abnormal features interfering with the semantic consistency of subsequent topic generation.
[0029] After density-based clustering and removing discrete feature points, the neural network decoder employs a sequence generation model. Using the mean vector of all multimodal feature embeddings within a single group as the initial latent state, it generates a word sequence constituting topic tags through autoregression. During execution, it first calculates the mean vector of all multimodal feature embeddings within a single group. The expression for calculating the mean vector is as follows: In the formula This represents the feature mean vector of the c-th group. This represents the number of multimodal feature embeddings within the c-th group. Let represent the i-th multimodal feature embedding within the c-th group. The mean vector dimension is exactly the same as the multimodal feature embedding dimension. This mean vector is used as the initial hidden state of the sequence generation model. The sequence generation model adopts the Transformer decoder architecture and has a built-in domain-specific vocabulary, a general Chinese vocabulary, and a sequence terminator. <eos>The model comprises a word embedding layer, multiple multi-head self-attention layers, a feedforward network layer, and an output layer. The word embedding layer maps discrete words to fixed-dimensional word vectors. The self-attention layer captures the semantic dependencies between words during generation. The feedforward network layer performs non-linear transformations of features. After the initial hidden state is input into the model, the autoregressive generation process is initiated. The first step calculates the generation probability of all words in the vocabulary through the output layer. The probability calculation expression is: In the formula Indicates the first generated word The probability, This represents the output layer weight matrix. This represents the output layer bias vector. The word with the highest probability is selected as the current generated word. The word vector corresponding to this word is input into the model to update the hidden state. The hidden state update expression is: In the formula This represents the updated implicit state. This represents the computation function of the Transformer decoder layer. Words The word vectors are repeatedly subjected to autoregressive steps of probability calculation, word selection, and latent state update until the end-of-sequence symbol is generated. <eos>The generation process is terminated, and all valid words generated are concatenated in order to form a text phrase. This text phrase serves as the topic label describing the common attributes of the group. Simultaneously, the confidence score of the topic label is calculated. The confidence score calculation expression is: In the formula This represents the confidence level of the topic label in the c-th group. This represents the total number of words in the generated text phrase. This represents the probability of the t-th generated word, used for subsequent filtering of low-confidence topic tags. After topic tag generation, the multi-layer neural network classifier sets both category output nodes and topic output nodes in the output layer, constructing a dual-branch parallel prediction structure. This allows for independent determination of the category probability distribution output and the topic relevance probability. During execution, the output layer node configuration is completed first. The number of category output nodes is consistent with the total number of predefined pipeline data categories, and each node uniquely corresponds to one predefined category. The number of topic output nodes is consistent with the total number of generated valid topic tags, and each node uniquely corresponds to one topic tag. The dual branches share the output vector of the last hidden layer of the classifier. However, each is configured with an independent weight matrix and bias vector. The category output node uses the Softmax activation function to map the hidden layer output vector to the probability distribution of each category. The probability calculation expression is as follows: In the formula Represents the category probability distribution vector, Represents the category branch weight matrix. This represents the category branch bias vector. The sum of all values in the probability distribution vector is 1. The category corresponding to the maximum probability is selected as the category determination result for the new input pipeline data. The topic output node uses the Sigmoid activation function to map the hidden layer output vector to independent correlation probabilities in the interval of 0 to 1. The probability calculation expression for each topic node is as follows: In the formula This represents the relevance probability of the k-th topic node. This represents the weight matrix of the k-th topic node. This represents the bias vector of the k-th topic node. The probabilities of each topic node are independent of each other. A threshold of 0.5 is set. All topic node probabilities are traversed, and topic labels with probabilities greater than 0.5 are selected to form multi-topic annotation results. Finally, the category judgment result and multi-topic annotation result are output simultaneously to complete the automated category classification and multi-topic semantic annotation of pipeline network data, providing complete structured annotation information for subsequent data archiving of distributed secure storage units.
[0030] After density-based clustering and removing discrete multimodal feature embedding points, the intelligent data classification unit uses density-based clustering to remove discrete points and feeds back the grouping results formed by the remaining valid multimodal feature embedding points to the topic generation model in a standardized structured data format. The grouping results contain two core data items: a unique identifier for each group and a set of all multimodal feature embeddings within the corresponding group. The unique identifier for each group is an integer number used to uniquely distinguish different feature clusters. The set of multimodal feature embeddings is a two-dimensional tensor structure, with the number of rows equal to the number of feature embeddings within the group and the number of columns equal to the dimension of the single-modal feature embedding. The feedback process is directly transmitted through the memory data channel without disk read / write, ensuring data transmission efficiency. After receiving the grouping results, the topic generation model starts the topic label generation process for each independent group, without mixing feature data across groups, ensuring that each topic label accurately corresponds to the semantic attributes of a single feature cluster.
[0031] The topic generation model adopts a sequence generation model architecture. For the multimodal feature embedding set within each group, it first calculates the feature mean vector of that group. The expression for calculating the mean vector is as follows: In the formula This represents the feature mean vector of the g-th group. This represents the number of multimodal feature embeddings within the g-th group. Let represent the i-th multimodal feature embedding within the g-th group. This mean vector is used as the initial decoding state of the sequence generation model to initiate the autoregressive word generation process. At each step of the generation process, the generation probability of each candidate word in the vocabulary is output. The probability calculation expression is: In the formula Let represent the probability of generating the k-th candidate word in the t-th generation step. This represents the decoder output layer weight matrix. This indicates the decoding state at step t-1. This represents the decoder output layer bias vector. The word with the highest probability is selected as the current generated word. The decoding state is updated, and the next generation step continues until the end-of-sequence marker is generated, completing the concatenation of the topic label word sequence. Simultaneously, the sequence generation model calculates the confidence score of each topic label. The confidence score is calculated using the geometric mean of the generation probabilities of all words in the generated sequence, expressed as: In the formula This represents the confidence score of the topic label for the g-th group. This represents the total number of words in the g-th topic tag. This represents the maximum generation probability of the selected word in generation step t. The confidence score ranges from 0 to 1; the closer the value is to 1, the higher the semantic reliability of the topic tag. After generation, the topic tag text, confidence score, group unique identifier, and multimodal feature embedding set within the group are bound one by one to form single-group topic generation result data. The intelligent data classification unit has a built-in independent tag filtering module, which is a configurable rule filtering subunit. The preset threshold is a pre-set confidence threshold value used to determine whether the topic tag has effective labeling value. The preset threshold ranges from 0.6 to 0.9, with a default setting of 0.7. After the tag filtering module starts, it iterates through the topic generation result data of all groups, extracts the confidence score corresponding to each topic tag, and performs a comparison between the confidence score and the preset threshold. The comparison logic expression is as follows: In the formula This indicates the validity of the g-th group topic tag; 1 indicates valid retention, and 0 indicates invalid filtering. This indicates a preset confidence threshold. If the confidence score is lower than the preset threshold, the topic label is deemed semantically ambiguous and the labeling reliability is insufficient. All data corresponding to this group, including topic labels, group labels, and multimodal feature embeddings, are directly filtered out to prevent low-quality data from entering the training process. If the confidence score is higher than or equal to the preset threshold, the topic label is deemed semantically clear and the labeling is effective. All associated data of this group are retained.
[0032] Subsequently, after the label filtering module completes the filtering operation, it integrates the associated data of all valid groups. The valid data contains three types of core information that correspond one-to-one: high-confidence topic labels, group labels, and multimodal feature embeddings. The group labels are unique identifiers assigned during the clustering stage to represent the inherent structural categories of the pipeline network data. The multimodal feature embeddings are high-dimensional numerical vectors used to represent the multimodal comprehensive features of the pipeline network data. During the integration process, data consistency checks are performed to ensure that the quantity of the three types of data is completely matched and the dimensional format is uniform. Invalid samples with missing data or abnormal dimensions are removed, and finally a standardized training sample set is formed. The training sample set has a triplet structure, and each sample has the format (multimodal feature embedding, group label, high-confidence topic label). This sample set is directly input into a multi-layer neural network classifier for subsequent supervised training to ensure that the feature-category-topic mapping relationship learned by the classification model is accurate and reliable, laying a high-quality training foundation for the automated classification and high-confidence topic labeling of new pipeline network data.
[0033] After completing the category determination and topic labeling of the pipeline network data, the distributed secure storage unit receives the pipeline network data with completed category determination and topic labeling. This data includes the original pipeline network data ontology, structured category determination results, and multi-topic labeling results. After data reception, the built-in SHA-256 hash algorithm module is invoked to generate a unique data fingerprint, which is used to accurately identify the integrity and uniqueness of the pipeline network data. During the generation process, the binary stream of the pipeline network data ontology, the category determination result string, and the topic labeling result string are first concatenated into a complete data sequence according to a fixed format, and then the SHA-256 hash operation is performed. The data fingerprint calculation expression is as follows: In the formula This represents the generated 256-bit data fingerprint. This represents the SHA-256 hash function. This represents the binary stream of the pipeline data body. This represents the string indicating the category determination result. This represents the string representing the topic annotation result. This indicates a data concatenation operation. This operation has an avalanche effect, meaning that any change in any byte of the pipeline data will completely change the data fingerprint, ensuring the uniqueness of the fingerprint and the ability to verify data integrity.
[0034] After the data fingerprint is generated, the distributed secure storage unit constructs an associated storage tuple. This tuple contains the pipeline data ontology, category determination result, topic labeling result, and data fingerprint. Subsequently, it calls the sharding and replica storage interface of the distributed database. The distributed database is a distributed data storage system composed of multiple independent storage nodes with multi-replica redundancy capabilities, used to improve data storage reliability and access efficiency. Storage allocation uses a consistent hashing algorithm, which minimizes data migration when nodes are added or removed. The node allocation calculation expression is as follows: In the formula Indicates the target storage node index. Represents a hash mapping function. This represents the total number of online available nodes in the distributed database. After the algorithm is executed, at least three different independent storage nodes are selected as target nodes. The associated storage tuple is completely written to each target node. During the writing process, data integrity verification is performed. After the verification is passed, the node storage is marked as successful, forming a multi-replica redundant storage structure to avoid data loss due to single node failure. After the associated storage is completed, the data stored on all nodes is completely consistent, and the corresponding data on any node can be quickly located through data fingerprint.
[0035] While completing distributed multi-node storage, the distributed secure storage unit initiates the blockchain verification and on-chain evidence storage process for pipeline data operation logs through a built-in integrated blockchain light node client. The blockchain light node client is a lightweight client program that does not require synchronizing the complete blockchain ledger, but only synchronizes the block header and core verification data. It can access the blockchain network at low cost and complete data verification and on-chain operations. The client pre-completes the blockchain network initialization configuration, which includes the blockchain network consensus node address, PBFT consensus algorithm parameters, client encryption key pair, and data on-chain permission certificate. After configuration, an encrypted communication connection with the blockchain network is established to complete identity authentication.
[0036] When performing four types of operations—add, modify, delete, and access—on stored pipeline network data, the blockchain light node client captures the operation behavior in real time and initiates the operation log recording and encapsulation process. The operation log record is a structured data unit used to completely record data operation behavior. During encapsulation, six core pieces of information are extracted sequentially: operation type, operation time, operator identity, associated data fingerprint, pipeline network data digest before operation, and pipeline network data digest after operation. The operation type is an enumerated value corresponding to the four types: add, modify, delete, and access. The operation time is a UTC standard timestamp, accurate to milliseconds. The operator identity is a unique user ID or service ID authenticated by the system. The associated data fingerprint is the unique data fingerprint of the currently operated pipeline network data. The pipeline network data digest before operation is the SHA-256 hash value of the pipeline network data body before operation, and the pipeline network data digest after operation is the SHA-256 hash value of the pipeline network data body after operation. The data digest calculation expression is: , In the formula This represents a summary of the data before the operation. This represents the binary stream of the pipeline data body before the operation. This represents a summary of the data after the operation. This represents the binary stream of the pipeline data body after the operation. After extraction, the six pieces of information are concatenated into plaintext log data according to a preset JSON data structure. Then, a digital signature is generated on the plaintext log data using the operator's private key. The digital signature calculation expression is: In the formula Indicates the digital signature of the log. This refers to the elliptic curve digital signature algorithm. This represents the operator's private key. The log is represented as plaintext. After a signature is generated, the plaintext log is merged with the digital signature to form a complete operation log record.
[0037] After the operation log is encapsulated, the blockchain light node client submits the log to the consensus node pool of the preset blockchain network via an encrypted communication link. The blockchain network is a distributed ledger network composed of multiple consensus nodes, which achieves data synchronization through a consensus algorithm and has the characteristic of data immutability. After receiving the log, the consensus node first performs a three-layer validity verification. The first layer verifies the legality of the digital signature by decrypting the signature with the operator's public key and verifying the integrity of the log plaintext. The second layer verifies the operator's identity and permissions, confirming that the operator has the current data operation permissions. The third layer verifies the compliance of the data fingerprint format and digest hash algorithm. After all three layers of verification pass, the consensus node broadcasts the log to... All consensus nodes within the network initiate the PBFT consensus algorithm process. Each consensus node verifies the log records and synchronizes its state. After reaching a consensus, multiple verified operation log records are packaged to generate a new block. The new block consists of a block header and a block body. The block header records the hash value of the previous block, the block generation timestamp, and the Merkle root hash value of the log record. The block body stores the complete set of operation log records. After packaging, the hash value of the new block is calculated and broadcast to the entire network. After all consensus nodes verify the new block, it is appended to the end of the blockchain ledger, completing the on-chain notarization of the operation log records. The log records on the chain cannot be tampered with or deleted, and can be traced and queried across the entire network through data fingerprints. After completing the multi-node data storage of the distributed database and the on-chain notarization of blockchain operation logs, a storage archive corresponding to the pipeline network data is constructed. The storage archive is a complete and secure storage certificate for the pipeline network data, consisting of two parts: the actual data in the distributed database and the operation log chain on the blockchain. The actual data in the distributed database consists of a complete copy of the pipeline network data ontology, category determination results, topic labeling results, and data fingerprints redundantly stored on multiple nodes, ensuring stable access and high availability of the data. The operation log chain on the blockchain is a chronologically linked and tamper-proof sequence of operation log records, completely recording all operations on the pipeline network data from creation, modification, deletion, and access throughout its entire lifecycle. The storage archive is uniquely associated through data fingerprints, allowing simultaneous retrieval of the actual data in the distributed database and the operation log chain on the blockchain via data fingerprints. Ultimately, a complete storage archive is formed, ensuring secure data storage, full traceability of operations, and immutability, providing a secure and reliable data foundation for subsequent data retrieval, access, and analysis by the interactive control unit.
[0038] After the operation log is encapsulated, the hash generation operation of the data fingerprint and the summary of the pipeline network data before and after the operation is first performed. This operation provides a unique identifier and integrity verification basis for blockchain verification, ensuring the authenticity and consistency of the data and log records. Subsequently, the blockchain network receives the log records and initiates the consensus verification, block packaging, and on-chain evidence storage process, ultimately forming an immutable time-series evidence storage chain, completing the full-link secure evidence storage of pipeline network data operation behavior. The data fingerprint is a unique hash identifier for the association information of pipeline network data, category determination results, and topic labeling results. During generation, the three types of information are first converted into standard binary format: pipeline network data is converted into a binary stream of the original data ontology; category determination results are converted into a UTF-8 encoded binary sequence of category text strings; and topic labeling results are converted into a UTF-8 encoded binary sequence of a concatenated string of topic tag sets. Then, the three segments of binary data are concatenated continuously in a fixed order to form a complete binary sequence to be hashed. The SHA-256 hash algorithm is used to perform irreversible hashing operations, generating a 256-bit fixed-length hexadecimal string as the data fingerprint, achieving unique data binding.
[0039] The pipeline data digests before and after the operation correspond to the hash verification values of the pipeline data body before and after the operation, respectively. These are used to verify the integrity and authenticity of the data before and after the operation. When generating the pipeline data digest before the operation, the binary stream of the pipeline data body stored in the distributed database before the operation is directly extracted, and the SHA-256 hash algorithm is used to complete the calculation, generating the data digest before the operation. When generating the pipeline data digest after the operation, the binary stream of the updated pipeline data body after the operation is extracted, and the same SHA-256 hash operation is performed to generate the data digest after the operation. The data digest is a 256-bit hexadecimal string, which allows for precise verification. To determine whether data has been tampered with, the blockchain network consists of multiple consensus nodes capable of data verification, consensus voting, and block storage. Consensus nodes are the core execution units of the blockchain network, responsible for log verification, block packaging, and consensus synchronization. After receiving operation log records submitted by the blockchain light node client, the network entry node first parses the log format, extracting the operation type, operation time, operator identity, data fingerprint, pre-operation data digest, post-operation data digest, and digital signature contained in the log. Then, the parsed log records are broadcast to all online consensus nodes in the network to initiate the log validity verification process.
[0040] The consensus node's verification of the validity of operation log records includes two core components: operator identity verification and data fingerprint consistency verification. Operator identity verification is completed in two steps: digital signature verification and permission verification. The first step extracts the digital signature and operator's public key from the log, decrypts the signature using an elliptic curve digital signature verification algorithm, restores the log plaintext corresponding to the signature, and compares the restored plaintext with the original log plaintext to verify the signature's authenticity. The second step queries the system's permission database to verify whether the operator's identity ID has the necessary operation permissions for the current network data, confirming the permission's validity. If both verifications pass, the operator's identity is deemed legitimate. Data fingerprint consistency verification involves extracting the data fingerprint from the log, calling the distributed database interface to query the data fingerprint in the corresponding network data storage tuple, and comparing the hexadecimal strings of the two data fingerprints to see if they match completely. If they match, then... The data fingerprint is determined to be valid. If either of the two verifications fails, the log record is directly deemed invalid and discarded. If the verification passes, the log record is marked as valid and awaiting on-chain status. After all consensus nodes complete the log validity verification, all valid operation log records are aggregated, and the new block packaging process is initiated. The new block consists of two parts: a block header and a block body. The block header stores key metadata, including the hash value of the previous block, the timestamp of the current block generation, the Merkle root hash value of the valid log record, and the consensus node signature. The hash value of the previous block is used to link historical blocks to form a chain structure. The Merkle root hash value is obtained by constructing a Merkle tree from the hash values of all valid log records and calculating the root hash, which is used to quickly verify the integrity of log records within the block. The block body stores all valid operation log records that have passed verification in sequence. During the packaging process, integrity verification is performed on the block header and block body data. After the verification passes, the original data of the new block is generated.
[0041] After a new block is packaged, the blockchain network uses the Practical Byzantine Fault Tolerance (PBFT) consensus algorithm to perform block consensus and on-chain operations. First, the master consensus node broadcasts the new block data to all slave consensus nodes. After receiving the block data, the slave consensus nodes re-execute log verification and block integrity verification. After the verification passes, they send an agreement voting message to the entire network. When more than two-thirds of the consensus nodes in the network send an agreement voting message, block consensus is achieved. Subsequently, all consensus nodes append the new block to the end of their local blockchain ledger, completing the block on-chain. After the new block is appended, the blockchain ledger forms a blockchain chain linked in chronological order, with each block containing operation log records for the corresponding time period. The resulting evidence storage chain possesses strict temporal order and immutability. Temporal order is reflected in the fact that blocks are linked sequentially according to their generation time, allowing for precise tracing of the entire lifecycle of pipeline data operations from creation to destruction. Immutability is reflected in the fact that any modification to block data will result in a change in the block hash value, disrupting the blockchain-like linking relationship. Furthermore, any modification requires control of more than two-thirds of the consensus nodes, resulting in extremely high technical and cost barriers. The evidence storage chain can quickly retrieve all operation log records of the corresponding pipeline data through data fingerprints, providing immutable and trustworthy credentials for pipeline data security auditing and accountability tracing.
[0042] After completing the distributed secure storage and blockchain-based archiving of pipeline network data, the interactive control unit immediately initiates the entire process of integrating and deploying the visual operation interface and knowledge graph-related retrieval functions, constructing the knowledge graph, responding to related retrieval commands, and visually presenting retrieval results. Relying on visual interaction technology, knowledge graph semantic association technology, and multi-hop traversal retrieval algorithms, it realizes intuitive visual operation and deep related retrieval of pipeline network data, providing users with convenient, efficient, and related data retrieval, access, and analysis services.
[0043] The interactive control unit first integrates the visual operation interface with the knowledge graph-based association retrieval function. The visual operation interface is a user-facing graphical interactive terminal that integrates a retrieval input module, a graph visualization rendering module, a retrieval result display module, a file details access module, and an interactive control button group. It is used to receive user operation commands and display data results. The knowledge graph-based association retrieval function is a backend retrieval engine that relies on the semantic relationships of the knowledge graph to achieve cross-data association queries. The integration process adopts a front-end and back-end separation architecture. The front-end interface integrates the D3.js graph rendering library through JavaScript component development, realizing the zooming, dragging, and highlighting interaction of the knowledge graph graphics. The back-end binds the interface interaction events with the association retrieval engine through RESTful API, establishing an event response chain of search term input, search parameter configuration, search request sending, search result reception, and result rendering. After the user enters a search term and submits the command, the interface automatically triggers a retrieval event, calls the engine to execute the retrieval logic, receives the returned results, and renders them to the interface in real time, achieving seamless integration of the visual interface and the association retrieval function. The interface supports interactive operations such as fuzzy matching of search terms, setting the number of search jumps, filtering of association relationships, paginated display of results, and viewing of stored file details.
[0044] After completing the functional integration, the interaction control unit initiates the knowledge graph construction process. The knowledge graph is a structured semantic network with pipeline network data at its core and semantic relationships as its links. During construction, the two core elements of the graph are first determined: entities and edges. Entities are pipeline network data stored in a distributed secure storage unit. Each piece of pipeline network data corresponds to a unique graph entity. The unique identifier of an entity uses the data fingerprint of the pipeline network data to ensure a one-to-one binding between the entity and the stored file. Entity attributes include key information about the pipeline network data ontology, category determination results, topic annotation results, data collection time, data collection geographical coordinates, equipment number, etc., used to describe the entity's own characteristics. The entity set expression is: In the formula Represents a set of entities. This represents the entity corresponding to the k-th pipeline data. This represents a unique identifier for an entity, i.e., a data fingerprint. This represents the set of entity attributes; edges represent the semantic relationships connecting entities, categorized into four types of directed or undirected edges: The first type is attribute-related edges, connecting entities with their own attributes; the second type is category-related edges, connecting entities with the same category classification; the third type is topic-related edges, connecting entities with overlapping topic annotation results; and the fourth type is spatiotemporal logical related edges, connecting entities whose data collection time difference is less than a preset time threshold (default 1 hour) or whose geographical distance is less than a preset spatial threshold (default 500 meters). The edge set expression is: In the formula Denotes the set of edges. Representing entities and The associated edges between them This indicates the edge type identifier.
[0045] When constructing a knowledge graph, the system first traverses all network data storage files in the distributed database, extracting data fingerprints and attribute information to generate an entity set. Then, it traverses the entity set according to four types of association rules, matching entity pairs that meet the association conditions and generating corresponding edges. After construction, the knowledge graph is stored in a graph database, establishing entity and edge indexes to support efficient retrieval and traversal operations. When the association retrieval function responds to a user's retrieval command, it first performs an initial entity location operation. The user enters search terms in the search input area of the visual interface. Search terms can be data fingerprints, category keywords, topic keywords, device numbers, geographical location information, etc. The user also sets the retrieval hop count (default 1 to 3 hops) and submits the retrieval command. The interface encapsulates the search terms and hop count into a standardized retrieval request and sends it to the association retrieval engine. After receiving the request, the engine first performs word segmentation, synonym matching, and semantic normalization on the search terms, then traverses the knowledge graph entity set, matching entities with unique identifiers and entity attributes containing the semantics of the search terms, generating an initial matching entity set. The expression for the initial matching entity set is: In the formula Represents the initial set of matched entities. This indicates the search terms entered by the user. This represents a semantic matching function used to determine the semantic matching relationship between search terms and entity information.
[0046] After initial entity localization, the association retrieval engine performs a multi-hop traversal along the edges of the knowledge graph. The traversal uses a breadth-first search algorithm to avoid loops and efficiency reductions caused by depth-first traversal. The traversal starts at all entities in the initial matching entity set. Each hop traverses all edges associated with each entity in the current entity set, extracting the associated entities at the other end of the edge, while excluding already traversed entities to prevent duplicate traversals. The expression for a single-hop traversal of the entity set is as follows: In the formula Represents the set of related entities in the nth hop. This represents the (n-1)th hop entity set, where n ranges from 1 to the user-defined hop number. After traversal, all hop-related entities are merged, and after deduplication, a final set of related entities is generated. This set contains all pipeline data entities that have a direct or indirect semantic relationship with the initial matching entity.
[0047] After the associated entity set is generated, the associated retrieval engine sends a storage file retrieval request to the distributed secure storage unit based on the unique identifier, i.e., data fingerprint, of each entity in the set. After receiving the request, the distributed secure storage unit calls the distributed database retrieval interface, accurately locates the storage file corresponding to each entity through the data fingerprint, extracts complete information such as pipeline data ontology, category determination results, topic annotation results, and blockchain operation log chain index from the file, and returns it to the associated retrieval engine. The engine encapsulates the associated entity set, key information of the storage file, and knowledge graph relationship data into retrieval results and sends them to the visual operation interface.
[0048] After receiving the search results, the visual operation interface performs result rendering and display operations. The front-end graph rendering module generates a visual knowledge graph based on the correlation data, marking the initial matching entities as highlighted core nodes. Related entities are displayed hierarchically according to the number of retrieval jumps, and different types of correlation edges are distinguished by different colors and line types, intuitively presenting the correlation between entities. At the same time, the results display area presents the core information of the corresponding storage files of each related entity in the form of a structured list, including data fingerprint, category judgment result, topic labeling result, data collection time, and device number. Users can click on list items to trigger the file details retrieval command. When retrieving, the interface jumps to the details page, displaying the complete content of the pipeline network data ontology, category and topic labeling details. At the same time, the blockchain operation log chain is loaded, generating an operation timeline in chronological order, allowing users to trace the operation records of the entire data lifecycle. Finally, the entire process of correlation retrieval and result visualization is completed, realizing accurate retrieval, correlation mining and intuitive visualization management of pipeline network data.
[0049] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.< / eos> < / eos>
Claims
1. A multimodal fusion intelligent management system for digital archives of pipeline network data, characterized in that, include: The multimodal data recognition unit (1) receives the multimodal raw data from the official website and preprocesses the multimodal raw data to obtain image, text and audio data. It extracts features from the image, text and audio data through parallel dedicated deep models and generates corresponding high-dimensional feature vectors. It maps each high-dimensional feature vector to the semantic space through cross-modal alignment technology and uses attention mechanism to allocate weights for weighted fusion and outputs multimodal feature embedding. The intelligent data classification unit (2) groups the multimodal feature embeddings by clustering algorithm to obtain the internal structure. At the same time, it combines the topic generation model to refine the internal structure groups and generate topic labels. The multimodal feature embeddings and topic labels are used as input features to train the classification model. The classification model is used to determine the category and label the topic of the newly input pipeline network data. The distributed secure storage unit (3) is used to store pipeline data that has been classified and labeled with a theme. The storage process introduces a blockchain verification mechanism to store the pipeline data operation log on the blockchain and obtain the storage file. The interactive control unit (4) provides a visual operation interface and integrates a knowledge graph-based association retrieval function to respond to user commands for retrieving, accessing and analyzing stored files.
2. The intelligent management system for digital archives of pipeline network data based on multimodal fusion as described in claim 1, characterized in that, The multimodal raw data includes on-site inspection images of the pipeline network, photos of equipment nameplate text, and on-site audio recordings; The parallel dedicated deep model includes a convolutional neural network, an optical character recognition model, and an audio feature extraction network; The convolutional neural network extracts visual feature vectors from the inspection images; the optical character recognition model identifies and outputs equipment parameter text from the equipment nameplate text photos; and then converts the equipment parameter text into text feature vectors through the text embedding model; the audio feature extraction network extracts audio feature vectors from the on-site recordings. The cross-modal alignment technique projects visual feature vectors, text feature vectors, and audio feature vectors into the semantic space through a semantic mapping layer, respectively, to obtain visual semantic vectors, text semantic vectors, and audio semantic vectors. The attention mechanism calculates the correlation weights among the visual semantic vector, text semantic vector, and audio semantic vector, and then performs a weighted summation of the visual semantic vector, text semantic vector, and audio semantic vector based on the correlation weights to output a multimodal feature embedding.
3. The intelligent management system for digital archives of pipeline network data based on multimodal fusion according to claim 2, characterized in that, During training, the semantic mapping layer is optimized using a contrastive learning loss function. This loss function reduces the distance between visual, text, and audio feature vectors generated from the same pipeline site event after projection, while increasing the distance between corresponding vectors generated from different pipeline site events. The attention mechanism employs a multi-head attention mechanism, using the visual, text, and audio semantic vectors as query vectors, key vectors, and value vectors, respectively, for multiple rounds of cross-attention calculation. Each round generates a set of weighted intermediate fusion vectors. Finally, all intermediate fusion vectors are concatenated and linearly transformed to generate the final multimodal feature embedding.
4. The intelligent management system for digital archives of pipeline network data based on multimodal fusion as described in claim 1, characterized in that, The clustering algorithm divides the feature space into feature spaces. Multimodal feature embeddings with a distance less than a preset threshold are grouped into the same group, and group labels representing the intrinsic structure of the pipeline network data are obtained. The topic generation model takes all multimodal feature embeddings within the same group as input, generates text phrases describing the common attributes of the group through a neural network decoder, and uses the text phrases as topic labels. The training process of the classification model is as follows: Multimodal feature embeddings with group labels and topic labels are used as training samples and input into a multilayer neural network classifier for supervised training. The multilayer neural network classifier learns the mapping relationship from the multimodal feature embeddings to the predefined pipeline data categories and topic labels. The trained classification model performs forward propagation calculation on the multimodal feature embeddings corresponding to the newly input pipeline data and outputs the category determination result and topic labeling result of the new pipeline data.
5. The intelligent management system for digital archives of pipeline network data based on multimodal fusion as described in claim 4, characterized in that: The clustering algorithm employs a density-based clustering method, which identifies dense regions in the feature space to form groups and excludes discrete multimodal feature embedding points during the clustering process. The neural network decoder uses a sequence generation model with the mean vector of multimodal feature embeddings within the group as the initial hidden state, and regressively generates word sequences that constitute topic labels. The multilayer neural network classifier sets both category output nodes and topic output nodes in the output layer. The category output nodes output the probability distribution of each predefined pipeline data category through the Softmax function, while the topic output nodes independently determine the relevance probability between new input data and each generated topic label through the Sigmoid function, thus completing category determination and multi-topic labeling.
6. The intelligent management system for digital archives of pipeline network data based on multimodal fusion according to claim 5, characterized in that: The density-based clustering method, after excluding discrete multimodal feature embedding points, feeds back the grouping results formed by the remaining multimodal feature embedding points to the topic generation model to generate topic labels. The sequence generation model outputs a confidence score for each generated topic label. The intelligent data classification unit (2) is equipped with a label filtering module. The label filtering module filters out topic labels corresponding to confidence scores lower than a preset threshold, and retains topic labels with confidence scores higher than the preset threshold, along with the corresponding group labels and multimodal feature embeddings, as training samples to train the multilayer neural network classifier.
7. The intelligent management system for digital archives of pipeline network data based on multimodal fusion as described in claim 1, characterized in that: After receiving the pipeline data that has been classified and labeled with a topic, the distributed secure storage unit (3) generates a unique data fingerprint for the pipeline data and stores the pipeline data, classification result, topic labeling result, and data fingerprint together on multiple nodes of the distributed database. The blockchain verification mechanism is as follows: The distributed secure storage unit (3) integrates a blockchain light node client. When adding, modifying, deleting or accessing the pipeline data, the blockchain light node client encapsulates the operation type, operation time, operator identity, associated data fingerprint and pipeline data summary information before and after the operation into an operation log record, and submits the operation log record to the preset blockchain network for consensus verification and on-chain storage. The storage file is composed of the actual data in the distributed database and the operation log chain on the blockchain.
8. The intelligent management system for digital archives of pipeline network data based on multimodal fusion according to claim 7, characterized in that: The data fingerprint is obtained by hashing the network data, category determination results, and topic labeling results. The official website data summary information before and after the operation is obtained by hashing the official website data before and after the operation, respectively. After receiving the operation log record, the blockchain network verifies the validity of the operation log record by multiple consensus nodes in the blockchain network. The verification includes the legitimacy of the operator's identity and the consistency of the data fingerprint. After the verification is passed, the operation log record is packaged into a new block and appended to the end of the blockchain through a consensus algorithm, forming a time-sequential and tamper-proof evidence chain.
9. The intelligent management system for digital archives of pipeline network data based on multimodal fusion according to claim 1, characterized in that: The visual operation interface of the interactive control unit (4) is integrated with the knowledge graph-based association retrieval function. The knowledge graph is constructed with pipeline data stored in the distributed secure storage unit (3) as entities and pipeline data attributes, category determination results, topic labeling results and spatiotemporal logical relationships between pipeline data as edges. When the association retrieval function responds to the user's retrieval command, it first locates the entity matching the search term in the knowledge graph, performs multi-hop traversal along the edges in the knowledge graph, discovers and returns other pipeline data entity sets associated with the initial entity, and the actual storage file corresponding to the entity set is obtained from the distributed database and finally presented to the user on the visual operation interface.
10. The intelligent management system for digital archives of pipeline network data based on multimodal fusion according to claim 1, characterized in that: When responding to the user's command to access and analyze the stored file, the interactive control unit (4) queries all historical operation log records associated with the target stored file and stored on the blockchain through the distributed secure storage unit (3), and generates a timeline in the visual operation interface according to the time sequence of the operation log records. The user can trace the historical process of the target stored file from its creation by operating the timeline.