Method and system for digitally classifying transmission and transformation project archives based on improved Transform model
By improving the Transformer model and combining external knowledge and feature fusion technology from power transmission and transformation engineering, the efficiency and accuracy issues in digital classification of archives were resolved, achieving efficient and accurate archive classification.
Patent Information
- Application Number
- CN202510823255.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies suffer from inefficiency and insufficient accuracy in the digital classification of power transmission and transformation engineering archives. They are particularly difficult to effectively capture correlations when dealing with long-distance dependencies and specialized fields, and existing methods are not well-suited for specific domains.
An improved Transformer model is adopted, which obtains the original features and external knowledge of power transmission and transformation engineering archives, generates vector representations by using a multi-head attention mechanism, and introduces a learnable low-rank approximation matrix and an adaptive cross-layer residual weighting mechanism. Combined with multi-scale Fourier transform, feature fusion and classification are optimized.
It achieves high-precision and high-efficiency digital classification of archives, and improves the model's ability to understand and generalize complex tasks, especially in the field of power transmission and transformation engineering, in terms of classification accuracy and inference accuracy.
Smart Images

Figure CN120951067A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to archival classification methods, belonging to the field of archival digitization classification technology, and particularly to a method and system for digitizing power transmission and transformation engineering archives based on an improved Transformer model. Background Technology
[0002] During the construction of power transmission and transformation projects, the large volume of archives generated renders traditional manual classification methods inefficient and inaccurate. In recent years, with the rapid development of deep learning technology, its application to archive classification has become a viable solution. However, in the field of power transmission and transformation engineering, research and application of digital archive classification remain relatively limited.
[0003] Existing technologies still have shortcomings in handling text-related tasks. On the one hand, while some methods, such as CNNs and bidirectional LSTMs, have reduced training and inference time to some extent, they have significant limitations in handling long-distance dependencies, especially with long sentences or complex contextual information, making it difficult to effectively capture the relationships within them. On the other hand, even if text classification accuracy can be improved while preserving semantics, it faces the problem of poor applicability in specific domains, such as specialized fields like power transmission and transformation engineering. Therefore, existing technologies are insufficient in both accuracy and efficiency, making it difficult to meet the high-precision and high-efficiency requirements of specific domains for data processing. Summary of the Invention
[0004] The purpose of this invention is to overcome the aforementioned defects and problems in the prior art and to provide a digital classification method and system for power transmission and transformation engineering archives with an improved Transformer model that has higher accuracy and efficiency in a specific field.
[0005] To achieve the above objectives, the technical solution of this invention is: a digital classification method for power transmission and transformation engineering archives based on an improved Transformer model, comprising:
[0006] The original features and external knowledge of power transmission and transformation project archive data are obtained, and vector representations of the original features and external knowledge are generated based on the Transformer multi-head attention mechanism; the external knowledge includes professional terms, equipment parameters and technical standards related to power transmission and transformation projects.
[0007] A learnable low-rank approximation matrix is introduced into the Transformer for training, and the original features and external knowledge are input into the Transformer to calculate attention weights and linearly fuse them to obtain the knowledge-enhanced feature representation.
[0008] The knowledge-enhanced feature representation is input into the MLP layer of the Transformer, and an adaptive cross-layer residual weighting mechanism is introduced into the MLP layer of the Transformer to generate optimized fused features.
[0009] In the KAN layer, the optimized fusion features are subjected to multi-scale Fourier transform to generate multi-scale feature representations at different frequency scales; the multi-scale feature representations are then weighted and fused with the optimized fusion features based on gradient weighting factors to obtain the final feature representation.
[0010] Based on the final feature representation, and according to the classification requirements of power transmission and transformation engineering archives, the classification results are output and corresponding classification labels are generated.
[0011] The process of acquiring the original features and external knowledge of power transmission and transformation engineering archive data, and generating vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism, specifically includes:
[0012] The original characteristics of the power transmission and transformation project archive data are represented as follows: , The dimension of the original feature;
[0013] The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector;
[0014] Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector Its expression is as follows:
[0015] ;
[0016] ;
[0017] ;
[0018] in: The weight matrix of the query vector. The dimension for the query; Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
[0019] The process involves introducing a learnable low-rank approximation matrix into the Transformer for training, and inputting the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain a knowledge-enhanced feature representation. Specifically, this includes:
[0020] Based on query vector With key vector Calculate the attention weights and introduce a learnable low-rank approximation matrix. , Replace the query vector and key vector to optimize the global loss function of the Transformer model. Simultaneously, use the global loss function to guide the training of the low-rank approximation matrix. Its expression is as follows:
[0021] ;
[0022] ; ;
[0023] in: For the global loss function, For classifying losses, To predict the output, For real labels, The Frobenius norm of a low-rank approximation matrix. , To adjust the weight matrix of the low-rank approximation matrix, For regularization hyperparameters;
[0024] Based on low-rank approximation matrix , With weight matrix , The weighted values are used to generate a learnable query matrix and a key matrix, resulting in learnable Transformer attention weights, expressed as follows:
[0025] ;
[0026] in: Scaling factor This is to convert the relevance scores into a probabilistic function;
[0027] Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows:
[0028] ;
[0029] ;
[0030] in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
[0031] The process involves inputting the knowledge-enhanced feature representation into the Transformer's MLP layer, and simultaneously introducing an adaptive cross-layer residual weighting mechanism into the Transformer's MLP layer to generate optimized fused features. Specifically, this includes:
[0032] Calculate the first based on cross-layer attention mechanism Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows:
[0033] ;
[0034] in: For cross-layer attention mechanisms, , The first Layer and first Features of the layer This is the attention weight matrix. The dimension of the feature;
[0035] The first Layer characteristics With the Layer characteristics The contributions of each layer are summed and adjusted using adaptive weighting coefficients and residual adjustment coefficients; simultaneously, the adaptive weighting coefficients for each layer are... and residual adjustment coefficient During backpropagation, adjustments are made based on the task loss function and regularization term. The expression for the loss function of the MLP layer is as follows:
[0036] ;
[0037] in: For the task loss function, , All are regularization terms. , The regularization coefficient is used.
[0038] The optimized fusion feature output by the MLP layer is expressed as follows:
[0039] ;
[0040] in: To optimize fusion features, For adaptive weighting coefficients, This is the residual adjustment coefficient.
[0041] The process involves performing a multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; then, based on a gradient weighting factor, weighting and fusing the multi-scale feature representations with the optimized fusion features to obtain the final feature representation. Specifically, this includes:
[0042] The optimized fusion features are subjected to Fourier transform to extract Fourier features at different frequency scales, as shown in the following expression:
[0043] ;
[0044] in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale;
[0045] The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows:
[0046] ;
[0047] in: Multi-scale feature representation;
[0048] Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows:
[0049] ;
[0050] ;
[0051] in: To optimize fusion features gradient information, Multi-scale feature representation gradient information;
[0052] The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows:
[0053] ;
[0054] in: For the final feature representation, , All of these are adjustable weights. , All are gradient weighting coefficients. This is for element-wise multiplication.
[0055] A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model, the system comprising:
[0056] The vector representation generation module is used to acquire the original features and external knowledge of power transmission and transformation engineering archive data, and generate vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism; the external knowledge includes professional terms, equipment parameters and technical standards related to power transmission and transformation engineering.
[0057] The knowledge enhancement module is used to introduce a learnable low-rank approximation matrix into the Transformer for training, and input the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain the knowledge-enhanced feature representation;
[0058] The feature optimization and fusion module is used to input the knowledge-enhanced feature representation into the MLP layer of the Transformer, and at the same time introduce an adaptive cross-layer residual weighting mechanism into the MLP layer of the Transformer to generate optimized fused features.
[0059] The final feature representation generation module is used to perform multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; the multi-scale feature representations are then weighted and fused with the optimized fusion features based on gradient weighting factors to obtain the final feature representation.
[0060] The classification output module is used to output classification results and generate corresponding classification labels based on the final feature representation and the classification requirements of power transmission and transformation engineering archives.
[0061] The vector representation generation module generates vector representations according to the following method:
[0062] The process of acquiring the original features and external knowledge of power transmission and transformation engineering archive data, and generating vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism, specifically includes:
[0063] The original characteristics of the power transmission and transformation project archive data are represented as follows: , The dimension of the original feature;
[0064] The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector;
[0065] Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector Its expression is as follows:
[0066] ;
[0067] ;
[0068] ;
[0069] in: The weight matrix of the query vector. The dimension for the query; Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
[0070] The knowledge enhancement module obtains the knowledge-enhanced feature representation according to the following method:
[0071] The process involves introducing a learnable low-rank approximation matrix into the Transformer for training, and inputting the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain a knowledge-enhanced feature representation. Specifically, this includes:
[0072] Based on query vector With key vector Calculate the attention weights and introduce a learnable low-rank approximation matrix. , Replace the query vector and key vector to optimize the global loss function of the Transformer model. Simultaneously, use the global loss function to guide the training of the low-rank approximation matrix. Its expression is as follows:
[0073] ;
[0074] ; ;
[0075] in: For the global loss function, For classifying losses, To predict the output, For real labels, The Frobenius norm of a low-rank approximation matrix. , To adjust the weight matrix of the low-rank approximation matrix, For regularization hyperparameters;
[0076] Based on low-rank approximation matrix , With weight matrix , The weighted values are used to generate a learnable query matrix and a key matrix, resulting in learnable Transformer attention weights, expressed as follows:
[0077] ;
[0078] in: Scaling factor This is to convert the relevance scores into a probabilistic function;
[0079] Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows:
[0080] ;
[0081] ;
[0082] in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
[0083] The feature optimization and fusion module generates optimized and fused features according to the following method;
[0084] The process involves inputting the knowledge-enhanced feature representation into the Transformer's MLP layer, and simultaneously introducing an adaptive cross-layer residual weighting mechanism into the Transformer's MLP layer to generate optimized fused features. Specifically, this includes:
[0085] Calculate the first based on cross-layer attention mechanism Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows:
[0086] ;
[0087] in: For cross-layer attention mechanisms, , The first Layer and first Features of the layer This is the attention weight matrix. The dimension of the feature;
[0088] Based on the cross-layer attention mechanism, the first layer... Layer characteristics With the Layer characteristics The contributions of each layer are summed and adjusted using adaptive weighting coefficients and residual adjustment coefficients; simultaneously, the adaptive weighting coefficients for each layer are... and residual adjustment coefficient During backpropagation, adjustments are made based on the task loss function and regularization term. The expression for the loss function of the MLP layer is as follows:
[0089] ;
[0090] in: For the task loss function, , All are regularization terms. , The regularization coefficient is used.
[0091] The optimized fusion feature output by the MLP layer is expressed as follows:
[0092] ;
[0093] in: To optimize fusion features, For adaptive weighting coefficients, This is the residual adjustment coefficient.
[0094] The final feature representation generation module obtains the final feature representation in the following manner:
[0095] The process involves performing a multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; then, based on a gradient weighting factor, weighting and fusing the multi-scale feature representations with the optimized fusion features to obtain the final feature representation. Specifically, this includes:
[0096] The optimized fusion features are subjected to Fourier transform to extract Fourier features at different frequency scales, as shown in the following expression:
[0097] ;
[0098] in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale;
[0099] The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows:
[0100] ;
[0101] in: Multi-scale feature representation;
[0102] Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows:
[0103] ;
[0104] ;
[0105] in: To optimize fusion features gradient information, Multi-scale feature representation gradient information;
[0106] The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows:
[0107] ;
[0108] in: For the final feature representation, , All of these are adjustable weights. , All are gradient weighting coefficients. This is for element-wise multiplication.
[0109] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0110] This invention discloses a digital classification method and system for power transmission and transformation engineering archives based on an improved Transformer model. The method first acquires the original features and external knowledge of the power transmission and transformation engineering archives, and generates vector representations based on the Transformer multi-head attention mechanism. Then, a low-rank approximation matrix is introduced for training, attention weights are calculated and linearly fused to obtain knowledge-enhanced features. These features are then input into an MLP layer and an adaptive cross-layer residual weighting mechanism is introduced to generate optimized fused features. Finally, multi-scale Fourier transforms are performed on the optimized fused features, and they are weighted and fused based on gradient weighting factors to obtain the final feature representation. The classification results are then output and labels are generated. In application, this design, by introducing external knowledge from a specific domain and fusing it with the original data, supplements the model's ability to understand unknown data. Furthermore, the introduction of low-rank approximation, adaptive cross-layer residual weighting, and multi-scale Fourier transform optimizes feature fusion and multi-band information understanding, achieving efficient and accurate digital classification of power transmission and transformation engineering archives with high stability and generalization ability. Attached Figure Description
[0111] Figure 1 This is a flowchart of the method of the present invention.
[0112] Figure 2 This is a diagram of the improved Transformer-MLP-KAN network framework in Embodiment 1 of the present invention.
[0113] Figure 3 This is a schematic diagram showing the comparison results of training losses of various neural network methods in Embodiment 1 of the present invention.
[0114] Figure 4 This is a system structure diagram of the present invention.
[0115] Figure 5 This is a structural diagram of the device of the present invention.
[0116] In the diagram: Vector representation generation module 1, knowledge enhancement module 2, feature optimization and fusion module 3, final feature representation generation module 4, classification output module 5, processor 6, memory 7, computer program code 71. Detailed Implementation
[0117] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0118] Example 1:
[0119] See Figure 1 An improved Transformer model-based digital classification method for power transmission and transformation engineering archives includes:
[0120] The original features and external knowledge of power transmission and transformation engineering archive data are obtained, and vector representations of the original features and external knowledge are generated based on the Transformer multi-head attention mechanism.
[0121] In this solution, to improve the model's performance in the field of power transmission and transformation engineering, external knowledge is incorporated into the Transformer in addition to the original input feature data. When facing unknown data, the introduction of external knowledge helps to compensate for the model's shortcomings in context understanding, enhancing its ability to understand and generalize complex tasks. For domain-specific tasks, external knowledge can help the model better understand technical terms, equipment, and technical standards, thereby improving classification accuracy and inference accuracy. The specific steps are as follows:
[0122] For the original features, the original data of the power transmission and transformation project archives are obtained, and its original features are represented as follows: , The dimension of the original feature;
[0123] External knowledge needs to be transformed into a vector representation, which includes text data and image data. For text data, a vector representation of the external text data can be generated using a SciBERT pre-trained model. For image information, a graph convolutional network (GCN) method can be used to embed graph information into the vector space, and then the text vector and image vector are concatenated to form a vector representation of the external knowledge.
[0124] The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector;
[0125] Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector ;
[0126] Query: The original features are represented by a matrix. Mapped into the query space, a query vector is generated to select relevant knowledge, and its expression is as follows:
[0127] ;
[0128] in: The weight matrix of the query vector. The dimension for the query;
[0129] Key and Value: Each vector in external knowledge Mapped to the key and value spaces, generating key vectors and value vectors respectively, as shown in the following expressions:
[0130] ;
[0131] ;
[0132] in: Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
[0133] The above steps combine the original features with external knowledge to create vector representations that are crucial for model decision-making, enabling the model to find the correlations between features in subsequent attention calculations.
[0134] A learnable low-rank approximation matrix is introduced into the Transformer for training, and the original features and external knowledge are input into the Transformer to calculate attention weights and linearly fuse them to obtain the knowledge-enhanced feature representation.
[0135] Low-rank approximation is an effective method to improve computational efficiency. It reduces the computational load by decreasing the rank of the feature matrix, thus significantly improving computational efficiency. Traditional low-rank approximation methods, such as Singular Value Decomposition (SVD), are static and usually rely on fixed mathematical formulas. This makes them unable to dynamically adjust according to task requirements, limiting their adaptability to complex and changing data patterns. Learnable low-rank approximation, by introducing trainable parameters (such as matrices or weights), allows the low-rank representation to adaptively optimize during training, thereby more flexibly capturing key task-related features. The specific steps are as follows:
[0136] Traditional methods are based on query vectors With key vector The attention weights are calculated to represent the relevance of the input features to each knowledge vector, and their expression is as follows:
[0137] ;
[0138] in: The dot product between the query vector and the key vector is used to obtain the correlation between the original features and each external knowledge. This is a scaling factor used to prevent values from becoming too large; This is to convert the relevance scores into a probabilistic function;
[0139] Introducing a learnable low-rank approximation matrix , Replace the query vector and key vector to optimize the global loss function of the Transformer model, and use the global loss function to guide the training of the low-rank approximation matrix;
[0140] The main goal of learnable low-rank approximations is to obtain a low-rank approximation matrix through learning. and To replace the original high-rank query vector and key vector The core idea of learnable low-rank approximations is to learn low-rank approximate matrices using the gradient descent algorithm. and The dimensions and structure of these low-rank approximation matrices change during training; and during training, this scheme not only optimizes the model's classification loss, but also controls the learning process of the low-rank approximation matrices through regularization terms to prevent overfitting and ensure the effectiveness of the low-rank approximation. and Specifically, this is achieved by optimizing the global loss function. To obtain it, its expression is as follows:
[0141] ;
[0142] ; ;
[0143] in: The global loss function; For classification loss; For predicting output; This is a real label; is the Frobenius norm of the low-rank approximation matrix; , To adjust the weight matrix of the low-rank approximation matrix, the aim is to apply a low-rank matrix to the query vector. and key vector Compression and optimization are performed to improve computational efficiency; For regularization hyperparameters;
[0144] After training, the low-rank approximate matrix learned through training will be... , With weight matrix , The weights are used to generate a learnable query matrix and a key matrix, ultimately yielding learnable Transformer attention weights, expressed as follows:
[0145] ;
[0146] in: Scaling factor This is to convert the relevance scores into a probabilistic function;
[0147] The above steps not only retain the advantages of traditional attention mechanisms, but also reduce the amount of computation through low-rank approximation, and improve computational efficiency through trainable low-rank matrices.
[0148] Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows:
[0149] ;
[0150] ;
[0151] in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
[0152] By connecting the outputs of multiple heads and applying a linear transformation matrix The concatenated representations are fused to map the concatenated representations to the desired output dimension, resulting in a knowledge-enhanced representation. Captured original features and external knowledge It gathers information from multiple perspectives and processes it in parallel through a multi-head attention mechanism. This integrates contextual information from different knowledge sources, enabling the model to better understand and utilize this knowledge.
[0153] The knowledge-enhanced feature representation is input into the MLP layer of the Transformer, and an adaptive cross-layer residual weighting mechanism is introduced into the MLP layer of the Transformer to generate optimized fused features.
[0154] The enhanced feature representation obtained after multi-head attention mechanism The input will be fed into the MLP layer for deeper feature learning. To further optimize feature fusion, an adaptive cross-layer residual weighting mechanism is introduced into the MLP layer. This mechanism dynamically learns the correlation between each layer and automatically adjusts the weighting coefficients and residual contributions of each layer's features. This allows each layer to flexibly adjust the interaction between the current layer and the previous layer according to task requirements during feature fusion. Simultaneously, through a cross-layer attention mechanism, the model can accurately adjust the weights of features in each layer, improving information flow efficiency and enhancing feature fusion performance. The specific steps are as follows:
[0155] Cross-layer attention mechanisms are used to calculate the similarity between layers and to weight the information flow of each layer based on this similarity. The calculation of the first layer's information flow is based on this cross-layer attention mechanism. Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows:
[0156] ;
[0157] in: This is a cross-layer attention mechanism used to calculate the correlation or dependency between different layers; , The first Layer and first Characteristics of the layer; This is the attention weight matrix; The dimension of the feature is used as a scaling factor;
[0158] The first Layer characteristics With the Layer characteristics The contributions of each layer are summed, and the contributions of each layer are adjusted using adaptive weighting coefficients and residual adjustment coefficients, as expressed below:
[0159] ;
[0160] in: To optimize fusion features; These are adaptive weighting coefficients used to control the impact of the current layer's output on the total output of the MLP. Contributions; This is the residual adjustment coefficient, used to control the information flow of the previous layer; the residual information of each layer... This will affect the features of the current layer. By weighted fusion of the features of the current layer and the weighted information of the previous layer, the contribution of each layer in generating the final features can be obtained.
[0161] During training, the adaptive weighting coefficients of each layer and residual adjustment coefficient The loss function will be optimized; these parameters will be adjusted during backpropagation based on the task loss and regularization term to ensure the effectiveness of the weights of each layer in the entire model. The expression for the loss function of the MLP layer is as follows:
[0162] ;
[0163] in: For the task loss function, , All are regularization terms. , This is the regularization coefficient.
[0164] To enhance the understanding of complex, multi-frequency information, this scheme introduces a multi-scale Fourier series to generate a frequency domain representation, and effectively fuses the Fourier representation with the original representation through a gradient weighting factor, thereby improving the multi-dimensional perception capability of information. The specific steps are as follows:
[0165] Optimize fusion features Perform a multi-scale Fourier transform to generate multi-scale feature representations at different frequency scales. Assume that... Different scales (e.g., small scales capture high-frequency information, large scales capture low-frequency information), for each scale ,right The frequency domain representation is obtained by performing a Fourier transform. Its expression is as follows:
[0166] ;
[0167] in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale;
[0168] The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows:
[0169] ;
[0170] in: It is a multi-scale feature representation, which includes Rich information at different frequency scales.
[0171] The multi-scale feature representation and the optimized fusion feature are weighted and fused based on the gradient weighting factor to obtain the final feature representation.
[0172] In this scheme, multi-scale Fourier representation will be used. and optimized fusion features Fusion can help preserve the complementarity of time-domain and frequency-domain information. Traditional weighted fusion methods typically set fixed weighting ratios for features based on experience. However, in reality, the contribution of different features to the task changes during training, and some features may become more important at different stages. By using gradients as weighting factors, the weights of each feature can be dynamically adjusted according to its contribution to the loss function. In this way, features with larger gradients will receive more weight, while features with less impact on the loss will be weakened. This allows the model to automatically adjust feature weights, ensuring that important features contribute more prominently to the final output. The specific steps are as follows:
[0173] Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows:
[0174] ;
[0175] ;
[0176] in: To optimize fusion features gradient information, Multi-scale feature representation The gradient information.
[0177] The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows:
[0178] ;
[0179] in: This is the final feature representation; , All weights are adjustable; , All are gradient weighting coefficients; Element-wise multiplication indicates that the magnitude of the gradient is used as a weighting factor to adjust the importance of the features.
[0180] Based on the final feature representation, and according to the classification requirements of power transmission and transformation engineering archives, the classification results are output and corresponding classification labels are generated.
[0181] like Figure 2 As shown, this scheme improves the network framework of Transformer-MLP-KAN and implements the above steps, so that the final output feature representation contains multi-scale frequency information, enabling the network to analyze the input features at different scales, thereby effectively modeling complex signals.
[0182] like Figure 3As shown in the figure, this embodiment compares the training losses of various neural network methods. Transformer-FR-KAN represents a Fourier transform of the feature variables based on Transformer-KAN. Analysis of the graph shows that the training loss of the Transformer-KAN model starts at around 0.70, gradually decreasing with the number of training epochs, but at a relatively slow rate. By the 50th epoch, the loss stabilizes at approximately 0.44. The initial loss of the Transformer-FR-KAN model is also close to 0.70, but its rate of decrease is much faster than that of Transformer-KAN. By the 10th epoch, the loss is already close to 0.50. As training continues, the loss steadily decreases, reaching approximately 0.38 by the 50th epoch. The improved Transformer-MLP-KAN model proposed in this solution has an initial training loss of approximately 0.71, but its loss decreases the fastest, rapidly dropping below 0.40 within the first 10 epochs and gradually approaching 0.34 in subsequent epochs. Overall, the improved Transformer-MLP-KAN model significantly outperforms Transformer-KAN and Transformer-FR-KAN. It exhibits rapid loss reduction in the early stages of training and reaches its lowest loss value in the later stages.
[0183] Example 2:
[0184] See Figure 4 A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model, comprising:
[0185] Vector representation generation module 1 is used to acquire the original features and external knowledge of power transmission and transformation engineering archive data, and generate vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism; the external knowledge includes professional terms, equipment parameters and technical standards related to power transmission and transformation engineering.
[0186] Furthermore, the vector representation generation module 1 generates the vector representation according to the following method:
[0187] The process of acquiring the original features and external knowledge of power transmission and transformation engineering archive data, and generating vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism, specifically includes:
[0188] The original characteristics of the power transmission and transformation project archive data are represented as follows: , The dimension of the original feature;
[0189] The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector;
[0190] Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector Its expression is as follows:
[0191] ;
[0192] ;
[0193] ;
[0194] in: The weight matrix of the query vector. The dimension for the query; Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
[0195] Knowledge enhancement module 2 is used to introduce a learnable low-rank approximation matrix into the Transformer for training, and input the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain the knowledge-enhanced feature representation;
[0196] Furthermore, the knowledge enhancement module 2 obtains the knowledge-enhanced feature representation in the following manner:
[0197] The process involves introducing a learnable low-rank approximation matrix into the Transformer for training, and inputting the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain a knowledge-enhanced feature representation. Specifically, this includes:
[0198] Based on query vector With key vector Calculate the attention weights and introduce a learnable low-rank approximation matrix. , Replace the query vector and key vector to optimize the global loss function of the Transformer model. Simultaneously, use the global loss function to guide the training of the low-rank approximation matrix. Its expression is as follows:
[0199] ;
[0200] ; ;
[0201] in: For the global loss function, For classifying losses, To predict the output, For real labels, The Frobenius norm of a low-rank approximation matrix. , To adjust the weight matrix of the low-rank approximation matrix, For regularization hyperparameters;
[0202] Based on low-rank approximation matrix , With weight matrix , The weighted values are used to generate a learnable query matrix and a key matrix, resulting in learnable Transformer attention weights, expressed as follows:
[0203] ;
[0204] in: Scaling factor This is to convert the relevance scores into a probabilistic function;
[0205] Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows:
[0206] ;
[0207] ;
[0208] in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
[0209] Feature optimization and fusion module 3 is used to input the knowledge-enhanced feature representation into the MLP layer of Transformer, and at the same time introduce an adaptive cross-layer residual weighting mechanism into the MLP layer of Transformer to generate optimized fused features.
[0210] Furthermore, the feature optimization and fusion module 3 generates optimized and fused features according to the following method;
[0211] The process involves inputting the knowledge-enhanced feature representation into the Transformer's MLP layer, and simultaneously introducing an adaptive cross-layer residual weighting mechanism into the Transformer's MLP layer to generate optimized fused features. Specifically, this includes:
[0212] Calculate the first based on cross-layer attention mechanism Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows:
[0213] ;
[0214] in: For cross-layer attention mechanisms, , The first Layer and first Features of the layer This is the attention weight matrix. The dimension of the feature;
[0215] Based on the cross-layer attention mechanism, the first layer... Layer characteristics With the Layer characteristics The contributions of each layer are summed and adjusted using adaptive weighting coefficients and residual adjustment coefficients; simultaneously, the adaptive weighting coefficients for each layer are... and residual adjustment coefficient During backpropagation, adjustments are made based on the task loss function and regularization term. The expression for the loss function of the MLP layer is as follows:
[0216] ;
[0217] in: For the task loss function, , All are regularization terms. , The regularization coefficient is used.
[0218] The optimized fusion feature output by the MLP layer is expressed as follows:
[0219] ;
[0220] in: To optimize fusion features, For adaptive weighting coefficients, This is the residual adjustment coefficient.
[0221] The final feature representation generation module 4 is used to perform multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; the multi-scale feature representations are then weighted and fused with the optimized fusion features based on the gradient weighting factor to obtain the final feature representation.
[0222] Furthermore, the final feature representation generation module 4 obtains the final feature representation in the following manner:
[0223] The process involves performing a multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; then, based on a gradient weighting factor, weighting and fusing the multi-scale feature representations with the optimized fusion features to obtain the final feature representation. Specifically, this includes:
[0224] The optimized fusion features are subjected to Fourier transform to extract Fourier features at different frequency scales, as shown in the following expression:
[0225] ;
[0226] in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale;
[0227] The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows:
[0228] ;
[0229] in: Multi-scale feature representation;
[0230] Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows:
[0231] ;
[0232] ;
[0233] in: To optimize fusion features gradient information, Multi-scale feature representation gradient information;
[0234] The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows:
[0235] ;
[0236] in: For the final feature representation, , All of these are adjustable weights. , All are gradient weighting coefficients. This is for element-wise multiplication.
[0237] The classification output module 5 is used to output classification results and generate corresponding classification labels based on the final feature representation and the classification requirements of power transmission and transformation engineering archives.
[0238] Example 3:
[0239] See Figure 5 A digital classification device for power transmission and transformation engineering archives based on an improved Transformer model, the device including a processor 6 and a memory 7;
[0240] The memory 7 is used to store computer program code 71 and to transmit the computer program code 71 to the processor 6;
[0241] The processor 6 is used to execute the digital classification method for power transmission and transformation engineering archives based on the improved Transformer model described in Embodiment 1 according to the instructions in the computer program code 71.
[0242] This embodiment also includes a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed on a computer, the digital classification method for power transmission and transformation engineering archives based on the improved Transformer model described in Embodiment 1 is implemented.
[0243] Generally, the computer instructions for implementing the method of the present invention can be carried on any combination of one or more computer-readable storage media. Non-transitory computer-readable storage media can include any computer-readable medium except for the signal itself, which is temporarily propagating.
[0244] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EKROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0245] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smarttalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. In particular, Python, suitable for neural network computation, and platform frameworks such as TensorFlow and PyTorch can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer or to an external computer (e.g., via the Internet using an Internet service provider) through any type of network, including a local area network (LAN) or a wide area network (WAN).
[0246] For details regarding the aforementioned equipment and non-transitory computer-readable storage media, please refer to the specific description of a digital classification method for power transmission and transformation engineering archives based on an improved Transformer model and its beneficial effects, which will not be repeated here.
[0247] Although embodiments of the present invention have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A digital classification method for power transmission and transformation engineering archives based on an improved Transformer model, characterized in that, include: The original features and external knowledge of power transmission and transformation project archive data are obtained, and vector representations of the original features and external knowledge are generated based on the Transformer multi-head attention mechanism; the external knowledge includes professional terms, equipment parameters and technical standards related to power transmission and transformation projects. A learnable low-rank approximation matrix is introduced into the Transformer for training, and the original features and external knowledge are input into the Transformer to calculate attention weights and linearly fuse them to obtain the knowledge-enhanced feature representation. The knowledge-enhanced feature representation is input into the MLP layer of the Transformer, and an adaptive cross-layer residual weighting mechanism is introduced into the MLP layer of the Transformer to generate optimized fused features. In the KAN layer, the optimized fusion features are subjected to multi-scale Fourier transform to generate multi-scale feature representations at different frequency scales; the multi-scale feature representations are then weighted and fused with the optimized fusion features based on gradient weighting factors to obtain the final feature representation. Based on the final feature representation, and according to the classification requirements of power transmission and transformation engineering archives, the classification results are output and corresponding classification labels are generated.
2. The method for digital classification of power transmission and transformation engineering archives based on the improved Transformer model according to claim 1, characterized in that: The process of acquiring the original features and external knowledge of power transmission and transformation engineering archive data, and generating vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism, specifically includes: The original characteristics of the power transmission and transformation project archive data are represented as follows: , The dimension of the original feature; The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector; Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector Its expression is as follows: ; ; ; in: The weight matrix of the query vector. The dimension for the query; Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
3. The method for digital classification of power transmission and transformation engineering archives based on the improved Transformer model according to claim 2, characterized in that: The process involves introducing a learnable low-rank approximation matrix into the Transformer for training, and inputting the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain a knowledge-enhanced feature representation. Specifically, this includes: Based on query vector With key vector Calculate the attention weights and introduce a learnable low-rank approximation matrix. , Replace the query vector and key vector to optimize the global loss function of the Transformer model. Simultaneously, use the global loss function to guide the training of the low-rank approximation matrix. Its expression is as follows: ; ; ; in: For the global loss function, For classifying losses, To predict the output, For real labels, The Frobenius norm of a low-rank approximation matrix. , To adjust the weight matrix of the low-rank approximation matrix, For regularization hyperparameters; Based on low-rank approximation matrix , With weight matrix , The weighted values are used to generate a learnable query matrix and a key matrix, resulting in learnable Transformer attention weights, expressed as follows: ; in: Scaling factor This is to convert the relevance scores into a probabilistic function; Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows: ; ; in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
4. The method for digital classification of power transmission and transformation engineering archives based on the improved Transformer model according to claim 3, characterized in that: The process involves inputting the knowledge-enhanced feature representation into the Transformer's MLP layer, and simultaneously introducing an adaptive cross-layer residual weighting mechanism into the Transformer's MLP layer to generate optimized fused features. Specifically, this includes: Calculate the first based on cross-layer attention mechanism Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows: ; in: For cross-layer attention mechanism, , The first Layer and first Features of the layer This is the attention weight matrix. The dimension of the feature; The first Layer characteristics With the Layer characteristics The contributions of each layer are summed and adjusted using adaptive weighting coefficients and residual adjustment coefficients; simultaneously, the adaptive weighting coefficients for each layer are... and residual adjustment coefficient During backpropagation, adjustments are made based on the task loss function and regularization term. The expression for the loss function of the MLP layer is as follows: ; in: For the task loss function, , All are regularization terms. , The regularization coefficient is used. The optimized fusion feature output by the MLP layer is expressed as follows: ; in: To optimize fusion features, For adaptive weighting coefficients, This is the residual adjustment coefficient.
5. The method for digital classification of power transmission and transformation engineering archives based on the improved Transformer model according to claim 4, characterized in that: The optimized fusion features are subjected to multi-scale Fourier transform in the KAN layer to generate multi-scale feature representations at different frequency scales. The multi-scale feature representation and the optimized fusion feature are weighted and fused based on a gradient weighting factor to obtain the final feature representation, specifically including: The optimized fusion features are subjected to Fourier transform to extract Fourier features at different frequency scales, as shown in the following expression: ; in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale; The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows: ; in: Multi-scale feature representation; Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows: ; ; in: To optimize fusion features gradient information, Multi-scale feature representation gradient information; The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows: ; in: For the final feature representation, , All of these are adjustable weights. , All are gradient weighting coefficients. This is for element-wise multiplication.
6. A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model, characterized in that, The system includes: The vector representation generation module (1) is used to obtain the original features and external knowledge of the power transmission and transformation project archive data, and generate vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism; the external knowledge includes professional terms, equipment parameters and technical standards related to power transmission and transformation projects. The knowledge enhancement module (2) is used to introduce a learnable low-rank approximation matrix into the Transformer for training, and input the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain the feature representation after knowledge enhancement. The feature optimization and fusion module (3) is used to input the knowledge-enhanced feature representation into the MLP layer of Transformer, and at the same time introduce an adaptive cross-layer residual weighting mechanism in the MLP layer of Transformer to generate optimized fusion features. The final feature representation generation module (4) is used to perform multi-scale Fourier transform on the optimized fusion features in the KAN layer to generate multi-scale feature representations at different frequency scales; the multi-scale feature representations are weighted and fused with the optimized fusion features based on the gradient weighting factor to obtain the final feature representation; The classification output module (5) is used to output the classification results and generate corresponding classification labels based on the final feature representation and the classification requirements of the power transmission and transformation project archives.
7. A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model as described in claim 6, characterized in that: The vector representation generation module (1) generates vector representations according to the following method: The process of acquiring the original features and external knowledge of power transmission and transformation engineering archive data, and generating vector representations of the original features and external knowledge based on the Transformer multi-head attention mechanism, specifically includes: The original characteristics of the power transmission and transformation project archive data are represented as follows: , The dimension of the original feature; The external knowledge includes A knowledge vector, denoted as Each knowledge vector is represented as: , The dimension of the external knowledge vector; Generate original features based on Transformer multi-head attention mechanism. and external knowledge query vector Key vector Sum value vector Its expression is as follows: ; ; ; in: The weight matrix of the query vector. The dimension for the query; Let be the weight matrix of the key vectors. The dimension of the key; The weight matrix is the value vector. The dimension of the value.
8. A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model as described in claim 7, characterized in that: The knowledge enhancement module (2) obtains the knowledge-enhanced feature representation in the following manner: The process involves introducing a learnable low-rank approximation matrix into the Transformer for training, and inputting the original features and external knowledge into the Transformer to calculate attention weights and linearly fuse them to obtain a knowledge-enhanced feature representation. Specifically, this includes: Based on query vector With key vector Calculate the attention weights and introduce a learnable low-rank approximation matrix. , Replace the query vector and key vector to optimize the global loss function of the Transformer model. Simultaneously, use the global loss function to guide the training of the low-rank approximation matrix. Its expression is as follows: ; ; ; in: For the global loss function, For classifying losses, To predict the output, For real labels, The Frobenius norm of a low-rank approximation matrix. , To adjust the weight matrix of the low-rank approximation matrix, For regularization hyperparameters; Based on low-rank approximation matrix , With weight matrix , The weighted values are used to generate a learnable query matrix and a key matrix, resulting in learnable Transformer attention weights, expressed as follows: ; in: Scaling factor This is to convert the relevance scores into a probabilistic function; Based on learnable Transformer attention weights, the original features are computed in parallel via multi-head computation. and external knowledge Then, a linear transformation matrix is applied for fusion to obtain the knowledge-enhanced feature representation. Its expression is as follows: ; ; in: This represents the calculation results for each head in a multi-head parallel computation. Let be a linear transformation matrix.
9. A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model as described in claim 8, characterized in that: The feature optimization and fusion module (3) generates optimized and fused features in the following manner; The process involves inputting the knowledge-enhanced feature representation into the Transformer's MLP layer, and simultaneously introducing an adaptive cross-layer residual weighting mechanism into the Transformer's MLP layer to generate optimized fused features. Specifically, this includes: Calculate the first based on cross-layer attention mechanism Layer and first The similarity between layers is used to dynamically adjust the contribution of the attention mechanism in each layer, and its expression is as follows: ; in: For cross-layer attention mechanism, , The first Layer and first Features of the layer This is the attention weight matrix. The dimension of the feature; Based on the cross-layer attention mechanism, the first layer... Layer characteristics With the Layer characteristics The contributions of each layer are summed and adjusted using adaptive weighting coefficients and residual adjustment coefficients; simultaneously, the adaptive weighting coefficients for each layer are... and residual adjustment coefficient During backpropagation, adjustments are made based on the task loss function and regularization term. The expression for the loss function of the MLP layer is as follows: ; in: For the task loss function, , All are regularization terms. , The regularization coefficient is used. The optimized fusion feature output by the MLP layer is expressed as follows: ; in: To optimize fusion features, For adaptive weighting coefficients, This is the residual adjustment coefficient.
10. A digital classification system for power transmission and transformation engineering archives based on an improved Transformer model as described in claim 9, characterized in that: The final feature representation generation module (4) obtains the final feature representation in the following manner: The optimized fusion features are subjected to multi-scale Fourier transform in the KAN layer to generate multi-scale feature representations at different frequency scales. The multi-scale feature representation and the optimized fusion feature are weighted and fused based on a gradient weighting factor to obtain the final feature representation, specifically including: The optimized fusion features are subjected to Fourier transform to extract Fourier features at different frequency scales, as shown in the following expression: ; in: The frequency domain representation of the Fourier features. For Fourier transform, To transform the scale; The Fourier features obtained at each scale are combined to form a multi-scale feature representation matrix, the expression of which is as follows: ; in: Multi-scale feature representation; Computing multi-scale feature representations and optimized fusion features The effect of the loss on the gradient information is expressed as follows: ; ; in: To optimize fusion features gradient information, Multi-scale feature representation gradient information; The weights of each feature are adjusted based on gradient information to obtain the final feature representation, which is expressed as follows: ; in: For the final feature representation, , All of these are adjustable weights. , All are gradient weighting coefficients. This is for element-wise multiplication.