A Feature Extraction and Fusion Method Based on Three-Way Hybrid Coding and MOE Architecture
Through a three-way hybrid encoding and MOE architecture, combined with shared encoder, private encoder and Transformer expert model, dynamically select feature modeling sub-tasks, solve the problem of feature fusion, improve the flexibility and accuracy of feature extraction, and is suitable for troubleshooting of complex systems.
Patent Information
- Application Number
- CN202510640720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The prior art is difficult to effectively integrate features from different sources or forms, cannot fully capture the context information of the task, and it is difficult to find a balance between computing efficiency and model complexity.
Using three-way hybrid encoding and MOE architecture, common features and private features are extracted through shared encoder and private encoder, combined with Transformer hybrid expert model and gating mechanism, dynamically select feature modeling sub-tasks, and fuse using feedforward neural networks.
It improves the flexibility and accuracy of feature extraction, enhances the generalization ability of the model, reduces the risk of overfitting, is better adaptability and robust, and is suitable for troubleshooting of complex systems.
Smart Images

Figure CN120180057B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of monitoring variable time series classification, and particularly relates to a feature extraction and fusion method based on three-way hybrid coding and MOE architecture. Background Art
[0002] With the continuous innovation of industrial technology and the accelerating development of centralization, feature extraction and fault diagnosis of complex systems are facing increasing challenges. In traditional methods, a single feature representation usually has difficulty in comprehensively reflecting the diversity of complex data, while multi-feature representation effectively improves the robustness of feature expression by fusing features from different sources or in different forms.
[0003] However, relying solely on multi-feature representation may not be able to fully capture the context information of the task. How to reasonably fuse the diversity, how to design a good working mechanism between different architectures, and how to find a balance between computational efficiency and model complexity are also some technical challenges faced. This requires combining a more flexible hybrid architecture method by combining different network architectures to make full use of their respective advantages and further improve the accuracy and efficiency of feature extraction. Therefore, a feature extraction method based on multi-feature representation and architecture hybridization is proposed, aiming to make full use of the advantages of different representation methods and architectures to achieve efficient extraction of complex data features. Summary of the Invention
[0004] In view of the above-mentioned defects of the prior art, the present invention proposes a feature extraction and fusion method based on three-way hybrid coding and MOE architecture. Through multi-feature representation and hybrid architecture, this method can enhance the generalization performance of the model while maintaining the feature expression ability, and show better adaptability and robustness when processing time series data of nonlinear and dynamic systems.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A feature extraction and fusion method based on three-way hybrid coding and MOE architecture, comprising the following steps:
[0007] S1: Obtain the three-axis vibration acceleration monitoring time series data and its N-dimensional feature data, and divide the N-dimensional feature data into N groups of subsequence data by rows;
[0008] S2: Perform the first mapping on the N groups of subsequence data, and extract N groups of common features through a common encoder with the same parameters;
[0009] S3: Perform the second mapping on the N groups of subsequence data, and extract N groups of private features through N private encoders with different parameters;
[0010] S4: Stack the common features column by column to obtain a common feature matrix, stack the private features column by column to obtain a private feature matrix, and stack the common feature matrix and the private feature matrix column by column to obtain a hybrid feature matrix;
[0011] S5: Build a three-way Transformer mixture-of-experts model with three branches; each branch includes an independent expert module and a gating mechanism; the expert module includes multiple Transformer experts;
[0012] S6: Input the common feature matrix, the private feature matrix, and the hybrid feature matrix into the three branches respectively, and use the gating mechanism to control the selection of the Transformer experts to calculate the feature modeling subtasks; the three branches respectively obtain corresponding output features;
[0013] S7: Use a feed-forward neural network to fuse each output feature to obtain a feature result.
[0014] Preferably, in step S2, both the common encoder and the private encoder include a one-dimensional convolutional neural network and an activation function.
[0015] Preferably, in step S5, the gating mechanism is a Top-K noise gating mechanism.
[0016] Preferably, in step S6, use a Gating network to calculate the unnormalized weights of each Transformer expert, and then calculate the output features.
[0017] Preferably, it further includes step S8: Use a KNN model to classify and verify the feature result.
[0018] Compared with the prior art, the beneficial effects of the present invention are reflected in:
[0019] Compared with traditional feature extraction methods, this method uses the mechanism mapping of different feature spaces to extract common features and private features respectively, deeply mines the effective information in the original data, and through a hybrid architecture and a dynamic selection mechanism, improves the flexibility and accuracy of feature extraction, providing scientific and reliable support for fault location in complex systems. The Top-K noise gating mechanism dynamically selects Transformer experts for each branch, accelerates the calculation process, has high flexibility, and allows flexible adjustment in the expert model; the Transformer experts can use the attention mechanism to capture global complex interrelationships, enhance the generalization ability of the model, and reduce the risk of overfitting. Description of the Drawings
[0020] Figure 1It is the flowchart of the method in Embodiment 1 of the present invention;
[0021] Figure 2 It is the mapping process of the shared space and the private space in Embodiment 1 of the present invention;
[0022] Figure 3 It is the multi-way hybrid architecture model diagram in Embodiment 1 of the present invention;
[0023] Figure 4 It is the accuracy confusion matrix diagram of the classification task after feature extraction in Embodiment 2 of the present invention. Detailed implementation manners
[0024] In order to make the technical means, creative features, achieved purposes and effects of the invention easy to understand, the present invention will be further described below in conjunction with specific drawings. However, the present invention is not limited to the following implemented cases.
[0025] It should be noted that the structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical substance significance. Any modification of the structure, change of the proportional relationship or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed by the present invention.
[0026] Embodiment 1:
[0027] Such as Figure 1 shown, a feature extraction and fusion method based on three-way hybrid coding and MOE architecture includes the following steps:
[0028] S1: Obtain the three-axis vibration acceleration monitoring time series data and its N-dimensional feature data, and divide the N-dimensional feature data into N groups of subsequence data by rows.
[0029] For a certain three-axis vibration acceleration monitoring time series monitoring variable (time series data) dimensional data is segmented and extracted, and the data of each dimension is used as an input subsequence. For dimensions, there are
[0030] (1)
[0031] Among them, is the number of dimensions, is the time series length or the number of sampling points corresponding to each dimension.
[0032] After data segmentation, the data of each dimension is extracted as:
[0033] (2)
[0034] S2: Perform the first mapping on the N groups of subsequence data, project them into a new space, and extract N groups of common features through a common encoder with the same parameters.
[0035] like Figure 2 As shown, The first mapping is performed, and a feature representation is performed, through a common encoder Extract common features .Total encoder It is composed of a one-dimensional convolutional neural network and an activation function ReLU, through a shared encoder Map the data to a common feature space, and use the same parameters for the common encoder for each set of inputs and , because the information mapped to this space has a similar distance.
[0036] (3)
[0037] (4)
[0038] in, Yes Perform a one-dimensional convolution operation, is the nonlinear activation function ReLU, is the convolution kernel weight, is the convolution kernel bias.
[0039] S3: Perform a second mapping on the N groups of subsequence data, project them into another new space, and extract N groups of private features through N private encoders with different parameters;
[0040] right Perform the second mapping through a private encoder Extract private features . Private Encoder It is also composed of a one-dimensional convolutional neural network and an activation function ReLU, through a private encoder Mapping the data into a private feature space, the private encoder for each set of inputs uses different parameters because the information mapped into this space is far away.
[0041] (5)
[0042] S4: Stack the common features column - by - column to obtain the common - feature matrix M1, stack the private features column - by - column to obtain the private - feature matrix M2, and stack the common - feature matrix M1 and the private - feature matrix M2 column - by - column to obtain the hybrid - feature matrix M3;
[0043] Stack the obtained group of common features column - by - column to obtain a matrix and stack the obtained group of private features column - by - column to obtain a matrix Then stack the matrices and obtained from the two mapping spaces column - by - column to obtain a new matrix .
[0044] (6)
[0045] (7)
[0046] (8)
[0047] S5: As Figure 3 shown, build a three - path Transformer mixture - of - experts model: The model contains three branches, each branch corresponding to processing different input features and serving as an independent feature - modeling path. Each branch integrates an expert module, which consists of N Transformer experts with the same structure but independent parameters, for realizing high - order feature modeling and representation enhancement; To achieve sparse activation and expert dynamic selection, each branch introduces a Top - K noise gating mechanism, which is used to score all Transformer experts through a Gating network and select the top K experts with the highest scores to participate in the calculation after injecting noise, enhancing the diversity and generalization ability of the model. The principle and control logic of the Top - K noise gating mechanism are prior arts.
[0048] S6: Input the common - feature matrix M1, the private - feature matrix M2, and the hybrid - feature matrix M3 into the three branches of the Transformer mixture - of - experts model respectively. The Top - K noise gating mechanism learns the matching relationship between the input features and the experts through the Gating network, scores all Transformer experts, realizes the adaptive selection of Transformer experts, and dynamically assigns the feature - modeling subtasks to the K most expressive, that is, the top - scoring Transformer experts. Each Transformer expert extracts the deep semantic features of the corresponding input of the branch according to the feature - modeling subtasks it calculates, realizing a more targeted and differentiated feature representation. For the th input , the Gating network calculates the unnormalized weights of each Transformer expert:
[0049] (9)
[0050] (10)
[0051] Where: is the feature extraction of the th expert for the input, is the weight matrix of the Gating network, is the bias vector of the Gating network, The output is the score of each expert, that is, the unnormalized weight score.
[0052] Secondly, inject noise and calculate the score after adding noise:
[0053] (11)
[0054] Where is the hyperparameter of the noise intensity.
[0055] For each sample, select the K experts with the largest score values row by row, and calculate the corresponding unnormalized weights of each expert:
[0056] (12)
[0057] For each sample, the selected set of Transformer experts is , which only processes the input corresponding to the sample, and obtains the output feature of the th path. The attention mechanism of the Transformer model in it enables each vector to be aware of other features, allowing each feature to induce potential information from other features, and obtaining the matrix .
[0058] (13)
[0059] (14)
[0060] Where is the weight of the th expert, is the th expert's feature extraction for the input, are the features extracted by each branch.
[0061] S7: Adaptive fusion using a feedforward neural network Obtain the final feature result .
[0062] First, concatenate along the feature dimension into a matrix :
[0063] (15)
[0064] Obtain the final output through adaptive fusion using a feedforward neural network.
[0065] (16)
[0066] where is the weight matrix, is the bias.
[0067] Example 2:
[0068] A feature extraction and fusion method based on three-way hybrid coding and the MOE architecture, taking the analysis of the three-axis vibration acceleration monitoring data of an elevator car as an example, includes the following steps:
[0069] S1: The three-axis vibration acceleration monitoring data of the elevator car is collected under the normal operation mode (denoted as D1) and six fault modes (denoted as D2 - D7), including guide wheel wear (D2), car top wheel wear (D3), wire rope wear (D4), wire rope breakage (D5), car top wheel bearing wear (D6), and guide wheel bearing wear (D7). For each mode, 500 samples are selected to construct the experimental dataset. The dataset is divided into 70% for training and 30% for testing. First, the dimensional information of the X, Y, and Z axes of the complete data is segmented into three groups of separate sample data .
[0070] S2: Perform the first mapping on through the shared encoder to extract the shared feature . The shared encoder is composed of a one-dimensional convolutional neural network and the activation function ReLU. Through this encoder, the data is mapped into the shared feature space, the shared encoder of and uses the same parameters
[0071] S3: Perform the first mapping on again through three private encoders Private features are extracted The private encoder is also composed of a one-dimensional convolutional neural network and the activation function ReLU. Through this encoder, data is mapped into the private feature space The three private encoders use different hidden layer dimensions because the information mapped into this space is at a relatively large distance
[0072] S4: Stack the three groups of common features obtained by columns to get a matrix Stack the three groups of private features obtained by columns to get a matrix Finally, stack the matrices M1 and M2 obtained from the two mapping spaces by columns to get a new matrix .
[0073] S5: Build a three-branch Transformer mixture-of-experts model. Each branch has its own expert module and gating mechanism. The expert module consists of 3 Transformer experts, and the gating uses the Top-K noise gating mechanism
[0074] S6: Use the three groups of matrix data obtained respectively as the inputs of the three-branch Transformer mixture-of-experts model. Use the Top-K noise gating mechanism to dynamically select the respective Transformer experts for the three branches. Each branch has 3 experts. For K = 2, select the results of the top 2 experts with the largest weights. Each Transformer expert processes the sequence assigned to itself. Calculate the unnormalized weights of the 3 experts for the inputs of the three branches respectively, add noise, select the 2 experts with the largest values, and calculate the corresponding weights. Finally, obtain the feature results corresponding to the three branches .
[0075] S7: Concatenate along the feature dimension to get a matrix and adaptively fuse through a feedforward neural network to obtain the final output features
[0076] S8: Use the KNN model to classify and verify the feature data extracted by the present invention. Five-fold cross-validation is adopted to avoid overfitting of the experimental results. The final results are represented by a confusion matrix, as shown in Figure 4 .
Claims
1. A feature extraction and fusion method based on three-way hybrid coding and MOE architecture, characterized in that It includes the following steps: S1: Obtain the time series data of three-axis vibration acceleration and its N-dimensional feature data, and split the N-dimensional feature data into N groups of subsequence data by rows; S2: Perform the first mapping on the N groups of subsequence data, and extract N groups of common features through a common encoder with the same parameters; S3: Perform the second mapping on the N groups of subsequence data, and extract N groups of private features through N private encoders with different parameters; S4: Stack the common features by columns to obtain a common feature matrix, stack the private features by columns to obtain a private feature matrix, and stack the common feature matrix and the private feature matrix by columns to obtain a mixed feature matrix; S5: Build a three-way Transformer mixture-of-experts model containing three branches; each of the branches includes an independent expert module and a gating mechanism; the expert module includes multiple Transformer experts; S6: Input the common feature matrix, the private feature matrix, and the mixed feature matrix into the three branches respectively, and use the gating mechanism to select the Transformer experts to calculate the feature modeling subtasks; the three branches respectively obtain corresponding output features; S7: Use a feed-forward neural network to fuse each of the output features to obtain a feature result.
2. The feature extraction and fusion method based on three-way hybrid coding and MOE architecture according to claim 1, wherein Both the common encoder and the private encoder include a one-dimensional convolutional neural network and an activation function.
3. The feature extraction and fusion method based on three-way hybrid coding and MOE architecture according to claim 2, wherein In step S5, the gating mechanism is a Top-K noise gating mechanism.
4. The feature extraction and fusion method based on three-way hybrid coding and MOE architecture according to claim 3, characterized in that In step S6, use a Gating network to calculate the unnormalized weights of each of the Transformer experts, and then calculate to obtain the output features.
5. The feature extraction and fusion method based on three-way hybrid coding and MOE architecture according to claim 1, wherein It also includes step S8: Use a KNN model to perform classification verification on the feature result.
Citation Information
Patent Citations
HRRP sequence identification method based on TSF-Transform-LS
CN117540184A
Abnormality detection method based on hybrid expert field adaptive industrial large model
CN120011862A