A multi-modal fusion state monitoring method and system for a hydrostatic spindle
By modeling modal reliability as a latent random variable and performing variational inference, adaptive estimation of modal weights and bilinear interaction are achieved, solving the robustness problem of multimodal monitoring of hydrostatic spindles under complex working conditions and improving the accuracy and stability of monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN UNIV
- Filing Date
- 2026-04-02
- Publication Date
- 2026-06-23
AI Technical Summary
Existing multimodal monitoring methods for hydrostatic spindles lack robustness under complex operating conditions. Single sensors are susceptible to interference from installation status and environment, resulting in insufficient diagnostic robustness. Multimodal fusion methods struggle to characterize the uncertainty of modal changes with operating conditions, leading to misjudgments.
By modeling modal reliability as a latent random variable, using variational inference for adaptive estimation of modal weights, and mining the collaborative features of multi-source information through bilinear interaction, a reliability-aware fusion model is constructed to achieve adaptive adjustment of modal weights and cross-modal collaborative interaction.
It improves the robustness and accuracy of hydrostatic spindle condition monitoring, maintains stable diagnostic performance in modal missing scenarios, and significantly enhances the interpretability and fault detection capability of the system.
Smart Images

Figure CN121959477B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mechanical condition monitoring technology, and more specifically, to a multimodal fusion condition monitoring method and system for a hydrostatic spindle. Background Technology
[0002] As a core functional component in high-end machine tools and precision machining equipment, the performance of hydrostatic spindles directly determines the machining accuracy, stability, and lifespan of the entire machine. In recent years, with the increasing demands for high speed, high rigidity, and high thermal stability in aerospace, optical processing, and ultra-precision manufacturing, the monitoring and data acquisition technology of machine tool spindle operation has become a research hotspot.
[0003] Traditional spindle condition monitoring methods often rely on single sensor signals, such as vibration, acoustic emission, or temperature signals. However, hydrostatic spindles operate under complex conditions and high noise levels for extended periods, making single sensors susceptible to installation variations and environmental interference, resulting in insufficient diagnostic robustness. While multimodal sensing can provide complementary information, existing multimodal fusion methods often assume equal reliability for each mode or employ deterministic gating, making it difficult to characterize the uncertainties of modal changes under varying operating conditions. When the quality of a particular modal signal deteriorates, these fusion methods can introduce low-quality information into the diagnostic results, leading to misdiagnosis.
[0004] Therefore, how to achieve adaptive fusion of multimodal signals under complex working conditions and improve the robustness and accuracy of hydrostatic spindle condition monitoring is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] The present invention aims to solve the above-mentioned problems in the prior art and provides a multimodal fusion state monitoring method and system for a hydrostatic spindle. By modeling modal reliability as a latent random variable and performing variational inference, adaptive estimation of modal weights is achieved, and the collaborative features of multi-source information are fully explored through bilinear interaction.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] According to a first aspect of the present invention, a multimodal fusion state monitoring method for a hydrostatic spindle is provided, comprising the following steps: acquiring multimodal sensing data during the operation of the hydrostatic spindle, wherein the multimodal sensing data includes at least vibration signals and acoustic signals; extracting features from the multimodal sensing data to obtain corresponding modal features; inputting the modal features into a pre-trained reliability-aware fusion model, and performing the following operations: constructing a variational posterior distribution of modal reliability based on each modal feature, and determining a statistical measure of each modal reliability from the variational posterior distribution, wherein the reliability is used to characterize the effective information intensity of the corresponding mode in the current sample; weighting the modal features using the statistical measure to obtain weighted modal features; performing cross-modal collaborative interaction modeling on the weighted modal features to obtain fusion representation features; and outputting the state monitoring result of the hydrostatic spindle based on the fusion representation features.
[0008] According to a second aspect of the present invention, a multimodal fusion state monitoring system for a hydrostatic spindle is provided, comprising: a data acquisition module for acquiring multimodal sensing data during the operation of the hydrostatic spindle, wherein the multimodal sensing data includes at least vibration signals and acoustic signals; a feature extraction module for extracting features from the multimodal sensing data to obtain corresponding modal features; and a fusion monitoring module internally configured with a pre-trained reliability-aware fusion model, wherein the reliability-aware fusion model includes: a reliability estimation unit for constructing a variational posterior distribution of modal reliability based on each modal feature, and determining a statistical measure of each modal reliability from the variational posterior distribution; a weighted fusion unit for weighting the modal features using the statistical measure to obtain weighted modal features; an interactive fusion unit for performing cross-modal collaborative interactive modeling on the weighted modal features to obtain fusion representation features; and an output unit for outputting the state monitoring results of the hydrostatic spindle based on the fusion representation features.
[0009] The beneficial effects achieved by this invention are as follows:
[0010] (1) By modeling modal reliability as a latent random variable and learning its posterior distribution, adaptive estimation of modal weights is achieved, which effectively suppresses the interference of low-quality modalities on the fusion results. It can still maintain stable diagnostic performance in modal missing scenarios, which significantly enhances the robustness and interpretability of the system.
[0011] (2) By fully exploring cross-modal collaborative features through bilinear interaction modeling, the fault discrimination capability is improved. In addition, by analyzing the contribution of cross-modal interaction to the prediction results, its importance in the diagnostic process is revealed. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:
[0013] Figure 1 This is a schematic diagram of the reliability-aware modal weighting mechanism in an embodiment of the present invention;
[0014] Figure 2 This is a schematic diagram of the structure of the multi-expert bilinear interactive fusion module in an embodiment of the present invention;
[0015] Figure 3 This is a schematic diagram of the overall network structure of the method proposed in this embodiment of the invention;
[0016] Figure 4 This is a comparison chart of diagnostic results under modal loss conditions according to an embodiment of the present invention;
[0017] Figure 5 This is a distribution diagram of the decision contribution values of cross-modal interaction under various health states of task T0 in an embodiment of the present invention;
[0018] Figure 6 This is a distribution diagram of the decision contribution values of cross-modal interaction under various health states in task T1, as shown in this embodiment of the invention.
[0019] Figure 7 This is a distribution diagram of the decision contribution values of cross-modal interaction under various health states in task T2, as shown in this embodiment of the invention.
[0020] Figure 8 This is a distribution diagram of the decision contribution values of cross-modal interaction under various health states in task T3, as shown in this embodiment of the invention.
[0021] Figure 9 This is a diagram showing the relationship between modal reliability and interaction contribution in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] Example 1: Reliability-Aware Fusion Method
[0024] This embodiment provides a multimodal fusion state monitoring method for a hydrostatic spindle. Its core lies in constructing a reliability-aware fusion model to achieve adaptive fusion of multimodal signals.
[0025] First, multimodal sensing data is acquired during the operation of the hydrostatic spindle. In specific implementations, the multimodal sensing data includes at least vibration signals and acoustic signals. Vibration signals can be acquired using an accelerometer, reflecting the mechanical vibration characteristics of the spindle; acoustic signals can be acquired using a microphone or acoustic emission sensor, reflecting the acoustic radiation characteristics of the spindle during operation. These two signals reflect the spindle's operating state from different physical perspectives and are naturally complementary.
[0026] Secondly, feature extraction is performed on the multimodal sensing data to obtain the corresponding modal features. Feature extraction can employ a neural network structure consisting of an embedding layer, a self-attention module, and a feedforward network to extract the structural features of the time-series signal and form a high-dimensional vector representation. Let the first... The output of the road signal after feature extraction is ,in For feature dimensions.
[0027] Subsequently, the modal features are input into a pre-trained reliability-aware fusion model. The core operations of this model include three parts: reliability estimation, weighted fusion, and interactive fusion.
[0028] (a) Reliability estimation steps
[0029] Under dynamic operating conditions and strong noise, the effective information content of multimodal signals fluctuates with rotational speed, load, installation status, and environmental interference. If all modes are assumed to be equally reliable during fusion, or if fixed weights are used, interference from low-quality modes can easily be introduced into the fusion result, leading to decreased diagnostic stability. Therefore, this embodiment designs a reliability-aware modal weighting mechanism to adaptively evaluate modal reliability. Figure 1 As shown, this mechanism achieves adaptive estimation of modal reliability by constructing a variational posterior distribution and a hierarchical conditional prior distribution, and constraining the KL divergence between the two.
[0030] Specifically, for the first One modality, introducing a positive reliability variable. This is used to characterize the effective information strength of the modality in the current sample. To ensure that the reliability variable is always positive and to characterize its random fluctuations under different sample conditions, this embodiment uses a log-normal distribution to model it, constructing a variational posterior distribution:
[0031] ;
[0032] in and Lightweight mapping network from input features Adaptive generation is used to characterize the estimation of modal reliability under the current sample conditions; This represents a log-normal distribution. Statistics on the reliability of each mode, such as the expected value and median, are determined from this posterior distribution and used as the basis for subsequent weighting.
[0033] However, data-driven posterior estimations alone are prone to extreme weighting under conditions of strong noise or transient perturbations, thus affecting fusion stability. To limit the fluctuation range of reliability variables, this embodiment designs a hierarchical conditional prior, modeling modal reliability as a combination of global prior and sample conditional bias:
[0034] ;
[0035] in and Let represent the mean and standard deviation of the log-normal prior in the logarithmic space, respectively, which are used to characterize the prior expectation level of modal reliability and its uncertainty range.
[0036] Furthermore, to balance global prior information with differences in sample conditions, the mean parameter and standard deviation are expressed as the sum of a global term and a sample bias term, respectively:
[0037]
[0038] in and This reflects the typical reliability level of the mode across the overall data, while and This is used to describe the reliability shift under specific sample conditions. To ensure the standard deviation is positive, the softplus function is used for transformation:
[0039] ;
[0040] in It is a smooth, non-negative activation function used to map unconstrained real numbers to positive values; This is a preset minimum standard deviation constant, used to avoid numerical instability caused by excessively small standard deviations.
[0041] Based on this, the distance between the posterior distribution and the conditional prior is constrained by KL divergence, thereby limiting extreme fluctuations in reliability while ensuring the model's adaptability. The optimization objective during the training phase can be expressed as:
[0042] ;
[0043] in These are the weighting coefficients for the KL regularization term; This represents an overall training loss. For task-related classification losses; Let represent the number of modes participating in the fusion. Since both the posterior and prior are log-normally distributed, the above KL equation can be directly transformed into a closed-form equation:
[0044] ;
[0045] This closed-form solution avoids additional numerical approximations and significantly improves optimization stability.
[0046] Preferably, to improve computational efficiency and maintain prediction stability, the posterior expected value is used as a statistic during the inference phase:
[0047] ;
[0048] During the training phase, to preserve the randomness in the posterior distribution and more fully characterize the uncertainty of modal reliability, Monte Carlo sampling is used to approximate the posterior expectation:
[0049] ;
[0050] in This represents random noise that follows a standard normal distribution and is used to achieve reparameterized sampling of the posterior distribution.
[0051] After obtaining the modal reliability, the fusion weights are calculated using the following normalization method. :
[0052] ;
[0053] Through the above mechanism, this embodiment achieves adaptive estimation of modal reliability, which can dynamically adjust the contribution of each mode according to the signal quality and effectively suppress the interference of low-quality modes; by constraining extreme weight fluctuations through hierarchical priors, the stability of the fusion process is improved; and the closed-form solution of KL divergence simplifies the training process.
[0054] (II) Weighted fusion steps
[0055] After obtaining the statistics for the reliability of each modality, the modal features are weighted using these statistics to obtain the weighted modal features. Let the statistics determined from the posterior distribution be... Then the weighted modal features It can be represented as:
[0056] ;
[0057] This weighting operation assigns greater contribution weights to high-quality modes through multiplicative modulation, while attenuating the influence of low-quality modes, thus providing more reliable feature inputs for subsequent fusion.
[0058] (III) Interactive Integration Steps
[0059] After obtaining the modal representation with reliability weighting, it is necessary to further establish the synergistic relationships between different modes. Although vibration signals and acoustic signals can reflect the spindle's operating state from different physical perspectives, there are often complex nonlinear coupling relationships between the two types of signals. For example, mechanical impact usually manifests as a transient pulse in vibration signals, while in acoustic signals it may manifest as spectral changes after energy propagation. Therefore, relying solely on linear fusion methods, such as feature splicing or weighted summation, can usually only characterize the complementary information between modes, but it is difficult to reflect the multiplicative coupling relationships between different modal features, thus limiting the expressive power of the fusion representation.
[0060] Therefore, this embodiment constructs as follows Figure 2 The multi-expert bilinear interactive fusion module shown performs cross-modal collaborative interactive modeling on the weighted modal features to obtain fused representation features. Specifically, it includes the following sub-steps:
[0061] First, the weighted modal features are linearly fused to obtain a first-order fusion term for preserving the linear complementarity information of the modes. Let the weighted vibration modes and acoustic modes be represented as follows: Then the first-order fusion term It can be represented as:
[0062] ;
[0063] in This represents vector concatenation. This is a learnable mapping matrix. The first-order fusion term preserves the linear complementary information of each modality, providing a stable basic representation for subsequent interaction modeling.
[0064] Secondly, bilinear interaction modeling is performed on the weighted modal features to obtain a second-order fusion term characterizing the multiplicative coupling relationship between modalities. Linear projections are then performed on the two modal features respectively, and cross-modal interaction representations are formed through element-wise multiplication.
[0065] ;
[0066] in and These are the representations of the two modes after linear projection; This represents a bilinear interactive feature constructed from two modal projection features through element-wise multiplication; and For learnable projection matrix; Represents element-wise product; Indicates different interaction branches; H Let be the number of interaction branches. The above form can characterize the cooperative response relationship between the two modes in different projection spaces, thereby learning richer cross-modal coupling patterns. Since the input to the interaction branches is a reliability-weighted modal representation, the second-order interaction term not only reflects the multiplicative coupling relationship between the two modes, but its response strength is also jointly modulated by the modal weights.
[0067] Subsequently, a multi-expert mapping module is designed within each interaction header. Assume there are a total of... Expert mapping function Then the first h The first interactive header in the k Output from an expert It can be represented as:
[0068] ;
[0069] A multi-expert structure enables different samples to adaptively call more appropriate mapping paths based on their cross-modal relationships, thereby improving the model's adaptability to complex coupling patterns and operating condition differences.
[0070] To enable dynamic selection of experts, a gated input is constructed. :
[0071] ;
[0072] in The representation layer is normalized. This gated input simultaneously includes modality self-representation, modality difference information, and cross-modality related information, helping to more comprehensively characterize the modal relationship state of the current sample. Based on this, expert combination weights are generated through the gated network:
[0073] ;
[0074] in This represents the normalization function, used to map the output of the gated network to a non-negative weight distribution that sums to 1; , Indicates the first k Normalized weights corresponding to each expert; This represents a gating network. Therefore, the first... h Second-order output of each interaction head Represented as:
[0075] ;
[0076] It can be seen that the above-mentioned gating mechanism does not directly participate in the construction of bilinear interaction terms, but rather improves the flexibility and robustness of the second-order fusion process by explicitly modeling the differences and correlations between the two modalities and adaptively adjusting the contributions of different experts to the interaction representation.
[0077] After obtaining the outputs of each interaction head, a weighted aggregation method is further used to form an overall second-order fusion. express:
[0078] ;
[0079] in Indicates the first h The model generates aggregate weights corresponding to each interaction head. By constructing multiple interaction heads in parallel, the model can model cross-modal coupling relationships from different projection subspaces, thereby learning richer cooperative patterns.
[0080] Finally, the first-order fusion term and the second-order fusion term are fused to obtain the fused representation feature. Considering that the contribution of the second-order interaction branch may differ in different training stages and on different samples, a learnable strength coefficient is introduced. By adjusting the second-order fusion term, the final fusion is obtained. express:
[0081] ;
[0082] The interactive fusion step explicitly models the cross-modal multiplicative coupling relationship through bilinear interaction, fully exploring the collaborative characteristics of multi-source information; multi-expert mapping and dynamic gating mechanism enable the model to adaptively select the mapping path according to the sample characteristics, improving the flexibility of fusion; the learnable intensity coefficient balances the contributions of the first-order fusion term and the second-order fusion term, enhancing the generalization ability of the model.
[0083] Finally, based on the fused representation features, the state monitoring results of the hydrostatic spindle are output through a linear classifier. :
[0084] ;
[0085] in W and b These represent the weight matrix and bias vector of the linear classifier, respectively.
[0086] The network structure of the entire reliability-aware fusion model is as follows: Figure 3 As shown, a multi-branch parallel architecture is adopted to extract and fuse features from the four sensor signals respectively.
[0087] Example 2: Multimodal Fusion Monitoring System
[0088] This embodiment provides a multimodal fusion state monitoring system for a hydrostatic spindle, used to implement the method described in Embodiment 1.
[0089] The system includes: a data acquisition module, a feature extraction module, and a fusion monitoring module.
[0090] The data acquisition module is used to acquire multimodal sensing data during the operation of the hydrostatic spindle. This multimodal sensing data includes at least vibration and acoustic signals. In a preferred embodiment, the multimodal sensing data includes at least one vibration signal and one acoustic signal. Multiple vibration signals can be acquired from different measurement points, providing more comprehensive vibration information; the acoustic signals supplement the acoustic radiation information that is difficult to capture with vibration signals. This configuration, through a combination of spatial and physical diversity, improves the ability to sense the spindle's state.
[0091] The feature extraction module is used to extract features from the multimodal sensing data to obtain corresponding modal features. In a preferred embodiment, the feature extraction module includes multiple parallel, parameter-discretionary feature extraction branches, each used to extract features from each signal. Each branch consists of an embedding layer, a self-attention module, and a feedforward network, used to extract the structural features of the time-series signal and form a high-dimensional vector representation. For multiple vibration signals, their features are aggregated to form a unified vibration feature. This parameter-discretionary branch design allows for specialized learning of the unique modes of different signals, preserving the unique information of each signal; intramodal aggregation integrates the complementary information of multiple vibration signals, improving the robustness of the vibration modal representation.
[0092] The fusion monitoring module is internally deployed with a pre-trained reliability perception fusion model, which includes a reliability estimation unit, a weighted fusion unit, an interactive fusion unit, and an output unit.
[0093] The reliability estimation unit is used to construct a variational posterior distribution of modal reliability based on each modal feature, and to determine the statistics of each modal reliability from the variational posterior distribution. In a preferred embodiment, the reliability estimation unit is further used to: construct a variational posterior distribution of modal reliability, the distribution parameters of which are generated by the corresponding modal features through a mapping network; construct a hierarchical conditional prior distribution of modal reliability, the distribution parameters of which are jointly determined by a global term reflecting the global reliability level of the modality and a sample offset term reflecting the reliability shift under specific sample conditions; and constrain the reliability estimation unit by minimizing the KL divergence between the variational posterior distribution and the hierarchical conditional prior distribution. This unit realizes probabilistic modeling and adaptive estimation of modal reliability, providing a reliable basis for subsequent weighted fusion.
[0094] The weighted fusion unit is used to weight the modal features using the statistics to obtain weighted modal features. This unit dynamically adjusts the contribution of each modality based on the estimated reliability, effectively suppressing interference from low-quality modes.
[0095] The interactive fusion unit is used to perform cross-modal collaborative interaction modeling on the weighted modal features to obtain fused representation features. In a preferred embodiment, the interactive fusion unit includes:
[0096] A first-order fusion subunit is used to linearly fuse the weighted modal features to generate a first-order fusion term. This subunit preserves the linear complementarity information of the modalities, providing a basis for subsequent interactions.
[0097] A second-order fusion subunit is used to perform bilinear interaction modeling on the weighted modal features to generate a second-order fusion term. The second-order fusion subunit further includes: a multi-head interaction module, used to perform element-wise multiplication operations on the weighted modal features in multiple different projection subspaces to generate multiple cross-modal interaction representations; a multi-expert mapping module, containing multiple expert mapping functions, used to process each cross-modal interaction representation in parallel; a dynamic gating module, used to generate expert combination weights based on the results of difference and correlation analysis of the current weighted modal features, and to perform weighted summation of the outputs of the multiple expert mapping functions according to the weights to obtain the output of each interaction head; and an aggregation module, used to perform weighted aggregation of the outputs of each interaction head to generate the second-order fusion term. This subunit characterizes the cross-modal multiplicative coupling relationship through bilinear interaction, and the multi-expert and gating mechanism achieves adaptive path selection, significantly improving the expressive power of the fusion representation.
[0098] The fusion subunit combines the first-order fusion term and the second-order fusion term to obtain the fused representation feature. This subunit balances the contributions of linear complementary information and nonlinear coupling information, making the fused representation both stable and expressive.
[0099] The output unit is used to output the state monitoring results of the hydrostatic spindle based on the fused representation features. This unit maps the high-dimensional fused features to specific state categories, realizing end-to-end fault diagnosis.
[0100] Example 3: Experimental Verification
[0101] To verify the technical effectiveness of this invention, this embodiment uses the Huazhong University of Science and Technology's multimodal motor dataset for experimental verification. This dataset was collected from the Spectra-Quest mechanical fault simulator. The test bench monitors six health conditions of the motor: healthy, bearing failure, rotor bending, rotor bar breakage, rotor misalignment, and voltage imbalance. Signals were collected at four different speeds (B1=5Hz, B2=10Hz, B3=20Hz, and B4=30Hz). Vibration and acoustic data were recorded simultaneously, with a sampling rate of 25.6kHz and a sample length of 1024 points for both signals. 1000 samples were prepared for each health condition, with 800 used for training and 200 for testing.
[0102] Four cross-speed diagnostic tasks were designed: T0 (source domain B2, B3, B4 → target domain B1), T1 (source domain B1, B3, B4 → target domain B2), T2 (source domain B1, B2, B4 → target domain B3), and T3 (source domain B1, B2, B3 → target domain B4). Each experiment was repeated five times, and the mean ± standard deviation was reported. Training consisted of 50 epochs with a batch size of 256, and an adaptive decay strategy was used for the learning rate.
[0103] The comparison models include: MUGTN (the current state-of-the-art multimodal fault diagnosis model), a feature-level fusion model based on ResNet18, and ResNet18-B (a decision-level fusion model based on ResNet18). The experimental results are shown in Table 1.
[0104] Table 1 Experimental Results
[0105]
[0106] As shown in Table 1, the proposed method outperforms the comparative models in all four cross-speed-range diagnostic tasks. MUGTN, as an advanced model for multimodal fault diagnosis, achieved a high accuracy of 82.16±0.70% in task T0, while the proposed method further improved this to 83.87±1.56%. In contrast, the ResNet18 model based on feature concatenation and the ResNet18-B model based on majority voting performed significantly worse in this task, indicating that simple feature-level or decision-level fusion is insufficient to fully model the collaborative relationships between modes. In tasks T1 and T2, the proposed method achieved the highest accuracies of 97.36±0.53% and 98.93±0.36%, respectively, demonstrating more stable cross-operating-condition identification capabilities. When the speed difference increases further (T3 task), the performance of all methods decreases, but the proposed method still achieves 91.04±0.76%, which is significantly better than other methods. This shows that the proposed fusion mechanism can more effectively model cross-modal feature interaction, thereby improving the generalization performance of cross-condition diagnosis.
[0107] This invention simulates a fault diagnosis task under modal incompleteness scenarios, setting up two cases: vibration mode incompleteness and acoustic mode incompleteness, to evaluate the diagnostic robustness of different methods under modal incompleteness conditions. The results are as follows: Figure 4 As shown, when vibration modes are missing, the overall diagnostic performance of each method decreases, but the proposed method still maintains a recognition accuracy of 44.82%, significantly better than ResNet18 and ResNet18-B, and close to the current advanced multimodal fault diagnosis model MUGTN. When acoustic modes are missing, the performance of each method improves, with the proposed method achieving the highest recognition accuracy of 92.61%, an improvement of approximately 18.66 percentage points compared to MUGTN, indicating that the proposed method can effectively utilize the remaining modal information even when modes are incomplete. The average results for both missing modes show that the proposed method achieves an average diagnostic accuracy of 68.71%, which is better than the comparison methods overall. This demonstrates that the proposed reliability-aware fusion strategy can adaptively adjust the contribution of different modes even under modal missing conditions, thereby improving the diagnostic stability and robustness of the model.
[0108] Figures 5 to 8 The distribution of decision contribution values under different health states is shown, corresponding to tasks T0 to T3 respectively. This represents the impact of cross-modal interactions on the final predicted probability. Looking at different fault types, the role of the interaction term varies significantly across categories. For bearing faults and rotor bending, the interaction term plays a more significant role in most samples. The presence of a clear positive distribution indicates that cross-modal interaction significantly enhances the model's discriminative ability. This may be because such faults typically cause simultaneous changes in structural vibration and enhanced acoustic radiation; the two modes have a strong coupling relationship in terms of physical mechanisms, and interactive modeling can more effectively extract their synergistic features. In contrast, the healthy state samples... The signals are mostly concentrated near zero, indicating that under normal circumstances, the complementary information between different modes is weak, and single-mode features are sufficient for effective discrimination. For fault types such as rotor bar breakage and voltage imbalance, The distribution of the parameters exhibits greater dispersion, which may be related to the different degrees of impact of the fault on vibration and acoustic signals under different working conditions. Therefore, the interaction terms can play a dynamic role in these scenarios according to the changes in modal information, thereby improving the model's adaptability to complex fault modes.
[0109] Further analysis of modal reliability product With interactive contributions Relationship discovery, such as Figure 9 As shown, some samples The distribution near zero indicates that cross-modal interactions do not significantly affect decision-making in all cases, but rather function adaptively based on modal reliability. The increase, The overall distribution gradually shifts towards the positive direction, while the number of negative contribution samples decreases significantly. This indicates that when both modalities have high reliability, cross-modal interaction tends to have a positive promoting effect on prediction, with fewer inhibitory effects. From a theoretical perspective, when the observation information of both modalities has high reliability, their joint representation is more likely to contain consistent fault discrimination information. Therefore, interaction modeling can effectively strengthen relevant features, thereby improving the prediction confidence of the model. This phenomenon statistically verifies that the proposed reliability-aware modal weighting mechanism can reasonably adjust the modal interaction intensity, enabling the model to adaptively utilize multimodal information under different operating conditions.
[0110] The above experimental results fully demonstrate the effectiveness and superiority of the method and system proposed in this invention.
[0111] Example 4: Industrial Application Case
[0112] This embodiment 4 provides a specific application example of the present invention in an industrial setting.
[0113] The multimodal fusion state monitoring system of this invention was deployed on the hydrostatic spindle test bench of a precision machine tool manufacturing company. Three accelerometers were arranged along the three axes to collect vibration signals, and a microphone was placed near the spindle to collect acoustic signals. The signal acquisition card simultaneously acquired four signals at a sampling rate of 25.6 kHz, with each 1024 points serving as an analysis sample.
[0114] During online system operation, signals are acquired in real time and input into the feature extraction module. The feature extraction module's four parallel branches process three vibration signals and one acoustic signal respectively, outputting a 128-dimensional feature vector. The three vibration signals are then averaged and aggregated to form the vibration features. Acoustic characteristics Common input reliability estimation unit.
[0115] During a certain operation, due to a loose sensor installation, abnormal noise appeared in one vibration signal. The reliability estimation unit automatically detected the degradation in the signal quality and estimated, using variational posterior distribution, that the reliability weight of the vibration mode decreased to 0.32, while the acoustic mode maintained a high weight of 0.68. Based on this, the weighted fusion unit attenuated the vibration characteristics, effectively suppressing interference from low-quality signals.
[0116] The interactive fusion unit performs bilinear interactive modeling on the weighted modal features, constructing cross-modal interactive representations in eight interactive heads. A second-order fusion term is generated through four expert mapping functions and a dynamic gating mechanism, combined with the first-order fusion term, and input into the classifier. The output spindle current state is "bearing failure (early stage)". The system issues an early warning, prompting the operator to check the bearing status.
[0117] After the operators stopped the machine for inspection, they found that the bearing did indeed have slight wear, consistent with the system diagnostic results. This timely warning prevented potentially serious losses from the malfunction worsening.
[0118] This application example demonstrates that the present invention can operate stably in actual industrial environments, automatically adapt to changes in sensor signal quality, accurately identify early faults, and has good engineering application value.
[0119] Example 5: Variant Implementation
[0120] This embodiment 5 provides several variations of the present invention.
[0121] In another embodiment, the multimodal sensing data includes one or more of the following, in addition to vibration and acoustic signals: temperature, pressure, flow, and current. Introducing more physical modes allows for a more comprehensive reflection of the spindle's operating state. Correspondingly, the reliability estimation unit needs to construct a corresponding variational posterior distribution for the newly added modes, and the interactive fusion unit needs to adapt to the multimodal inputs, which can be achieved by modeling multimodal collaborative relationships using methods such as pairwise bilinear interaction or higher-order tensor interaction.
[0122] In another embodiment, the feature extraction module replaces the self-attention module with a convolutional neural network (CNN) or a Transformer structure to adapt to the characteristics of different types of signals. For example, CNNs may be more effective for signals with obvious local patterns, while Transformers may be more advantageous for long-term temporal dependencies.
[0123] In another implementation, the statistic uses the median or mode instead of the expected value. In noisy environments, the median may be more robust than the expected value; under specific distribution patterns, the mode may be more suitable. The most suitable statistic can be selected based on the specific application scenario.
[0124] In another embodiment, the weights of the KL divergence constraint An adaptive adjustment strategy is adopted, which dynamically adjusts the performance based on the validation set during the training process to balance fitting ability and generalization ability.
[0125] In another embodiment, the learnable strength coefficients in the second-order fusion term It can be designed as input-related dynamic coefficients, generated by the gating network based on the characteristics of the current sample, to more finely adjust the balance between the first-order and second-order fusion terms.
[0126] These variations and implementations are all within the protection scope of this invention.
[0127] This invention provides a multi-modal fusion state monitoring method and system for hydrostatic spindles, which can adaptively fuse information from multiple sources under complex working conditions to accurately identify the health status of the spindle, and has the following industrial applicability:
[0128] First, the reliability sensing modal weighting mechanism of this invention can automatically adapt to changes in sensor signal quality, solving the practical problem of sensors being susceptible to interference in industrial settings and improving the reliability of the monitoring system.
[0129] Secondly, the bilinear interactive fusion module of the present invention fully exploits cross-modal collaborative features, enabling earlier and more accurate identification of early faults, providing a basis for predictive maintenance decisions, and avoiding economic losses caused by unplanned downtime.
[0130] Furthermore, the present invention can maintain stable performance even in the case of modal absence, which enhances the fault tolerance of the system and reduces the risk of false alarms or missed alarms caused by sensor failure.
[0131] Finally, the method of the present invention can be deployed on edge computing devices to achieve real-time online monitoring, or it can be integrated into machine tool CNC systems to provide intelligent solutions for spindle health management.
[0132] In summary, this invention has significant industrial practical value and can be widely used in fields such as high-end machine tools and precision machining equipment.
[0133] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art can make various improvements and modifications without departing from the spirit and principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multi-modal fusion state monitoring method for a hydrostatic spindle, characterized in that, Includes the following steps: Acquire multimodal sensing data during the operation of the hydrostatic spindle, wherein the multimodal sensing data includes at least vibration signals and acoustic signals; Feature extraction is performed on the multimodal sensing data to obtain the corresponding modal features; The modal features are input into a pre-trained reliability-aware fusion model, and the following operations are performed: Based on the features of each modality, the reliability weights of each modality are determined. This determination includes: constructing a variational posterior distribution of modal reliability and a hierarchical conditional prior distribution of modal reliability; constraining the model by minimizing the KL divergence between the variational posterior distribution and the hierarchical conditional prior distribution during training; and determining the statistical measure of modal reliability from the variational posterior distribution output by the trained model. The reliability is used to characterize the effective information strength of the corresponding modality in the current sample. The modal features are weighted using the statistical measures to obtain weighted modal features; Cross-modal collaborative interaction modeling is performed on the weighted modal features to obtain fused representation features. The method includes: linearly fusing the weighted modal features to obtain a first-order fusion term; performing bilinear interaction modeling on the weighted modal features to obtain a second-order fusion term; and fusing the first-order fusion term and the second-order fusion term to obtain the fused representation features. Based on the fusion representation features, the status monitoring results of the hydrostatic spindle are output.
2. The multi-modal fusion state monitoring method for a hydrostatic spindle according to claim 1, characterized in that, The statistic is the expected value of the variational posterior distribution.
3. The multi-modal fusion state monitoring method for a hydrostatic spindle according to claim 1, characterized in that, Both the variational posterior distribution and the hierarchical conditional prior distribution are modeled using a log-normal distribution, and the KL divergence is calculated using a closed-form solution.
4. The multi-modal fusion state monitoring method for a hydrostatic spindle according to claim 1, characterized in that, The distribution parameters of the hierarchical conditional prior distribution are jointly determined by a global term reflecting the global reliability level of the modality and a sample offset term reflecting the reliability shift under specific sample conditions.
5. The multi-modal fusion state monitoring method for a hydrostatic spindle according to claim 1, characterized in that, The step of performing bilinear interaction modeling on the weighted modal features to obtain a second-order fusion term for characterizing the multiplicative coupling relationship between modalities includes: The weighted modal features are projected onto multiple interaction heads, and cross-modal interaction representations are constructed in each interaction head by element-wise multiplication. The cross-modal interaction representation is processed using a multi-expert mapping module, and the contribution weights of each expert are dynamically adjusted based on the results of the difference and correlation analysis of the current weighted modal features through a gating network to obtain the output of each interaction head. The outputs of each interaction head are weighted and aggregated to obtain the second-order fusion term.
6. A multi-modal fusion state monitoring system for a hydrostatic spindle, characterized in that, include: The data acquisition module is used to acquire multimodal sensing data during the operation of the hydrostatic spindle, wherein the multimodal sensing data includes at least vibration signals and acoustic signals; The feature extraction module is used to extract features from the multimodal sensing data to obtain corresponding modal features; The fusion monitoring module has a pre-trained reliability-aware fusion model deployed inside it, the reliability-aware fusion model including: The reliability estimation unit, used to determine the reliability weights of each modality based on each modal feature, is configured to: construct a variational posterior distribution of modal reliability and a hierarchical conditional prior distribution of modal reliability; constrain the model by minimizing the KL divergence between the variational posterior distribution and the hierarchical conditional prior distribution during training; and determine the statistics of modal reliability from the variational posterior distribution output by the trained model. The weighted fusion unit is used to weight the modal features using the statistics to obtain the weighted modal features; An interactive fusion unit is configured to perform cross-modal collaborative interaction modeling on the weighted modal features to obtain fused representation features. The unit is configured to: perform linear fusion on the weighted modal features to obtain a first-order fusion term; perform bilinear interaction modeling on the weighted modal features to obtain a second-order fusion term; and fuse the first-order fusion term and the second-order fusion term to obtain the fused representation features. The output unit is used to output the status monitoring results of the hydrostatic spindle based on the fused representation features.
7. The multi-modal fusion state monitoring system for a hydrostatic spindle according to claim 6, characterized in that, The reliability estimation unit is further used for: A variational posterior distribution for modal reliability is constructed, the distribution parameters of which are generated by the corresponding modal features through a mapping network; A hierarchical conditional prior distribution for modal reliability is constructed. The distribution parameters of the hierarchical conditional prior distribution are jointly determined by a global term reflecting the global reliability level of the modality and a sample offset term reflecting the reliability shift under specific sample conditions.
8. The multi-modal fusion state monitoring system for a hydrostatic spindle according to claim 6, characterized in that, The interactive fusion unit includes: The first-order fusion subunit is used to linearly fuse the weighted modal features to generate a first-order fusion term. A second-order fusion subunit is used to perform bilinear interaction modeling on the weighted modal features to generate a second-order fusion term. The second-order fusion subunit includes: The multi-head interaction module is used to perform element-wise multiplication of the weighted modal features in multiple different projection subspaces to generate multiple cross-modal interaction representations. The multi-expert mapping module contains multiple expert mapping functions for parallel processing of each cross-modal interaction representation; The dynamic gating module is used to generate expert combination weights based on the results of differential and correlation analysis of the current weighted modal features, and to perform weighted summation of the outputs of multiple expert mapping functions according to the weights to obtain the output of each interaction head; The aggregation module is used to perform weighted aggregation of the outputs of each interaction head to generate the second-order fusion term; The fusion subunit is used to combine the first-order fusion term and the second-order fusion term to obtain the fusion representation feature.
9. The multimodal fusion state monitoring system for a hydrostatic spindle according to any one of claims 6 to 8, characterized in that, The multimodal sensing data includes at least one vibration signal and one acoustic signal. The feature extraction module contains multiple parallel feature extraction branches with non-shared parameters, which are used to extract features from each signal and aggregate the features of at least one vibration signal to form a unified vibration feature.
Citation Information
Patent Citations
CN121765842A
WO2026031415A1