Method and system for intelligent monitoring based on ai multi-modal and internet of things

CN122863267APending Publication Date: 2026-10-02XIAOWEI TECH (ZHUHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610967988.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

[0005]本发明提供了一种基于AI多模态和物联网的智能监测的方法及系统,以解决负载工况下阈值漂移导致误报漏报的问题

Benefits of technology

(1)本发明以振动加速度序列为统一时间基准,对温度标量序列采用线性插值、对电流频率分量采用插值重采样,使三类数据在同一采样时刻一一对应,并将对齐后的数据按时间窗口分段写入存储数据块,同时生成时空关联索引表。通过该处理,温度与电流的低采样数据被映射到振动采样时刻,避免多通道不同采样率导致的错位分析,使后续特征计算基于同一时间片完成,降低跨通道对齐误差并提升检索定位效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122863267A_ABST
    Figure CN122863267A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of Internet of Things chips, and discloses a method and system for intelligent monitoring based on AI multi-modal and Internet of Things, which comprises the following steps: acquiring vibration acceleration time series, temperature scalar and current frequency component, synchronizing the vibration acceleration as a reference time series, constructing a multi-dimensional signal matrix and segmenting and indexing encapsulation to obtain an original data set; obtaining a standardized feature set through feature extraction and standardization, obtaining a multi-modal fusion embedding vector through impact mapping and spectrum weighted fusion; obtaining an enhanced embedding vector through feature weighted fusion and weighted increment, screening semantic correlation features based on feature decoupling and mutual information, and then obtaining an abnormal feature set through sparse reconstruction, residual impact elimination and spectrum clustering encapsulation; determining a fault type according to the abnormal feature set; determining a damage degree through amplitude mapping grading; and generating complete early warning information. The method can solve the problem of false positives and false negatives caused by threshold drift under load working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of IoT chip technology, and in particular to a method and system for intelligent monitoring based on AI multimodal and IoT. Background Technology

[0002] In the field of sensor network chips, applications such as industrial equipment status sensing, edge data acquisition, and intelligent monitoring are targeted, focusing on solving the problems of accessing, converting, synchronizing, processing, and transmitting multiple types of sensor signals. During the operation of motors, bearings, pumps, and other rotating equipment, signals such as vibration, temperature, and current typically exhibit different sampling frequencies, data formats, and variation characteristics. Therefore, sensor network chips are needed to achieve unified acquisition, timing alignment, feature extraction, and preliminary edge-side analysis of multi-source heterogeneous signals, thereby providing fundamental support for subsequent fault identification, status assessment, and early warning output.

[0003] In existing technologies, health indicators and alarm thresholds are pre-set for each monitoring channel, such as vibration, temperature, and current. During operation, window statistics for each channel are extracted at fixed intervals and compared with corresponding rules one by one. An alarm event is output when the conditions are met. However, this method of independently judging based on the statistics of a single channel and relying on fixed thresholds to trigger alarms is based on the premise that the characteristic distribution of each channel is relatively stable. When the equipment is under changing operating conditions such as load fluctuations, changes in ambient temperature, or process switching, the statistics of each channel will shift with the overall operating conditions, which can easily lead to false alarms or missed alarms.

[0004] In summary, existing technologies suffer from the problem of false alarms and missed alarms due to threshold drift under load conditions. Summary of the Invention

[0005] This invention provides a method and system for intelligent monitoring based on AI multimodal and IoT to solve the problem of false alarms and missed alarms caused by threshold drift under load conditions.

[0006] In a first aspect, to address the aforementioned technical problems, this invention provides a method for intelligent monitoring based on AI multimodal and IoT, comprising: Obtain the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor; Based on the vibration acceleration sequence, the temperature scalar sequence and the current frequency component are subjected to time synchronization processing to obtain a multidimensional signal matrix. The multidimensional signal matrix is ​​then segmented, indexed, and encapsulated to obtain the original data set. The original dataset is subjected to temperature feature extraction processing to obtain the temperature change rate, and the temperature change rate is then standardized to obtain a standardized feature set. Based on the standardized feature set, a three-channel feature map is constructed, and a trend dependency vector sequence is obtained by recursive calculation based on the three-channel feature map; based on the trend dependency vector sequence, a spectral weighted fusion process is performed to obtain a multimodal fusion embedding vector. Based on the multimodal fusion embedding vector, feature weighted fusion is performed to obtain the original weight value vector. Based on the original weight value vector, the multimodal fusion embedding vector is subjected to weighted incremental processing to obtain the enhanced embedding representation vector. Based on the enhanced embedding representation vector, feature node construction processing is performed to obtain independent feature nodes in the semantic association graph, and association feature extraction processing is performed on the independent feature nodes to obtain semantic association features. Based on the semantic association features, sparse reconstruction extraction is performed to obtain the repeated impact interval component; and residual impact removal is performed on the repeated impact interval component to obtain the second residual vector. Based on the second residual vector, spectral clustering is performed to encapsulate the result and obtain the abnormal feature set. Based on the set of abnormal features, the fault type is determined to obtain the fault type; Based on the fault type, amplitude mapping is performed to determine the damage level, and based on the damage level, periodic retrieval is performed to generate complete early warning information.

[0007] Secondly, the present invention provides an intelligent monitoring system based on AI multimodal and IoT, comprising: The data acquisition module is used to acquire the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor. The data processing module is used to perform time-series synchronization processing on the temperature scalar sequence and current frequency component based on the vibration acceleration sequence to obtain a multidimensional signal matrix, and to perform segmented indexing and encapsulation processing on the multidimensional signal matrix to obtain the original data set. The standardization module is used to extract temperature features from the original dataset to obtain the temperature change rate, and to standardize the temperature change rate to obtain a standardized feature set. The embedding vector module is used to construct a three-channel feature map based on the standardized feature set, and to perform recursive calculation based on the three-channel feature map to obtain a trend dependency vector sequence; and to perform spectral weighted fusion processing based on the trend dependency vector sequence to obtain a multimodal fusion embedding vector. The enhancement vector module is used to perform feature weighted fusion based on the multimodal fusion embedding vector to obtain a weighted original value vector, and to perform weighted incremental processing on the multimodal fusion embedding vector based on the weighted original value vector to obtain an enhanced embedding representation vector. The associated feature module is used to perform feature node construction processing based on the enhanced embedding representation vector to obtain independent feature nodes in the semantic association graph, and to perform associated feature extraction processing on the independent feature nodes to obtain semantic association features. An anomaly feature module is used to perform sparse reconstruction extraction based on the semantic association features to obtain a repeated impact interval component; and to perform residual impact removal on the repeated impact interval component to obtain a second residual vector; and to perform spectral clustering encapsulation based on the second residual vector to obtain an anomaly feature set. The fault determination module is used to determine the fault type based on the set of abnormal features to obtain the fault type; The early warning generation module is used to perform amplitude mapping and classification based on the fault type to obtain the degree of damage, and to perform periodic retrieval and generation based on the degree of damage to obtain complete early warning information.

[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention uses the vibration acceleration sequence as a unified time reference, employs linear interpolation for the temperature scalar sequence and interpolation resampling for the current frequency component, so that the three types of data correspond one-to-one at the same sampling time. The aligned data is then segmented into storage data blocks according to time windows, and a spatiotemporal correlation index table is generated simultaneously. Through this processing, the low-sampling data of temperature and current are mapped to the vibration sampling time, avoiding misalignment analysis caused by different sampling rates in multiple channels. This allows subsequent feature calculations to be completed based on the same time slice, reducing cross-channel alignment errors and improving retrieval and positioning efficiency.

[0009] (2) This invention calculates the peak factor, root mean square value, temperature change rate, current fundamental frequency amplitude, harmonic energy ratio, and sideband amplitude on the original dataset, and performs Z-score standardization on the above features using the mean and standard deviation of historical normal data, so that features with different dimensions are transformed into a unified scale that can be directly compared. Through this processing, the degree of abnormality of vibration, temperature, and current features is expressed in the same numerical space, reducing the problem of threshold dependence on single-channel experience and improving the consistent characterization ability of early weak anomalies on multiple channels.

[0010] (3) This invention performs sparse coding on semantic association features and reconstructs the repeated impact interval component. Then, the repeated impact interval component is subtracted from the semantic association features to obtain a residual vector. Impact point determination is performed on the residual vector and set to zero to obtain a second residual vector. Subsequently, the spectral energy ratio of the second residual vector is calculated to identify the broadband noise rise component. Adjacent impact points are clustered within a preset low amplitude range to obtain a low amplitude impact cluster. Finally, the clusters are encapsulated to form an abnormal feature set. Through this processing, periodic impacts, non-periodic impacts, broadband noise, and low amplitude impacts are separated and quantized, reducing misjudgments caused by signal aliasing and making the anomaly representation closer to the fault mode.

[0011] (4) This invention calculates the cosine of the angle between the abnormal feature set and the preset fault feature template vector to determine the fault type. After locating the target position based on the fault type, the peak value is extracted from the vibration signal at the target position as the feature amplitude. The damage degree is obtained by classifying the amplitude range. Then, the inspection cycle is generated by consulting the rule table according to the damage degree, and complete early warning information is output. Through this processing, the output result is transformed from "abnormal value" to an executable conclusion of "defect type, defect location, damage level and inspection cycle", which reduces the workload of manual secondary judgment and improves the operability of the early warning. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the intelligent monitoring method based on AI multimodal and IoT provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of an intelligent monitoring system based on AI multimodal and IoT provided in the second embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] Reference Figure 1 The first embodiment of the present invention provides an intelligent monitoring method based on AI multimodal and IoT, comprising the following steps: S11, obtain the vibration acceleration sequence, temperature scalar sequence and current frequency component of the motor; S12, using the vibration acceleration sequence as a reference, perform time synchronization processing on the temperature scalar sequence and current frequency component to obtain a multidimensional signal matrix, and perform segmented indexing and encapsulation processing on the multidimensional signal matrix to obtain the original data set. S13, perform temperature feature extraction processing on the original data set to obtain the temperature change rate, and perform standardization processing on the temperature change rate to obtain a standardized feature set; S14. Based on the standardized feature set, construct a three-channel feature map, and perform recursive calculation based on the three-channel feature map to obtain a trend dependency vector sequence; based on the trend dependency vector sequence, perform spectral weighted fusion processing to obtain a multimodal fusion embedding vector. S15, based on the multimodal fusion embedding vector, perform feature weighted fusion to obtain the original weight value vector, and based on the original weight value vector, perform weighted incremental processing on the multimodal fusion embedding vector to obtain the enhanced embedding representation vector; S16, Based on the enhanced embedding representation vector, feature node construction processing is performed to obtain independent feature nodes in the semantic association graph, and association feature extraction processing is performed on the independent feature nodes to obtain semantic association features; S17. Based on the semantic association features, perform sparse reconstruction extraction to obtain the repeated impact interval component; and perform residual impact removal on the repeated impact interval component to obtain the second residual vector; and perform spectral clustering encapsulation based on the second residual vector to obtain an abnormal feature set. S18. Based on the set of abnormal features, determine the fault type to obtain the fault type; S19. Based on the fault type, amplitude mapping is performed to determine the damage level, and based on the damage level, periodic retrieval is performed to generate complete early warning information.

[0015] In step S11, the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor are obtained.

[0016] It is worth noting that the vibration acceleration sequence is generated by an ICP piezoelectric accelerometer deployed at the motor bearing housing, which outputs an analog acceleration signal; the temperature scalar sequence is generated by a platinum resistance temperature probe deployed at the stator winding, which outputs a corresponding temperature electrical signal to form a temperature scalar sequence; and the current frequency component is generated by a closed-loop Hall current sensor deployed at the stator input line, which outputs a current signal.

[0017] In step S12, based on the vibration acceleration sequence, the temperature scalar sequence and the current frequency component are subjected to time-series synchronization processing to obtain a multidimensional signal matrix. The multidimensional signal matrix is ​​then segmented, indexed, and encapsulated to obtain the original data set, including: Using the vibration acceleration sequence as a reference, the temperature scalar sequence is linearly interpolated to obtain an aligned temperature value sequence; Based on the vibration acceleration sequence, the current frequency components are subjected to FIR interpolation to obtain an aligned current value sequence. The vibration acceleration sequence, the aligned temperature value sequence, and the aligned current value sequence are concatenated column by column to obtain a time-synchronized multidimensional signal matrix; The multidimensional signal matrix is ​​segmented according to a preset time window to obtain storage data blocks, and a spatiotemporal correlation index table is added to the storage data blocks to obtain the original data set.

[0018] It is worth noting that the temperature scalar sequence is linearly interpolated based on the vibration acceleration sequence to obtain the aligned temperature value sequence. The linear interpolation uses the sampling time and temperature value of two adjacent temperature sampling points to calculate the temperature value corresponding to the vibration sampling time through proportional calculation.

[0019] To obtain an aligned current value sequence, the current frequency component undergoes a 4x FIR interpolation upsampling. Specifically, three zero-value samples are inserted between two adjacent sampling points to form an intermediate sequence. This intermediate sequence is then reconstructed using pre-set symmetrical low-pass filter coefficients, resulting in current values ​​aligned with the vibration sampling time. Finally, the vibration acceleration values, linearly interpolated temperature values, and FIR-interpolated current values ​​corresponding to the same target time are concatenated column-wise to form a time-synchronized multidimensional signal matrix.

[0020] In the FIR convolution reconstruction, the weights are pre-set according to the symmetric low-pass reconstruction kernel. First, a symmetric integer template with the largest weight at the center and decreasing weights on both sides with increasing distance is selected, such as [1,2,4,2,1]. Then, the template is subjected to maximum and minimum normalization to make the sum of all weights equal to 1, ensuring that the signal amplitude after interpolation retains its original dimensions.

[0021] Subsequently, the time-synchronized multidimensional signal matrix is ​​segmented into 10-minute windows to obtain the first segment of the stored data block, corresponding to the time range of 10:00:00–10:10:00. A spatiotemporal correlation index table is added to the stored data block to obtain the original data set. The structure of the preset spatiotemporal correlation index table is predefined during system initialization and includes fields such as device number, window start and end timestamps, storage address, and data length. During segmentation, the vibration, temperature, and current statistical characteristics within the current window are calculated and written into the index table along with the aforementioned fields to obtain the original data set.

[0022] The duration of the window segmentation is used to process historical data and evaluate the accuracy of fault diagnosis with different window lengths, and the window length with the highest accuracy is selected as the preset value; the accuracy of each group of data is obtained by calculating the ratio of the number of correctly classified samples to the total number of samples; for example, a 10-minute window can be selected as the optimal time window duration.

[0023] In step S13, temperature feature extraction processing is performed on the original dataset to obtain the temperature change rate, and the temperature change rate is standardized to obtain a standardized feature set, including: The absolute peak value and root mean square value are calculated for the vibration acceleration sequence of the original dataset, and the peak factor is calculated using the absolute peak value and the root mean square value. The temperature scalar sequence of the original dataset is calculated with a weighted moving average temperature value according to a preset feature extraction window, and the difference between the weighted moving average temperature values ​​adjacent to the preset feature extraction window is calculated to obtain the temperature change rate. The continuous interval where the temperature change rate is less than the preset temperature stability threshold is taken as the stable operating condition period. During the stable operating condition period, the aligned current value sequence of the original data set is sliced ​​according to the preset spectrum analysis duration to perform spectrum calculation, so as to obtain the fundamental frequency amplitude, harmonic component energy ratio and sideband characteristic amplitude. The peak factor, root mean square value, temperature change rate, fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude are summarized into a feature vector. The feature vector is then subjected to Z-score standardization based on the mean and standard deviation of historical normal data to obtain a standardized feature set.

[0024] It is worth noting that the vibration acceleration sequence, temperature scalar sequence, and current frequency component within the target time window of the original data are read first. The peak factor and root mean square value are calculated for the vibration acceleration sequence in slices of fixed duration. For example, the slice length is 1 second. First, all sampling points within this 1 second are traversed to obtain the maximum absolute value as the absolute peak value. Then, the average of the squares of all sampling points within the same slice is calculated and the root mean square value is obtained. The peak factor is obtained by dividing the absolute peak value by the root mean square value.

[0025] The temperature scalar sequence of the original dataset is then calculated to obtain the rate of temperature change. Specifically, the arithmetic mean of temperature sampling points within a 300-second window is obtained. Then, an exponentially weighted moving average is performed between adjacent windows, with an exponential smoothing factor of 0.05. The calculation method is that the current output temperature equals the previous output temperature multiplied by 0.95 plus the current window mean multiplied by 0.05, thus obtaining a continuous exponentially weighted moving average temperature value. The rate of temperature change is obtained by dividing the difference between adjacent output temperatures by the corresponding time interval.

[0026] The exponential smoothing factor of 0.05 is determined based on the time scale of temperature signal changes. Temperature, as a slow variable, changes slowly on a minute-by-minute basis in industrial operations. In this step, after averaging over a 300-second window, to avoid occasional fluctuations in a single window directly altering the trend judgment, the weight of the single-window mean in the exponential weighting is fixed at 5%. This ensures that the output requires changes in the same direction over multiple consecutive windows to show a significant shift, which is equivalent to using the cumulative results of approximately 20 windows as the primary trend basis, thereby stabilizing the trend judgment at approximately 100 minutes.

[0027] Subsequently, current features are extracted based on the original dataset. Current features are extracted only during stable operating periods, which are divided into continuous intervals where the rate of temperature change is less than a preset temperature stability threshold. Within these intervals, the current sampling sequence is sliced ​​in 1-second increments for spectral calculation. Specifically, a Discrete Fourier Transform is performed on the current samples within each slice, and the amplitude at 50 Hz is extracted as the fundamental frequency amplitude. The energy ratio of the harmonic components is calculated by summing the squares of the amplitudes at the second and third harmonic frequencies and then comparing them with the energy ratio across the entire frequency band. Sideband frequency amplitudes are extracted at slip frequency intervals on both sides of 50 Hz as sideband feature amplitudes.

[0028] The preset temperature stability threshold can be set to 0.1℃ / s. This threshold uses the temperature change rate as a quantitative indicator of operating condition stability and is determined using the normal distribution statistical results of historical stable operating data. Specifically, within the historically marked stable operating time period, a sample set of temperature change rates is statistically analyzed, and its mean and standard deviation are calculated. It is found that the temperature change rate follows a normal distribution. The threshold is set to the mean plus three times the standard deviation, corresponding to approximately 99.7% of the stable sample coverage. The mean temperature change rate of the stable samples is 0.02℃ per minute, and the standard deviation is 0.026℃ per minute. Therefore, the mean plus three times the standard deviation is 0.098℃ per minute. In engineering implementation, this is rounded to 0.1℃ per minute as the judgment threshold, thereby ensuring that the interval classified as stable operating condition belongs to a high confidence range of stable distribution in a statistical sense.

[0029] Subsequently, the peak factor and root mean square value of the vibration acceleration sequence, the temperature change rate, the fundamental frequency amplitude of the current, the harmonic component energy ratio, and the sideband characteristic amplitude are summarized into a feature vector for the same time slice. Z-score standardization is then performed on each feature. The mean and standard deviation used for standardization are taken from the statistical results of historical normal datasets. The calculation method is to subtract the historical mean from the current feature value and then divide by the historical standard deviation, resulting in a standardized feature set with uniform dimensions. These six standardized values ​​are then packaged into a single standardized feature record according to the same timestamp. The final output standardized feature set is saved in units of time slices.

[0030] The preset spectrum analysis duration, through quantitative analysis, balances frequency resolution and harmonic energy ratio. Based on Fourier transform theory, the frequency resolution of the spectrum is inversely proportional to the duration; a longer duration provides higher resolution but misses rapid changes in the device. Comparative experiments show that a 1-second window can capture instantaneous changes in the device in real time while maintaining high resolution, especially performing best in rapidly changing fault modes.

[0031] In step S14, a three-channel feature map is constructed based on the standardized feature set, and a trend dependency vector sequence is obtained by recursive calculation based on the three-channel feature map; based on the trend dependency vector sequence, a spectral weighted fusion process is performed to obtain a multimodal fusion embedding vector, including: The fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude of the standardized feature set are extracted to form the current standardized spectrum features; Peak factors and root mean square values ​​are extracted from the standardized feature set to form a vibration array. The vibration array is then slid along the time axis with three preset window lengths to form three sets of sliding window data. The three sets of sliding window data are multiplied and summed point by point according to preset weights to obtain the convolution output sequence of the three sets of sliding window data. The convolution output sequence is then stacked by channel at the same timestamp to obtain a three-channel feature map. The temperature change rate sequence is read based on the timestamp of the three-channel feature map, and the temperature change rate sequence is recursively calculated in chronological order to obtain the trend dependence vector sequence. The high-frequency harmonic energy ratio is calculated based on the trend dependency vector of the current normalized spectrum feature. When the high-frequency harmonic energy ratio exceeds a preset baseline, the current normalized spectrum feature is dimensionally scaled to obtain the scaled current normalized spectrum feature. The three-channel feature map, the trend dependency vector sequence, and the scaled current-normalized spectral features are subjected to weighted fusion compression processing to obtain a multimodal fusion embedding vector.

[0032] It is worth noting that the fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude are first extracted from the standardized feature set and then arranged in order to obtain the current standardized spectrum. For example, in a certain time slice, the standardized fundamental frequency amplitude is 2.00, the harmonic component energy ratio is 3.00, and the sideband characteristic amplitude is 2.80. Then, by arranging the fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude in a fixed order, the current standardized spectrum [2.00, 3.00, 2.80] corresponding to that time slice is obtained. For multiple time slices, the above three standardized features can be combined in sequence to form the current standardized spectrum features.

[0033] Subsequently, the peak factor and root mean square value are extracted from the standardized feature set to form a vibration array, and a multi-scale one-dimensional convolution is performed on this array to obtain a three-channel feature map. Specifically, the multi-scale one-dimensional convolution takes a sliding window of length L at each time position t, multiplies the L sampled values ​​within the window with the preset L weights point by point, and then sums them to obtain the convolution output value at that position. This output value is used to quantify the concentration of vibration changes within the window. To simultaneously cover impact characteristics of different durations, L was selected as 3, 9, and 27 points, and the calculation was performed point by point along the time axis to obtain three convolutional output sequences corresponding to the scales. By analyzing the historical vibration signals of the target equipment (such as a motor), the duration of its fault impact was mainly distributed in three typical ranges. The 3-point scale showed the strongest response to a single spike, used to highlight instantaneous impacts; the 9-point scale maintained a continuous response to recurring pulse sequences, used to describe periodic impacts; and the 27-point scale maintained a stable response to slow rises or changes in impact density over a longer period, used to characterize trend anomalies. The specific window length was adjusted according to the actual sampling frequency of the target equipment and typical fault mechanisms, with the principle being that each window could completely encompass the impact waveform of the corresponding scale. The three sequences were aligned at the same time and then stacked by channel to form a local impact mode feature mapping map.

[0034] The preset weights of each point within the convolution window are generated offline and solidified by a set of fixed symmetric low-pass weighted templates. The generation process first determines the window length L, and then constructs an integer template with the largest weight at the center and decreasing weights on both sides with decreasing distance, and is symmetrical to ensure that time alignment does not produce offset. For example, when L is 3 points, the integer template [1,2,1] is taken, and when L is 9 points, the integer template [1,2,3,4,5,4,3,2,1] is taken. Then, the integer templates are normalized so that the sum of all weights is equal to 1. The normalized weights are used as the fixed weights of the convolution at this scale.

[0035] Next, using the timestamps of the three-channel feature map as a reference, the temperature change rate sequence is read, and then the temperature change rate is recursively calculated point by point in chronological order to obtain the trend dependence vector. Specifically, the recursive calculation uses three gating coefficients to control the retention ratio of historical accumulated information, the writing ratio of current new information, and the final output ratio, respectively, so that short-term fluctuations are suppressed while the continuous heating or cooling trend is accumulated and preserved. In the specific processing, the temperature change rate at each moment is concatenated with the output state at the previous moment, and the gating coefficients of the forget gate, input gate, and output gate are calculated respectively. The gating coefficients are first obtained by linear weighted summation, and then compressed to the range of zero to one by an sigmoid function to form the proportional coefficients. Then, the forget gate coefficient is used to proportionally retain the accumulated state at the previous moment, and the input gate coefficient is used to proportionally write the candidate information at the current moment. The two are added together to obtain the updated accumulated state. Finally, the updated accumulated state is subjected to hyperbolic tangent compression and multiplied by the output gate coefficient to obtain the trend output vector at the current moment. The continuous output vectors constitute the trend dependence vector sequence. For example, the cumulative state at the previous time step is 0.60, the output state at the previous time step is 0.20, and the current temperature change rate is 0.01 degrees Celsius per minute. After linear weighting and sigmoid function, the forget gate coefficient is 0.90, the input gate coefficient is 0.20, and the output gate coefficient is 0.80. After linear weighting and hyperbolic tangent, the candidate information is 0.70. Then the updated cumulative state is 0.90 multiplied by 0.60 plus 0.20 multiplied by 0.70 to get 0.68. The current trend output is 0.80 multiplied by hyperbolic tangent 0.68 to get 0.47.

[0036] It should be noted that the linear transformation weights of the forget gate, input gate, and output gate used in the recursive calculation are obtained through training on historical normal operation data. By collecting temperature change rate sequences under historical stable operating conditions as training samples, and aiming to predict the temperature change rate at the next moment, the weight matrix is ​​optimized using a backpropagation algorithm over time. After training convergence, fixed weights are obtained for online recursive calculation. The multiple sets of linear weights used in the weighted fusion compression processing are obtained through training on historical multimodal fusion embedding vectors and their corresponding fault labels. The multimodal fusion embedding vectors of each time slice in the historical data are collected as input, and their corresponding fault type labels are used as supervision signals. A fully connected network is trained to map the embedding vectors to fixed-length output vectors. The trained network weights are the linear weights. By setting different output dimensions or using multiple different networks, multiple sets of linear weights can be obtained.

[0037] After obtaining the trend-dependent vector sequence, it is directly concatenated with the current-normalized spectral features of the normalized feature set in sequence to form a joint vector containing temperature trend information and current spectrum information. Then, the proportion of high-frequency harmonic energy is calculated within the current spectrum features. Specifically, the amplitude of each harmonic is squared and summed to obtain the high-frequency energy. Simultaneously, the amplitudes of all frequency points within the entire frequency band are squared and summed to obtain the total energy. The high-frequency energy is divided by the total energy to obtain the high-frequency proportion. When the proportion of high-frequency harmonic energy exceeds a preset baseline of 0.05, a dimension-by-dimensional scaling process is performed on the current spectrum features. Specifically, a corresponding scaling factor is pre-set for each dimension of the spectrum features. All scaling factors are arranged diagonally. During calculation, the first spectral dimension is multiplied by the first factor, the second spectral dimension by the second factor, and so on. The factor for the dimension corresponding to the high-frequency harmonic band is 2.2, and the factor for the other dimensions is 1.0.

[0038] The scaled, normalized current spectral features are obtained, and then weighted fusion compression is performed by combining the vibration and shock feature map and the trend dependence vector. Specifically, the three vibration and shock channels are multiplied by 1.8, and each dimension of the trend dependence vector is multiplied by 1.4. The scaled spectral features are kept unchanged and not scaled again. Then, the three types of data are concatenated into a long vector in a fixed order: vibration and shock channels first, trend dependence vector second, and spectral features third. Finally, linear compression is performed on this long vector. Linear compression uses a set of pre-stored linear weights. Each dimension of the long vector is multiplied by its corresponding linear weight and summed to obtain the first output value. Then, another set of linear weights is used to multiply the same long vector dimension by dimension and summed to obtain the second output value. This process is repeated until a fixed-length output sequence is obtained. This fixed-length output sequence is the multimodal fusion embedding vector. For example, the trend dependence vector within a certain time slice is taken as four-dimensional values ​​of 0.10, 0.25, 0.40 and 0.35, and the current normalized spectral features are taken as six-dimensional values ​​of 1.2, 0.8, 0.5, 2.0, 1.6 and 0.9, where the latter three dimensions correspond to the high-frequency harmonic band. First, calculate the high-frequency energy as 2.0 squared, 1.6 squared, and 0.9 squared to get 7.37. The total energy is the sum of the squares of the six dimensions to get 9.70. The high-frequency proportion is 7.37 divided by 9.70 to get 0.76. Then, scale the spectral features dimension by dimension. Multiply the last three dimensions by 2.2 to get 4.4, 3.52, and 1.98 respectively, while multiplying the other dimensions by 1.0 to keep them unchanged. The scaled spectral vectors are 1.2, 0.8, 0.5, 4.4, 3.52, and 1.98. Then, multiply the vibration and shock three-channel values ​​of 0.60, 0.90, and 0.50 by 1.8 to get 1.08, 1.62, and 0.90 respectively. Multiply the four-dimensional trend dependence vector by 1.4 to get 0.14, 0.35, 0.56, and 0.49 respectively. The vectors are then concatenated in a fixed order to form a long vector: 1.08, 1.62, 0.90, 0.14, 0.35, 0.56, 0.49, 1.2, 0.8, 0.5, 4.4, 3.52, 1.98. During linear compression, the first set of linear weights is used to multiply the long vector dimension-by-dimensional and sum the results to obtain the first output value. For example, if the weights are 0.10, 0.05, 0.05, and the rest are 0, the first output value is 0.10 x 1.08 + 0.05 x 1.62 + 0.05 x 0.90, resulting in 0.234. Similarly, the second set of linear weights is multiplied dimension-by-dimensional and summed to obtain the second output value. For example, if the weights are 0.02, 0.02, and the rest are 0 (corresponding to dimensions 4.4 and 3.52), the second output value is 0.02 x 4.4 + 0.02 x 3.52, resulting in 0.1584. Repeat the calculation in the same way with different weights until a fixed-length output sequence is obtained. This output sequence is the multimodal fusion embedding vector of that time slice.

[0039] The preset baseline of 0.05 is obtained from the statistical analysis of the proportion of high-frequency harmonic energy in historical normal operating data. Specifically, the proportion of high-frequency harmonic energy is calculated segment by segment from current spectrum samples marked as normal operating conditions in history, forming a proportion sample set. The mean and standard deviation of this set are then calculated. Since the proportion samples satisfy the normal distribution condition, the baseline is set as the mean plus twice the standard deviation to cover approximately 95% of the normal samples. For example, if the mean of the high-frequency harmonic energy proportion of the normal samples is 0.03 and the standard deviation is 0.01, then the mean plus twice the standard deviation equals 0.03 plus 0.02, resulting in 0.05. This 0.05 is then set as the preset baseline.

[0040] The reason for setting the frequency scaling factor to 2.2 for the high-frequency dimension and 1.0 for the other dimensions is based on the distinguishability between historical fault samples and normal samples. In the labeled fault samples such as inter-turn short circuits and imbalances, the mean increase in the high-frequency harmonic dimension is 2.2 times that of the normal samples, while the mean increase in the low-frequency fundamental frequency and low-order harmonic dimensions is less than 1.2 times. Therefore, the high-frequency dimension scaling factor is fixed at 2.2 to match this mean multiple, while the other dimensions are kept at 1.0 to avoid amplifying irrelevant components. The reason for setting the overall scaling factor of 1.8 for the three vibration and shock channels is to align the numerical scale of the vibration channels before fusion with the scaled scale of the current high-frequency dimension. Specifically, the root mean square amplitude of the three-channel convolution output is 0.62, while the root mean square amplitude of the scaled current high-frequency dimension is 1.12. The ratio of 1.12 to 0.62 is divided to obtain 1.81, and an approximate value of 1.8 is taken as the fixed factor. The scaling factor of the trend dependence vector is set to 1.4 so that the output amplitude of the temperature change rate sequence after trend accumulation is on the same order of magnitude as the amplitude of the vibration mesoscale channel. In the historical sample statistics, the root mean square amplitude of the trend dependence vector is 0.50, and the root mean square amplitude of the vibration mesoscale channel is 0.70. Dividing 0.70 by 0.50 gives 1.40, and this is set to 1.4.

[0041] To determine the appropriate window length, experiments were conducted to compare the impact of different window lengths (3 points, 9 points, and 27 points) on the accuracy of equipment fault diagnosis. In the experiments, the equipment sampling frequency was 1 Hz, with data collected once per second. Spectral analysis and convolution calculations were performed on vibration acceleration signals, temperature change rates, and current spectra using 3-point, 9-point, and 27-point windows, respectively. First, when using a 3-point window, three consecutive data points of a vibration acceleration signal were selected. Using symmetric weights Convolution was performed on the signal, yielding an output value of 1.375. Calculations showed that the accuracy of the 3-point window was 80%, effectively capturing instantaneously changing signals, but it was insufficient in capturing periodic pulses and long-term trend changes.

[0042] Next, a 9-point window was used to capture periodic pulse signals. Assuming the nine consecutive data points of the vibration signal are 0.2, 0.1, 0.3, 2.5, 0.2, 0.1, 0.2, 0.3, and 0.4, convolution was performed using symmetric weights of 0.05, 0.10, 0.15, 0.20, 0.20, 0.15, 0.10, and 0.05, resulting in a convolution output value of 0.67. Experimental results show that the 9-point window can effectively capture periodically changing signals with an accuracy of 85%, exhibiting better stability compared to the 3-point window, especially in capturing periodic pulse signals.

[0043] Finally, a 27-point window was used to capture long-term trend signals. The 27-point window smooths out long-term changes in device signals, making it particularly suitable for capturing slow changes in temperature or device performance. For example, in signals with long-term trend changes, the convolution calculation of 27 data points effectively removes short-term fluctuations, accurately reflecting the gradual change in the device's state. Experiments showed that the 27-point window achieved an accuracy of 90%, performing best in capturing long-term trend changes in devices and effectively reflecting gradual changes in device status.

[0044] Experimental results show that the accuracy of fault diagnosis gradually increases with the increase of window length. A 3-point window can quickly respond to instantaneous impacts, but it is lacking in capturing periodic pulses and long-term trends, with an accuracy of 80%. A 9-point window is suitable for capturing periodically changing signals, with an accuracy of 85%. The 27-point window can most comprehensively capture different fault characteristics of the equipment, with an accuracy of 90%. Therefore, the 27-point window was ultimately determined to be the optimal window length, as it balances short-term changes and long-term trends in signals, providing the best fault diagnosis accuracy.

[0045] In step S15, feature weighted fusion is performed based on the multimodal fusion embedding vector to obtain a weighted original value vector. Based on the weighted original value vector, the multimodal fusion embedding vector is subjected to weighted incremental processing to obtain an enhanced embedding representation vector, including: Pooling calculations are performed on each channel in the multimodal fusion embedding vector to obtain the channel statistical descriptor for each channel; The channel statistical descriptor is subjected to a first set of linear transformations to obtain an intermediate vector. At the same time, the intermediate vector is subjected to a second linear transformation to obtain the original weight value vector. The original weight vector is compressed using a sigmoid function to obtain the final channel weights. By combining the final channel weights, each channel in the multimodal fusion embedding vector is weighted and adjusted to obtain the enhanced embedding representation vector.

[0046] It is worth noting that pooling calculations are performed on the vibration channel segment, temperature change rate channel segment, and current frequency component channel segment in the multimodal fusion embedding vector. Pooling calculations involve directly calculating the arithmetic mean of the values ​​of all dimensions within a channel segment to obtain a channel statistical value. For example, if the vibration channel segment contains 64 dimensions, the sum of these 64 dimensions is divided by 64 to obtain the vibration statistical value; if the temperature change rate channel segment contains 16 dimensions, the sum is divided by 16 to obtain the temperature statistical value; and if the current spectrum channel segment contains 32 dimensions, the sum is divided by 32 to obtain the current statistical value. The three are combined to form the channel statistical descriptor.

[0047] Next, two linear transformations are performed on the channel statistical descriptors, followed by superimposed nonlinear compression to obtain the channel weight vector. Specifically, the first set of linear weights is used to multiply the channel statistical descriptors term by term and sum them to obtain an intermediate vector. Then, terms less than 0 in the intermediate vector are set to 0 to remove negative contributions. Subsequently, the second set of linear weights is used to multiply the intermediate vector term by term and sum them to obtain the original weight values. Finally, a sigmoid function is used to compress the original weight values ​​to the range of 0 to 1, which are used as the contribution weights for each mode. For example, given the channel statistical descriptor [0.65, 0.42, 0.31], the first set of linear weights is [0.5, 0.3, 0.2], and the second set of linear weights is [0.515, 0.295, 0.191]. The calculated original weight value vector is [0.167, 0.037, 0.012]. After compression by the sigmoid function, the final channel weights are [0.542, 0.509, 0.503].

[0048] The first set of linear weights is quantified by calculating the correlation between each modal characteristic and the fault type. First, the correlation value between each mode and the fault is calculated. Then, the correlation value of each mode is compared with the sum of the total correlations, and weights are allocated proportionally. Specifically, the sum of the correlations of all modes is first calculated, and then the weight of each mode is determined based on the proportion of its correlation to the total correlation. For example, assuming the correlation between vibration signal and fault is 0.85, the temperature change rate is 0.65, the current signal is 0.45, and the total correlation is 1.95, then this proportion is rounded to two significant figures. The weight of the vibration signal is 0.436 (rounded to 0.44), the weight of the temperature change rate is 0.333 (rounded to 0.33), and the weight of the current signal is 0.231 (rounded to 0.23).

[0049] The second set of linear weights is adjusted based on the first set of weights and the variation amplitude of each modal feature. First, the variation amplitude of each mode is calculated according to the standard deviation changes of the vibration signal, temperature change rate, and current signal in historical data. For example, the variation amplitude of the vibration signal is 108%, the temperature change rate is 67%, and the current signal is 20%. Based on these variation amplitudes, amplification coefficients are assigned to each feature: 1.1 for the vibration signal, 1.05 for the temperature change rate, and 1.02 for the current signal. The adjusted weights are 0.55, 0.315, and 0.204, respectively. Next, these adjusted weights are normalized so that their sum is 1, resulting in the final second set of linear weights [0.515, 0.295, 0.191].

[0050] Subsequently, each parameter of each channel in the multimodal fusion embedding vector is multiplied by the corresponding channel weight in the final channel weight, thereby weighting and adjusting the features of each channel. For example, if the channel weights are vibration 0.542, temperature 0.509, and current 0.503, then when the original value of any dimension of the vibration segment is 0.50, the enhanced value is 0.50 multiplied by 0.542 to obtain 0.271; when the original value of any dimension of the temperature segment is 0.40, the enhanced value is 0.40 multiplied by 0.509 to obtain 0.204; when the original value of any dimension of the current segment is 0.90, the enhanced value is 0.90 multiplied by 0.503 to obtain 0.453. The remaining dimensions are multiplied dimension by dimension according to their respective channel weights to obtain the entire enhanced embedding representation vector.

[0051] In step S16, feature node construction is performed based on the enhanced embedding representation vector to obtain independent feature nodes in the semantic association graph. Then, association feature extraction is performed on these independent feature nodes to obtain semantic association features, including: The enhanced embedding representation vector is mapped to feature nodes in the semantic association graph; Calculate the mutual information between the feature nodes. If the mutual information is greater than a preset association threshold, establish connection edges based on the independent feature nodes to obtain a semantic topology structure. The semantic topology is input into a pre-trained graph convolutional neural network, which outputs semantic association features.

[0052] It is worth noting that the enhanced embedding representation vector is composed of multiple feature components, such as vibration signals, temperature, and current spectra. In steps S14 and S15, although the various parts of the vector undergo weighted increments and graph convolution processing, their structure remains unchanged. These processes primarily establish connections between features and adjust the weights of the features. Therefore, even in subsequent steps, the structure of the vector still consists of its individual feature components, and each feature component still represents its corresponding physical quantity.

[0053] The enhanced embedding representation vector is decomposed into multiple feature vectors according to feature categories, and these multiple feature vectors are mapped to feature nodes in the semantic association graph. For example, the enhanced embedding representation vector is [0.5,0.8,1.2,0.3,0.6,0.9,1.1], where the first three elements represent the features of the vibration signal, the next two elements represent the features of the temperature change rate, and the last two elements represent the features of the current spectrum. This vector is first decomposed into multiple feature vectors: [0.5,0.8,1.2] (vibration signal), [0.3,0.6] (temperature change rate), and [0.9,1.1] (current spectrum). Then, these feature vectors are mapped to the corresponding nodes in the semantic association graph, which is a feature relationship representation graph. The independent feature node corresponds to a single-dimensional feature sub-vector with clear physical meaning separated from the enhanced embedding representation vector, such as a vibration high-frequency energy node, a temperature trend node, and a current third harmonic amplitude node. Semantic association refers to the physical coupling or statistical correlation between these feature nodes, which is quantified by calculating the mutual information between node feature sequences and used as the basis for constructing the connection edges between nodes.

[0054] Next, the mutual information between these nodes is calculated. Mutual information quantifies the degree of information sharing between features by calculating the joint probability distribution and marginal probability distribution between them. First, the joint probability of each pair of features is calculated. and marginal probability The joint probability represents the probability that features X and Y occur simultaneously, while the marginal probability is the probability of each feature occurring individually. Mutual information is calculated using the following formula. ,in, Let p(x) be the mutual information between features X and Y, and p(Y) be the marginal probability of feature X. When calculating the mutual information between different nodes, the feature vector sequence of each node within a time window is treated as a sample of a random variable. The probability distribution is estimated using the histogram method and then substituted into the formula. For example, the joint probability between the peak factor and the root mean square value of the vibration signal is 0.4, the marginal probabilities are both 0.4, and the calculated mutual information is 0.3665. If the mutual information between two nodes is greater than the preset association threshold of 0.30, a connection edge is established between these two nodes. The mutual information of all node pairs in the semantic association graph is calculated, and connection edges are established between all node pairs with mutual information greater than the preset association threshold, thus forming the initial semantic topology.

[0055] The preset association threshold of 0.30 was set based on statistical analysis of historical data. Specifically, the mutual information of each pair of features was extracted from historical normal operation data, and its mean and standard deviation were calculated. Assuming that the mutual information of the vibration signal and other features follows a normal distribution, the calculated mean is 0.15 and the standard deviation is 0.075. To ensure that the threshold can cover most normal data, the threshold is set to the mean plus twice the standard deviation, i.e., 0.15 + 2 × 0.075 = 0.30. This threshold can guarantee that the relationship between 95% of normal samples is captured, while avoiding the erroneous establishment of connections between too many irrelevant features.

[0056] Subsequently, the initial semantic topology is input into a pre-trained graph convolutional neural network (GCN) to output semantic association features. The graph convolutional neural network structure is mainly divided into an input layer, a graph convolutional layer, and an output layer.

[0057] The input layer of a graph convolutional neural network contains initial features for each node of an initial semantic topology. Each node represents a feature (e.g., vibration signal, rate of temperature change, current spectrum, etc.), and these features are represented as feature vectors.

[0058] The graph convolutional layers of a graph convolutional neural network obtain a new feature representation for a node through linear transformation. A non-linear activation function, such as ReLU, Sigmoid, or Tanh, is applied to the new node features. The computation process of graph convolution is as follows: For each node... Its new features Calculated using the following formula, in, Let i be the new feature vector of node i. It is the feature vector of node i. Let i be the set of neighboring nodes. Let W be the set of neighboring nodes of node j, and W be the weight matrix. The graph convolution operation aggregates the information of the neighboring nodes onto the current node i. The formula only aggregates the features of the neighboring nodes. excluding its own node. .

[0059] The output layer of the graph convolutional neural network uses the Softmax activation function to process the features of each node. The Softmax function maps the feature vector of each node to the probabilities of different categories, such that the sum of the probabilities of each node belonging to a certain category is 1. Specifically, for the feature vector of each node, the Softmax function calculates the probability of each category and obtains the final classification result by calculating these probabilities; then, the model outputs semantic association features through a fully connected layer.

[0060] The graph convolutional neural network employs two graph convolutional layers, each followed by a ReLU activation function, and the output layer is a fully connected layer. It maps the latent features of each node to a new feature vector, which serves as the semantic association feature. The network training adopts an unsupervised reconstruction method, using the node feature reconstruction error as the loss function, so that the output features retain the semantic information of the original features, which is convenient for subsequent sparse reconstruction processing. If fault labels exist, supervised learning can also be used for training.

[0061] The training process of a Graph Convolutional Neural Network (GCN) relies on labeled graph-structured data as the training set samples. Node features are derived from sensor data, and the label of each node is used for supervised learning tasks. In forward propagation, the GCN processes node features layer by layer through graph convolutional layers, aggregating information from neighboring nodes into the current node. Next, a loss function (such as cross-entropy) is used to calculate the difference between the model output and the true label. Through backpropagation, the GCN updates the network weights based on the gradient information of the loss function, using optimization algorithms such as Adam for weight updates. Training continues until the loss function converges, i.e., the loss value changes steadily, or the preset maximum number of iterations is reached. During training, the convergence accuracy is set to... The maximum number of iterations is set to 500. During GCN training, the convergence accuracy is set to... This means that when the change in the loss function is less than this value, the model is considered to have reached sufficient optimization, and further training will have little effect on improving performance. Setting the maximum number of iterations to 500 ensures that the model can fully learn the features within a reasonable iteration range, avoiding excessive training time and wasted computational resources.

[0062] The training set samples are obtained from the device's sensor data. Features of different modes (such as vibration signals, temperature changes, current spectra, etc.) are mapped to nodes in a graph, and edges between nodes are established through physical or temporal correlations. Each node corresponds to a feature, and the edges represent the relationships between features. By pairing these features with the device's fault labels (such as normal, fault type, etc.), a labeled graph structure dataset is formed, which serves as the input for GCN training.

[0063] In step S17, sparse reconstruction extraction is performed based on the semantic association features to obtain the repeated impact interval component; and residual impact removal is performed on the repeated impact interval component to obtain a second residual vector. Based on the second residual vector, spectral clustering is performed to obtain an abnormal feature set, including: Dictionary learning is performed on the semantic association features to obtain dictionary basis vectors, and sparse encoding is performed on the dictionary basis vectors to obtain a sparse feature matrix; Based on the sparse feature matrix, sparse coefficient groups with equal intervals of repeated activation are selected, and the dictionary basis vectors of the sparse coefficient groups are retained. Matrix multiplication reconstruction is then performed to obtain the repeated impact interval component. Subtract the repeated impact interval component from each of the semantic association features to obtain the first residual vector, and then determine the impact point on the first residual vector to obtain the second residual vector. The second residual vector is used to calculate the spectral energy proportion to obtain the broadband noise rise component. Simultaneously, the second residual vector is clustered into low-amplitude clusters according to a preset low-amplitude range to obtain low-amplitude impact clusters. The broadband noise uplift component is encapsulated with the low-amplitude impact cluster to obtain an abnormal feature set.

[0064] It's worth noting that the dictionary learning method (K-SVD) is used to extract regular features from semantic association features, removing noise and redundant information. Specifically, K-SVD learns a dictionary to ensure that each sample of the signal can be accurately represented by a few basic elements in the dictionary. First, K-SVD initializes the dictionary, generating initial elements randomly. Then, for each signal sample, K-SVD uses the current dictionary to sparsely encode it, decomposing the signal sample into a linear combination of dictionary bases, where most coefficients are zero, achieving the goal of sparse representation. For example, if the fundamental frequency of the periodic impact pattern, a semantic association feature, is 50Hz and the second harmonic amplitude is 10A, then dictionary learning can extract these two features and clearly separate them from the vibration signal, forming a new sparse representation. Finally, through this sparse representation method, the periodic impact pattern is accurately extracted, thereby removing redundant information and noise, resulting in a clear sparse feature matrix.

[0065] Subsequently, the repeating impact interval component is extracted from the semantic association features using a sparse reconstruction method. Specifically, the semantic association features are sparsely encoded using a sparse basis matrix to obtain the coefficient set of each feature sequence on each basis vector. Then, the set of coefficients with an equal-interval repetition pattern is selected from the coefficient set, and only the basis vectors corresponding to this set of coefficients are retained for reconstruction. The remaining coefficients are set to zero, and a matrix multiplication reconstruction is performed again. The reconstruction result is the repeating impact interval component. For example, if the semantic association features form a sequence of length 8 [1.20,0.10,0.05,0.90,0.08,0.06,1.10,0.09] within a certain time window, the coefficients of the three basis vectors obtained after sparse encoding are as follows: the coefficients of the first basis vector are 0.80, 0.75, and 0.78 at positions 1, 4, and 7, and 0 at the other positions; the coefficients of the second basis vector are 0.20 and 0.18 at positions 2 and 5, and 0 at the other positions; and the coefficients of the third basis vector are 0.15 and 0.14 at positions 3 and 6, and 0 at the other positions. Since the non-zero coefficients of the first basis vector appear at positions 1, 4, and 7, with an interval of 3 sampling points and repeated occurrences, it is determined to be the basis vector corresponding to the repeated impact interval. Only the coefficients of the first basis vector are retained, and the coefficients of the second and third basis vectors are cleared to zero. Then, the sparse basis matrix is ​​multiplied with the retained coefficients to reconstruct the repeated impact interval components as [0.80,0,0,0.75,0,0,0.78,0].

[0066] Next, the difference between the repeated impact interval component and the semantic association feature is calculated to obtain the first residual vector. For example, the semantic association feature is [1.20,0.10,0.05,0.90,0.08,0.06,1.10,0.09], and the repeated impact interval component obtained by sparse reconstruction is [0.80,0,0,0.75,0,0,0.78,0]. The first residual vector is calculated by subtracting the repeated impact interval component from the semantic association feature, resulting in [0.40,0.10,0.05,0.15,0.08,0.06,0.32,0.09].

[0067] When the magnitude of an impact at a certain position in the first residual vector minus the magnitude of its adjacent position exceeds a preset impact threshold, it is determined to be an asymmetric impact point. The asymmetric impact point is then set to zero in the first residual vector to obtain the second residual vector. For example, the first residual vector is [0.40, 0.10, 0.05, 0.15, 0.08, 0.06, 0.32, 0.09], and the preset impact threshold is 0.20. It is found that the differences between positions 1 and 7 and their adjacent positions both exceed 0.20; therefore, positions 1 and 7 are determined to be impact points and set to zero. Finally, the second residual vector is obtained as [0, 0.10, 0.05, 0.15, 0.08, 0.06, 0, 0.09].

[0068] The preset impact threshold is set based on the amplitude distribution of residual vectors in historical data. Residual vectors are extracted from historical data, and the mean and standard deviation of their amplitudes are calculated. It is found that the residual amplitudes follow a normal distribution with a mean of 0.1 and a standard deviation of 0.05. Therefore, based on the characteristics of the normal distribution, the impact threshold is set to the mean plus twice the standard deviation. This ensures that approximately 95% of normal data falls within the threshold range.

[0069] Subsequently, the remaining energy components in the second residual vector are split according to frequency band and amplitude clusters, and the splitting results are encapsulated by category. Specifically, the spectrum of the second residual vector is first calculated by performing a discrete Fourier transform on the vector within a preset time window to obtain the amplitude at each frequency point, and then using the sum of squares of the amplitudes across the entire frequency band as the total energy. When continuous rises occur in both the high-frequency and low-frequency bands and the energy distribution is not concentrated in a few frequency points, this part is marked as a broadband noise rise component, and its energy proportion and corresponding frequency band range are used as the parameters of this component. Then, low-amplitude impact clusters are extracted from the second residual vector in the time domain. The extraction method is to first set an upper and lower limit for the amplitude threshold, and group impact points with amplitudes within this range and densely appearing in a short period of time into the same cluster. The parameters of the cluster are represented by the number of impact points, the duration of the cluster, and the maximum amplitude within the cluster. For example, the second residual vector is [0, 0.10, 0.05, 0.15, 0.08, 0.06, 0, 0.09]. Performing a Discrete Fourier Transform on this vector yields the spectral amplitudes [0.15, 0.18, 0.20, 0.28, 0.34, 0.28, 0.22, 0.18]. Summing the squares of the amplitudes at each frequency point gives a total energy of 0.4481. Then, summing the squares of the corresponding amplitudes (0.34, 0.28, 0.22, 0.18) in the high-frequency band gives a high-frequency energy of 0.2748. The proportion of high-frequency energy is 0.2748 divided by 0.4481, resulting in 0.61, which is higher than... When the preset noise line is 0.50, this component is recorded as a broadband noise rise component. Then, in the time domain, the low amplitude range is set to 0.05 to 0.15. Points of 0.10, 0.05, 0.15, 0.08, 0.06 and 0.09 that fall within this range are screened out from the second residual vector and clustered into a cluster according to their adjacent positions to obtain a low amplitude impact cluster. The maximum amplitude within the cluster is 0.15 and contains 6 impact points. Finally, the repeated impact interval component obtained in the previous step, the impact sequence obtained after elimination, and the broadband noise rise component identified in this step are encapsulated together with the low amplitude impact cluster to form the current abnormal mode set of the bearing.

[0070] The preset noise threshold of 0.50 is set based on the statistical results of the high-frequency energy ratio under historical normal operating conditions. The high-frequency energy ratio is calculated segment by segment from the second residual vector sample that was marked as normal operation in history to form a set of ratio samples. The mean and standard deviation of the set are calculated. After confirming that the ratio samples approximately follow a normal distribution, the threshold is set to the mean plus twice the standard deviation to cover about 95% of the normal samples. For example, the mean of the high-frequency energy ratio of the normal samples is 0.38 and the standard deviation is 0.06. The mean plus twice the standard deviation is 0.38 plus 0.12 to get 0.50. Thus, 0.50 is used as the broadband noise rise judgment line.

[0071] The low amplitude range of 0.05 to 0.15 is defined based on the statistical distribution of impact point amplitudes from historical samples. First, impact point amplitude sets are extracted from normal samples and early minor damage samples, and their mean and standard deviation are calculated respectively. The low amplitude impact cluster is defined as the amplitude range between the upper limit of normal background fluctuations and the lower limit of significant impacts. For example, the mean of normal background fluctuation amplitude is 0.04 and the standard deviation is 0.005. Then, the upper limit of normal is the mean plus twice the standard deviation, which is 0.04 plus 0.01, resulting in 0.05. The mean of significant impact amplitude in historical fault samples is 0.18 and the standard deviation is 0.015. Then, the lower limit of significant impact is the mean minus twice the standard deviation, which is 0.18 minus 0.03, resulting in 0.15. Therefore, 0.05 to 0.15 is set as the low amplitude impact identification range.

[0072] Finally, the repeated impact interval component obtained in the previous sequence, the impact sequence extracted under the premise of setting the impact point to zero, and the broadband noise rise component and low amplitude impact cluster identified from the second residual vector are uniformly encapsulated into an abnormal mode set.

[0073] In step S18, based on the set of abnormal features, a fault type determination is performed to obtain the fault type, including: Calculate the cosine of the angle between the abnormal feature set and each fault feature template vector in the preset fault feature template library; When the cosine value of the included angle is greater than a preset similarity threshold, the defect type to which the fault feature template vector belongs is taken as the fault type.

[0074] It is worth noting that the cosine value of the angle between the abnormal pattern set and each template in the fault feature template library is calculated. When the cosine value is greater than a preset similarity threshold, it can be confirmed that the current abnormal pattern is highly correlated with a certain type of fault, thereby marking the fault type. For example, if an abnormal pattern vector in the abnormal pattern set is A=[0.60,0.20,0.10,0.10], and the "inner ring defect" template vector in the fault feature template library is B=[0.55,0.25,0.10,0.10], the cosine value of the angle is calculated by dividing the vector dot product by the product of the magnitudes of the two vectors. The cosine value is 0.996, which is greater than the preset similarity threshold of 0.52, thus marking the abnormal pattern as an inner ring defect fault type.

[0075] The preset similarity threshold is set based on the statistical distribution of cosine similarity under historical normal operating conditions. Specifically, a set of abnormal patterns historically marked as normal operation is selected as the sample set. The cosine value of the angle between each sample and each template in the fault feature template library is calculated, and the maximum cosine similarity for each sample is recorded, forming a maximum similarity sample set. The mean and standard deviation of this sample set are calculated. After finding that it approximately follows a normal distribution, the similarity threshold is set to the mean plus twice the standard deviation to cover approximately 95% of the upper bound of normal samples. For example, if the mean maximum cosine similarity of historical normal samples is 0.40 and the standard deviation is 0.06, then the mean plus twice the standard deviation equals 0.40 plus 0.12, resulting in 0.52. Therefore, 0.52 is set as the preset similarity threshold.

[0076] The fault feature template library is constructed from standard data, which comes from online monitoring data of confirmed defect types in historical maintenance records, or sample data manually prepared under the test bench according to the location and degree of defects. For each type of defect, no less than a preset number of sample segments are collected, and the abnormal pattern extraction steps consistent with the online process are performed on the sample segments to obtain abnormal pattern vectors of the same dimension. Then, the arithmetic mean of multiple abnormal pattern vectors of the same defect type is calculated according to the dimension to obtain the template vector of the defect type. The template vector is then normalized to ensure the scale consistency of the subsequent cosine similarity calculation. The template vectors of each defect type are indexed into the library according to the defect label, thus forming a fault feature template library containing inner ring defect harmonic order templates, outer ring repetitive impact interval templates, rolling element asymmetric sequence templates, and early pitting low-amplitude impact cluster templates.

[0077] In step S19, amplitude mapping is performed to determine the damage level based on the fault type, and periodic retrieval is performed based on the damage level to obtain complete early warning information, including: Based on the fault type, locate the target location and obtain the vibration amplitude at the target location; Based on the vibration amplitude, the degree of damage of the fault type is determined; Based on the degree of damage and in conjunction with a preset rule table, the required inspection cycle for the fault type is determined. Based on the target location, the degree of damage, and the inspection cycle, a complete early warning message is generated.

[0078] It's worth noting that the target location is determined based on the fault type; for example, if the fault type is an inner ring defect, the location is the inner ring raceway surface. Subsequently, the peak value at the target location is extracted from the bearing's vibration signal as the characteristic amplitude. For instance, suppose that in a segment of vibration signal data, the vibration signal on the bearing's inner ring raceway surface is [0.1, 0.3, 0.15, 0.18, 0.1]. In this signal, the maximum vibration value is 0.18g. Therefore, 0.18g is the characteristic amplitude at this target location.

[0079] Next, the degree of damage is defined based on the characteristic amplitude. The degree of damage is divided into amplitude ranges according to a set standard. For example, a characteristic amplitude between 0.1g and 0.3g is defined as moderate damage. Based on a characteristic amplitude of 0.18g, the degree of damage for this fault type is moderate damage. Below 0.1g is defined as mild damage, and above 0.3g is defined as severe damage.

[0080] Next, the recommended inspection cycle is retrieved based on the degree of damage. The preset rule table is as follows: minor damage is recommended to be inspected every 30 days, moderate damage is recommended to be inspected every 15 days, and severe damage is recommended to be inspected weekly. Based on the degree of damage being moderate, the retrieved inspection cycle is once every 15 days.

[0081] Finally, combining the defect location information "inner raceway surface", the damage level "moderate damage", and the recommended inspection cycle "every 15 days", a complete early warning message is generated. For example, an inner race defect with a characteristic amplitude of 0.18g indicates moderate damage, and an inspection every 15 days is recommended. This information provides detailed guidance for subsequent maintenance and fault diagnosis.

[0082] The amplitude range of damage severity is defined based on the quantiles of historical samples and failure verification records. First, historical samples of the same defect type and at the same measuring point location are divided into three groups—mild, moderate, and severe—according to disassembly inspection or bench wear levels. For each group, the characteristic amplitude peak value at the target location is extracted within a unified time window, forming three amplitude sequences. Then, each amplitude sequence is sorted from smallest to largest. The 95th percentile of the mild group is taken as the upper limit for mild damage, and the 5th percentile of the severe group is taken as the lower limit for severe damage. The interval between these two is defined as the moderate damage interval. The resulting threshold is directly provided by the sample ranking and does not depend on distribution assumptions. For example, the peak amplitude values ​​extracted from the mild group samples were 0.05g, 0.06g, 0.07g, 0.08g, 0.09g, and 0.10g, respectively. The value corresponding to the 95th percentile was taken as 0.10g as the upper limit of mild damage. The peak amplitude values ​​extracted from the severe group samples were 0.30g, 0.32g, 0.35g, 0.38g, and 0.40g, respectively. The value corresponding to the 5th percentile was taken as 0.30g as the lower limit of severe damage. Thus, less than 0.10g was defined as mild damage, 0.10g to 0.30g was defined as moderate damage, and greater than 0.30g was defined as severe damage. The interval boundaries were derived from the ranking percentile of the peak amplitude values ​​of historical samples and corresponded one-to-one with the disintegration or test bench level labels.

[0083] The recommended inspection cycle is based on historical degradation rate statistics and operational tolerance risk thresholds. Specifically, it involves calculating the growth rate of statistical characteristic amplitudes of samples at different damage levels over time, and using the average time required for the amplitude to jump from the current level to the next as the upper limit of the reference cycle. A safety factor is then introduced to compress the inspection cycle to half or less of the jump time to avoid missed detections. For example, if the average jump time for a mild sample from 0.06g to 0.10g is 60 days, the inspection cycle is 30 days; if the average jump time for a moderate sample from 0.18g to 0.30g is 30 days, the inspection cycle is 15 days; and if a severe sample shows a significant jump within 7 days accompanied by a significant increase in the proportion of abnormal temperature rise, the inspection cycle is 7 days, forming a rule table for mild cases every 30 days, moderate cases every 15 days, and severe cases weekly.

[0084] In summary, this invention discloses an intelligent monitoring method based on AI multimodal and IoT, which can solve the problem of false alarms and missed alarms caused by threshold drift under load conditions.

[0085] Reference Figure 2 The second embodiment of the present invention provides an intelligent monitoring system based on AI multimodal and IoT, comprising: The data acquisition module is used to acquire the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor. The data processing module is used to perform time-series synchronization processing on the temperature scalar sequence and current frequency component based on the vibration acceleration sequence to obtain a multidimensional signal matrix, and to perform segmented indexing and encapsulation processing on the multidimensional signal matrix to obtain the original data set. The standardization module is used to extract temperature features from the original dataset to obtain the temperature change rate, and to standardize the temperature change rate to obtain a standardized feature set. The embedding vector module is used to construct a three-channel feature map based on the standardized feature set, and to perform recursive calculation based on the three-channel feature map to obtain a trend dependency vector sequence; and to perform spectral weighted fusion processing based on the trend dependency vector sequence to obtain a multimodal fusion embedding vector. The enhancement vector module is used to perform feature weighted fusion based on the multimodal fusion embedding vector to obtain a weighted original value vector, and to perform weighted incremental processing on the multimodal fusion embedding vector based on the weighted original value vector to obtain an enhanced embedding representation vector. The associated feature module is used to perform feature node construction processing based on the enhanced embedding representation vector to obtain independent feature nodes in the semantic association graph, and to perform associated feature extraction processing on the independent feature nodes to obtain semantic association features. An anomaly feature module is used to perform sparse reconstruction extraction based on the semantic association features to obtain a repeated impact interval component; and to perform residual impact removal on the repeated impact interval component to obtain a second residual vector; and to perform spectral clustering encapsulation based on the second residual vector to obtain an anomaly feature set. The fault determination module is used to determine the fault type based on the set of abnormal features to obtain the fault type. The early warning generation module is used to perform amplitude mapping and classification based on the fault type to obtain the degree of damage, and to perform periodic retrieval and generation based on the degree of damage to obtain complete early warning information.

[0086] It should be noted that the intelligent monitoring system based on AI multimodal and IoT provided in this embodiment of the invention is used to execute all the process steps of the intelligent monitoring method based on AI multimodal and IoT in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.

[0087] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.

[0088] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for intelligent monitoring based on AI multimodal and IoT, characterized in that, include: Obtain the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor; Based on the vibration acceleration sequence, the temperature scalar sequence and the current frequency component are subjected to time synchronization processing to obtain a multidimensional signal matrix. The multidimensional signal matrix is ​​then segmented, indexed, and encapsulated to obtain the original data set. The original dataset is subjected to temperature feature extraction processing to obtain the temperature change rate, and the temperature change rate is then standardized to obtain a standardized feature set. Based on the standardized feature set, a three-channel feature map is constructed, and a trend dependency vector sequence is obtained by recursive calculation based on the three-channel feature map; based on the trend dependency vector sequence, a spectral weighted fusion process is performed to obtain a multimodal fusion embedding vector. Based on the multimodal fusion embedding vector, feature weighted fusion is performed to obtain a weighted original value vector. Based on the weighted original value vector, the multimodal fusion embedding vector is subjected to weighted incremental processing to obtain an enhanced embedding representation vector. Based on the enhanced embedding representation vector, feature node construction processing is performed to obtain independent feature nodes in the semantic association graph, and association feature extraction processing is performed on the independent feature nodes to obtain semantic association features. Based on the semantic association features, sparse reconstruction extraction is performed to obtain the repeated impact interval component; and residual impact removal is performed on the repeated impact interval component to obtain the second residual vector. Based on the second residual vector, spectral clustering is performed to encapsulate the result and obtain the abnormal feature set. Based on the set of abnormal features, the fault type is determined to obtain the fault type; Based on the fault type, amplitude mapping is performed to determine the damage level, and based on the damage level, periodic retrieval is performed to generate complete early warning information.

2. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The process involves using the vibration acceleration sequence as a reference, performing time-series synchronization processing on the temperature scalar sequence and the current frequency component to obtain a multidimensional signal matrix, and then performing segmented indexing and encapsulation processing on the multidimensional signal matrix to obtain the original data set, including: Using the vibration acceleration sequence as a reference, the temperature scalar sequence is linearly interpolated to obtain an aligned temperature value sequence; Based on the vibration acceleration sequence, the current frequency components are subjected to FIR interpolation to obtain an aligned current value sequence. The vibration acceleration sequence, the aligned temperature value sequence, and the aligned current value sequence are concatenated column by column to obtain a time-synchronized multidimensional signal matrix; The multidimensional signal matrix is ​​segmented according to a preset time window to obtain storage data blocks, and a preset spatiotemporal correlation index table is added to the storage data blocks to obtain the original data set.

3. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The original dataset is subjected to temperature feature extraction processing to obtain the temperature change rate, and the temperature change rate is then standardized to obtain a standardized feature set, including: The absolute peak value and root mean square value are calculated for the vibration acceleration sequence of the original dataset, and the peak factor is calculated using the absolute peak value and the root mean square value. The temperature scalar sequence of the original dataset is calculated with a weighted moving average temperature value according to a preset feature extraction window, and the difference between the weighted moving average temperature values ​​adjacent to the preset feature extraction window is calculated to obtain the temperature change rate. The continuous interval where the temperature change rate is less than the preset temperature stability threshold is taken as the stable operating condition period. During the stable operating condition period, the aligned current value sequence of the original data set is sliced ​​according to the preset spectrum analysis duration to perform spectrum calculation, so as to obtain the fundamental frequency amplitude, harmonic component energy ratio and sideband characteristic amplitude. The peak factor, root mean square value, temperature change rate, fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude are summarized into a feature vector. The feature vector is then subjected to Z-score standardization based on the mean and standard deviation of historical normal data to obtain a standardized feature set.

4. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The process involves constructing a three-channel feature map based on the standardized feature set, and performing recursive calculations based on the three-channel feature map to obtain a trend dependency vector sequence; then, based on the trend dependency vector sequence, performing spectral weighted fusion processing to obtain a multimodal fusion embedding vector, including: The fundamental frequency amplitude, harmonic component energy ratio, and sideband characteristic amplitude of the standardized feature set are extracted to form the current standardized spectrum features; Peak factors and root mean square values ​​are extracted from the standardized feature set to form a vibration array. The vibration array is then slid along the time axis with three preset window lengths to form three sets of sliding window data. The three sets of sliding window data are multiplied and summed point by point according to preset weights to obtain the convolution output sequence of the three sets of sliding window data. The convolution output sequence is then stacked by channel at the same timestamp to obtain a three-channel feature map. The temperature change rate sequence is read based on the timestamp of the three-channel feature map, and the temperature change rate sequence is recursively calculated in chronological order to obtain the trend dependence vector sequence. The high-frequency harmonic energy ratio is calculated based on the trend dependency vector of the current normalized spectrum feature. When the high-frequency harmonic energy ratio exceeds a preset baseline, the current normalized spectrum feature is dimensionally scaled to obtain the scaled current normalized spectrum feature. The three-channel feature map, the trend-dependent vector sequence, and the scaled current-normalized spectral features are weighted, fused, and compressed to obtain a multimodal fusion embedding vector.

5. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The step of performing feature weighted fusion based on the multimodal fusion embedding vector to obtain a weighted original value vector, and then performing weighted incremental processing on the multimodal fusion embedding vector based on the weighted original value vector to obtain an enhanced embedding representation vector, includes: Pooling is performed on each channel in the multimodal fusion embedding vector to obtain the channel statistical descriptor for each channel; The channel statistical descriptor is subjected to a first set of linear transformations to obtain an intermediate vector, and the intermediate vector is subjected to a second linear transformation to obtain the original weight value vector. The original weight vector is compressed using an S-shaped function to obtain the final channel weights. By combining the final channel weights, the weighted adjustment of each channel in the multimodal fusion embedding vector is performed to obtain the enhanced embedding representation vector.

6. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The step involves constructing feature nodes based on the enhanced embedding representation vector to obtain independent feature nodes in the semantic association graph, and then extracting association features from these independent feature nodes to obtain semantic association features, including: The enhanced embedding representation vector is mapped to feature nodes in the semantic association graph; Calculate the mutual information between the feature nodes. If the mutual information is greater than a preset association threshold, establish connection edges based on the independent feature nodes to obtain a semantic topology structure. The semantic topology is input into a pre-trained graph convolutional neural network, which outputs semantic association features.

7. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The process involves sparse reconstruction extraction based on the semantic association features to obtain a repeated impact interval component; residual impact removal is then performed on the repeated impact interval component to obtain a second residual vector; and spectral clustering is performed based on the second residual vector to obtain an anomaly feature set, including: Dictionary learning is performed on the semantic association features to obtain dictionary basis vectors, and sparse encoding is performed on the dictionary basis vectors to obtain a sparse feature matrix; Based on the sparse feature matrix, sparse coefficient groups with equal intervals of repeated activation are selected, and the dictionary basis vectors of the sparse coefficient groups are retained. Matrix multiplication reconstruction is then performed to obtain the repeated impact interval component. Subtract the repeated impact interval component from each of the semantic association features to obtain the first residual vector, and then determine the impact point on the first residual vector to obtain the second residual vector. For the second residual vector, the spectral energy ratio is calculated to obtain the broadband noise rise component. At the same time, the second residual vector is clustered into low-amplitude clusters according to the preset low-amplitude range to obtain low-amplitude impact clusters. The broadband noise uplift component is encapsulated with the low-amplitude impact cluster to obtain an abnormal feature set.

8. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The step of determining the fault type based on the set of abnormal features to obtain the fault type includes: Calculate the cosine of the angle between the abnormal feature set and each fault feature template vector in the preset fault feature template library; When the cosine value of the included angle is greater than a preset similarity threshold, the defect type to which the fault feature template vector belongs is taken as the fault type.

9. The intelligent monitoring method based on AI multimodal and IoT according to claim 1, characterized in that, The step involves performing amplitude mapping and classification based on the fault type to obtain the damage level, and then performing periodic retrieval and generation based on the damage level to obtain complete early warning information, including: Based on the fault type, locate the target location and obtain the vibration amplitude at the target location; Based on the vibration amplitude, the degree of damage of the fault type is determined; Based on the degree of damage and in conjunction with a preset rule table, the required inspection cycle for the fault type is determined. Based on the target location, the degree of damage, and the inspection cycle, a complete early warning message is generated.

10. An intelligent monitoring system based on AI multimodal and IoT, characterized in that, include: The data acquisition module is used to acquire the vibration acceleration sequence, temperature scalar sequence, and current frequency component of the motor. The data processing module is used to perform time-series synchronization processing on the temperature scalar sequence and current frequency component based on the vibration acceleration sequence to obtain a multidimensional signal matrix, and to perform segmented indexing and encapsulation processing on the multidimensional signal matrix to obtain the original data set. The standardization module is used to extract temperature features from the original dataset to obtain the temperature change rate, and to standardize the temperature change rate to obtain a standardized feature set. The embedding vector module is used to construct a three-channel feature map based on the standardized feature set, and to perform recursive calculation based on the three-channel feature map to obtain a trend dependency vector sequence; and to perform spectral weighted fusion processing based on the trend dependency vector sequence to obtain a multimodal fusion embedding vector. The enhancement vector module is used to perform feature weighted fusion based on the multimodal fusion embedding vector to obtain a weighted original value vector, and to perform weighted incremental processing on the multimodal fusion embedding vector based on the weighted original value vector to obtain an enhanced embedding representation vector. The associated feature module is used to construct feature nodes based on the enhanced embedding representation vector to obtain independent feature nodes in the semantic association graph, and to extract associated features from the independent feature nodes to obtain semantic association features. An anomaly feature module is used to perform sparse reconstruction extraction based on the semantic association features to obtain a repeated impact interval component; and to perform residual impact removal on the repeated impact interval component to obtain a second residual vector; and to perform spectral clustering encapsulation based on the second residual vector to obtain an anomaly feature set. The fault determination module is used to determine the fault type based on the set of abnormal features to obtain the fault type; The early warning generation module is used to perform amplitude mapping and classification based on the fault type to obtain the degree of damage, and to perform periodic retrieval and generation based on the degree of damage to obtain complete early warning information.