Convolutional neural network gearbox fault diagnosis method based on attention mechanism

By using a convolutional neural network method based on an attention mechanism, the problem of complex fault mode recognition under multiple operating conditions in gearbox fault diagnosis was solved, achieving more stable fault type recognition and clearer feature attribution judgment.

CN121834545APending Publication Date: 2026-04-10CRRC IND INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient for identifying complex fault modes under various operating conditions in gearbox fault diagnosis. They have fixed feature representation dimensions, coarse response area coverage, and are prone to ignoring subtle changes. The classification output is slow to respond to changes in input, resulting in significant lag and risk of misjudgment, leading to confusion in fault type labels and ambiguity in identification.

Method used

A convolutional neural network method based on attention mechanism is adopted. By acquiring gearbox vibration data, the frame segments are divided according to the meshing period, the amplitude response of the main frequency position is extracted, and the edge region is located by combining the attention distribution map. A multi-channel structure attribution index table is constructed to track the fault type identification results.

Benefits of technology

It enhances the ability to extract feature differences under multiple operating conditions, improves the ability to distinguish fault attribution and the stability of classification results, reduces the risk of misjudgment, and improves the clarity of fault mode identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834545A_ABST
    Figure CN121834545A_ABST
Patent Text Reader

Abstract

The invention discloses a convolutional neural network gearbox fault diagnosis method based on an attention mechanism, and relates to the technical field of convolutional neural networks, and the method specifically comprises the following steps: obtaining vibration data, segmenting the vibration data into frame segments according to a meshing period, extracting a dominant frequency amplitude mapping time sequence, matching attention region coordinates to form a response frame segment focusing region list; and positioning layer edges to construct a boundary feature structure set, generating a multi-channel structure attribution index table in combination with spatial positions, and outputting a gearbox fault type identification result by associating tags with frame segments. According to the invention, through mapping of dominant frequency response and time sequence, connection of amplitude change and attention area coordinates, construction of a frame segment focusing area, association of a layer boundary and a gear ring position, and extraction of a multi-channel response structure with a consistent spatial position, a mapping relation between a frame segment and a label identifier is formed, and the extraction capability of differences under multiple working conditions is enhanced. And the distinguishing capability of feature affiliation judgment and the stability of a classification result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of convolutional neural network, in particular to a convolutional neural network gearbox fault diagnosis method based on attention mechanism. BACKGROUND

[0002] The technical field of convolutional neural network belongs to one of the deep learning methods in artificial intelligence, mainly including convolution operation, feature map generation, pooling operation, nonlinear activation function and full connection output and other core matters, through constructing multi-layer network structure to extract features and classify and distinguish the input structured data, widely applied in image recognition, speech processing, industrial fault detection and other fields, the methodological path of this technical field usually includes format processing of data input layer, local feature extraction of convolution layer, feature dimension compression of pooling layer, nonlinear expression ability of activation function, and the determination result of full connection layer output, and the network parameters are trained and optimized through the back propagation algorithm, wherein the traditional gearbox fault diagnosis method refers to collecting the vibration signal of the gearbox in the industrial transmission system as the input data, determining the frequency domain or time-frequency domain analysis method such as fast Fourier transform, wavelet packet decomposition or envelope spectrum analysis according to experience to extract feature indexes, selecting energy spectrum, root mean square value or kurtosis as classification basis, and then inputting into the support vector machine or artificial neural network classifier to complete fault type determination, which depends on manual selection of features and classifier performance, and is difficult to cope with complex fault pattern recognition under multi-working condition environment.

[0003] The prior art highly depends on fixed parameter setting and unified processing flow in the process of vibration signal feature extraction, and it is difficult to complete stable recognition in the face of abnormal signals with dramatic frequency fluctuations or short response duration, the feature expression dimension is fixed, the response region is roughly covered, and the details of the associated changes are easily ignored, in the scene of multiple fault superposition or working condition interlacing, the classification output reacts slowly to the input changes, there is significant lag and misjudgment risk, which easily causes label confusion and response position recognition ambiguity of fault type, limits the overall recognition efficiency and the discrimination of fault mode output, and is not conducive to realizing structure attribution tracking and clear display of fault path. SUMMARY

[0004] In order to solve the above problems existing in the prior art, the embodiment of the present application provides a convolutional neural network gearbox fault diagnosis method based on attention mechanism, and the specific technical scheme is as follows: A convolutional neural network gearbox fault diagnosis method based on attention mechanism, comprising the following steps: S1: obtaining the vibration data sequence of the gearbox under steady state working condition, dividing into frame segments according to the meshing period, extracting the amplitude response of the main frequency position in the frame segment, corresponding the frame segment and the frequency band position in the time sequence, and obtaining a period response frame segment index group; S2: based on the feature map content corresponding to the frame segment in the periodic response frame segment index group, extracting the response region coordinates in time sequence, and obtaining a response frame segment focusing region list by matching the coordinates with the corresponding region position in the attention distribution map; S3: based on the response frame segment focusing region list, extracting the mutation region in the convolution channel map and positioning the edge, and then matching the output shaft meshing gear response position, associating the graph layer coordinates with the frame segment time, and obtaining a boundary feature structure set; S4: based on the position of the response region in the channel map in the boundary feature structure set, searching for the region with the same spatial position, combining the corresponding graph layer position and time point, and obtaining a multi-channel structure attribution index table; S5: based on the multi-channel structure attribution index table, tracking the corresponding region position and label of the attention output graph, and obtaining a gear box fault type identification result corresponding to the label position and frame segment time.

[0005] As a further scheme of the present application, the periodic response frame segment index group includes a frequency band position, a response amplitude, and a time sequence mapping, the response frame segment focusing region list includes an amplitude change position, an attention response region, and a time sequence label, the boundary feature structure set includes a graph layer edge coordinate, a meshing gear boundary, and a frame segment time point, the multi-channel structure attribution index table includes a channel graph number, a graph layer response position, and a time point record, and the gear box fault type identification result includes a label identification, a label number mapping times, and a label content.

[0006] As a further scheme of the present application, the meshing period indicates the time interval required for the gear to complete meshing during operation. The amplitude response refers to the response strength amplitude of the vibration signal at each frequency point extracted in the frequency band range, which is an energy form of the main frequency region.

[0007] As a further scheme of the present application, the meshing gear response position refers to the response boundary coordinate point corresponding to the output shaft meshing gear in the convolution channel map according to boundary feature recognition and positioning. The attention output graph refers to a visual heat map used by a neural network to indicate attention at a predetermined position, which displays the response strength of a key feature region.

[0008] As a further scheme of the present application, the specific steps of S1 are: S101: obtaining a vibration data sequence of the gear box under steady state working condition, extracting each group of continuous vibration data in time sequence, identifying a gear meshing period signal segment, dividing a period length corresponding time range, and obtaining a meshing period time segment interval group; S102: based on the time range content in the meshing period time interval group, the vibration data stream is cut based on each period as a boundary, the main frequency component corresponding band in the time interval is extracted, and the amplitude data under the band position is associated with the frame segment index to obtain a main frequency amplitude frame segment corresponding relationship set; S103: based on each group of data in the main frequency amplitude frame segment corresponding relationship set, the band position and the frame segment index are mapped in a unified timeline, the mapped data is aggregated, and a period response frame segment index group is obtained.

[0009] As a further scheme of the application, the specific steps of S2 are: S201: based on the frame segment position in the period response frame segment index group, the channel coordinates of the position of the amplitude change point in the feature map are extracted, the channel index, spatial position and time point of the coordinates are associated, and the processing result is stored in a sequence structure to obtain an amplitude change coordinate sequence; S202: based on each group of coordinates in the amplitude change coordinate sequence, the response position of the corresponding channel in the attention distribution map is searched, the coordinate points with the same spatial position are extracted, and the time point and the layer channel index are compared to obtain a consistent coordinate sequence; S203: based on the coordinate data in the consistent coordinate sequence, the associated frame segment time and layer channel index are extracted, the position, time information of each group of data and the response area in the graph are written into the sequence structure, and the frame segment set is expanded according to the time sequence to obtain a response frame segment focus area list.

[0010] As a further scheme of the application, the specific steps of S3 are: S301: based on the frame segment number in the response frame segment focus area list, the corresponding frame segment layer in the convolution channel graph is extracted, the response coordinate position of the amplitude jump in the layer is located, and a mutation region coordinate sequence is obtained; S302: based on the coordinate content in the mutation region coordinate sequence, it is searched whether the surrounding area of the jump point surrounds the edge position of the layer, the boundary coordinates of the meshing gear ring area corresponding to the output shaft are output, and the position relationship of the coordinates in the channel is compared to obtain a gear ring boundary corresponding coordinate set; S303: based on each coordinate in the gear ring boundary corresponding coordinate set, the corresponding frame segment time point is extracted, the channel serial number and spatial index in the layer are read, the time, position and channel data are combined into the same coordinate sequence according to the frame segment number, and a boundary feature structure set is obtained.

[0011] As a further scheme of the application, the specific steps of S4 are: S401: Based on the spatial coordinates of the response region in the convolutional channel map within the boundary feature structure set, locate the horizontal and vertical axis values ​​of the response point in the layer, and obtain the channel response distribution coordinate set by matching the coordinates with the corresponding channel number. S402: Based on the horizontal and vertical axis values ​​in the channel response distribution coordinate group, scan for positional repetitions in the cross-channel graph group, and determine whether the response appears in the channel according to the relationship between the horizontal and vertical axis values ​​to obtain the inter-channel repetition response index group. S403: Based on the inter-channel repeat response index group, index the frame segment number and channel number in each corresponding channel, and pair them with time information to obtain a multi-channel structure attribution index table.

[0012] As a further aspect of the present invention, taking each pair of horizontal and vertical axis values ​​in the channel response distribution coordinate set as the center position, the corresponding scanning area is expanded according to the preset horizontal and vertical ranges, and the response position is extracted within the area. During the acquisition of the inter-channel repeating response index group, the horizontal axis values ​​and vertical axis values ​​in the channel response distribution coordinate group are compared item by item. When the two values ​​are completely consistent, they are associated with and matched with the channel number. During the acquisition of the multi-channel structure attribution index table, the frame segment number and time information corresponding to the channel number in the inter-channel repeat response index group are extracted, and the correspondence between the channel number and the frame segment number is associated according to the order of the time information.

[0013] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Based on the map group number in the multi-channel structure attribution index table, locate the corresponding region of the map group in the attention map of the convolution output channel, extract the coordinate information of the region in the attention response map according to the channel number, and pair the coordinate data with the number to obtain the attention region number sequence. S502: Based on the response position corresponding to the number in the attention region numbering sequence, retrieve the label block information in the label mapping table, identify the time point number of the frame segment corresponding to the label block, match the mapping relationship between the number and the label, and obtain the label time mapping sequence. S503: Based on the tag time mapping sequence, analyze the degree of correlation between the tag and the frame number, extract the content of the associated tag and the descriptive information in the time sequence, and obtain the gearbox fault type identification result.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by mapping the main frequency response with the time series and connecting the amplitude change with the coordinates of the attention region, a frame segment focusing region is constructed and associated with the layer boundary and the position of the tooth ring. A multi-channel response structure with consistent spatial position is extracted to form a mapping relationship between the frame segment and the label, thereby enhancing the ability to extract differences under multiple working conditions and improving the discrimination ability of feature attribution judgment and the stability of classification results. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] Please see Figure 1 This invention provides a convolutional neural network-based gearbox fault diagnosis method based on an attention mechanism, comprising the following steps: S1: Obtain the vibration data sequence of the gearbox under steady-state conditions, divide it into frames according to the meshing period, extract the amplitude response of the main frequency position in the frame segment, and match the frame segment with the frequency band position in the time series to obtain the periodic response frame segment index group. Specifically, the vibration data sequence of the gearbox under steady-state conditions is obtained. Based on the meshing period segmentation data, the amplitude response of the frequency band position corresponding to the main frequency in the frame segment is extracted. Each frame segment is mapped to the extracted frequency band amplitude position time series to obtain the periodic response frame segment index group.

[0023] S2: Based on the feature map content corresponding to the frame segment in the periodic response frame segment index group, extract the response region coordinates according to the time order, match the coordinates with the corresponding region positions in the attention distribution map, and obtain the response frame segment focus region list. Specifically, based on the feature map content corresponding to the frame segments in the periodic response frame segment index group, the position coordinates of the amplitude change area in the map are extracted according to the time order. Then, the coordinates are matched point by point with the response area coordinates in the attention distribution map. The matched frame segments and the position information in the map are sent into the target sequence to obtain the list of response frame segment focus areas.

[0024] S3: Based on the list of focused regions of the response frame segments, extract the abrupt change regions and locate the edges in the convolution channel map, and then compare the response position of the output shaft meshing gear ring with the layer coordinates and frame time to obtain the set of boundary feature structures. Specifically, based on the frame segment number in the list of focused regions of the response frame segment, the abrupt response region inside each frame segment is extracted in the convolutional channel map, the region is located at the edge position in the map, and then the corresponding boundary of the output shaft meshing gear ring in the map is compared with the corresponding boundary, and the layer coordinates of the response point are associated with the frame segment time point to obtain the boundary feature structure set.

[0025] S4: Based on the position of the response region in the channel diagram in the boundary feature structure set, search for regions with consistent spatial positions, combine the corresponding layer positions with time points, and obtain a multi-channel structure attribution index table. Specifically, based on the positional distribution of response regions in the boundary feature structure set within the differential channels of the convolutional network, the response status of corresponding regions in the channel map is retrieved, it is determined whether there are layer responses with the same spatial location, and then the time point and spatial location of the region in the differential channel are written into the output sequence to obtain a multi-channel structure attribution index table.

[0026] S5: Based on the multi-channel structure attribution index table, track the corresponding region position and label of the attention output map, the corresponding label position and frame time, and obtain the gearbox fault type identification result; Specifically, based on the map group number in the multi-channel structure attribution index table, the position of the attention distribution area in the output channel of the convolutional network is tracked, the label identifier corresponding to the numbered area is extracted from the attention map, the number of repeated mappings between the frame segment number and the label number is associated, and the label content is extracted according to the order of the label appearance to obtain the gearbox fault type identification result.

[0027] In this embodiment of the invention, the periodic response frame index group includes frequency band position, response amplitude, and time series mapping; the response frame focus area list includes amplitude change position, attention response area, and time series label; the boundary feature structure set includes layer edge coordinates, meshing gear ring boundary, and frame time point; the multi-channel structure attribution index table includes channel map number, layer response position, and time point record; and the gearbox fault type identification result includes label identifier, label number mapping count, and label content.

[0028] Please see Figure 2 The specific steps of S1 are as follows: S101: Obtain the vibration data sequence of the gearbox under steady-state conditions, extract each group of continuous vibration data in chronological order, identify the gear meshing cycle signal segment, divide the time range corresponding to the cycle length, and obtain the meshing cycle time interval group. First, vibration signals along a specific direction are acquired using an accelerometer mounted on the gearbox housing. The acquired analog signals are converted into digital signals via an A / D converter. Then, the continuously sampled digital signals are stored at fixed time intervals, with a sampling frequency of 12,800 times per second to ensure coverage of the gear meshing frequency range. Each set of sampled data is then sorted by time to maintain data continuity in the recording. Based on the gear meshing characteristics corresponding to the sensor acquisition points, the dominant frequency signal corresponding to the gear meshing frequency is identified in the frequency components. A short-time Fourier transform is performed, dividing the sampled data into several time windows. The spectral characteristic curve of the vibration signal within each window is extracted. By observing the periodic characteristics of amplitude changes in specific frequency bands, the period length corresponding to the meshing frequency is determined. The observed peak periodic positions are then mapped to the time axis. The time interval between any two adjacent peak points is marked as a gear meshing cycle. Then, the time interval between all adjacent main frequency peaks in the entire data sequence is extracted using a sliding window method, and its mean and standard deviation are calculated. Assuming that a set of extracted time intervals are 4.68 ms, 4.66 ms, 4.67 ms, 4.65 ms and 4.69 ms, the mean is 4.67 ms and the standard deviation is 0.014 ms, the judgment condition is set as the cycle interval value within plus or minus two times the standard deviation of the mean is considered as a normal cycle segment. Therefore, each cycle time segment is determined to be between 4.642 ms and 4.698 ms, which is marked as a cycle time range. These cycle segments are then numbered and summarized in chronological order, and the set of start and end time points of each gear meshing cycle in the original vibration data is output, finally obtaining the meshing cycle time interval group.

[0029] S102: Based on the time range content in the meshing cycle time interval group, the vibration data stream is divided into segments with each cycle segment as the boundary. The frequency band corresponding to the main frequency component in the time interval is extracted, and the amplitude data under the frequency band position is associated with the frame segment index to obtain the main frequency amplitude frame segment correspondence set. First, the start and end times of each cycle time segment identified in the previous steps need to be extracted one by one. Using a time interval setting method of "closed before open," the original vibration data stream is segmented with the start and end times of each cycle segment as the dividing points. For the case where 12,800 data points are collected per second in the original vibration data stream, if the start time of a certain cycle segment is 0.093 seconds and the end time is 0.0977 seconds, the corresponding sampling point range is calculated to be 1190 to 1249. Then, all acceleration sampling values ​​within this interval are extracted, completing the first step of cycle segmentation. Subsequently, short-time frequency analysis is performed on the data subsequence extracted from each cycle segment: the data segment is divided into windows of 256 points each, with a window interval of 128 points, and a Fast Fourier Transform (FFT) is performed within each window to obtain the amplitude sequence corresponding to all frequency points in the current data segment. With a sampling rate of 12800Hz and a window size of 256, the FFT resolution is 50Hz, meaning the frequency point intervals are 0Hz, 50Hz, 100Hz, 150Hz, etc. Considering the gear meshing frequency is 213.4Hz, to match the actual frequency point positions, adjacent frequency points containing 213.4Hz are selected to form the main frequency band, which is determined to be the set of frequency amplitudes between 200Hz and 250Hz. The amplitudes corresponding to each frequency point within this frequency band are extracted to form the main frequency amplitude set for the current period segment. An index relationship is established between the current period segment number and its corresponding main frequency amplitude set, thereby constructing the corresponding links between the main frequency amplitude frame segments. In the example, if the extracted main frequency amplitude values ​​for the 5th period segment are 0.342, 0.375, 0.391, 0.356, and 0.377, then an index pair is established between it and period segment number 5, recorded as (frame segment index 5, main frequency amplitude sequence [0.342, 0.375, 0.391, 0.356, 0.377]). This process is repeated for all other period segments to ensure that each period segment corresponds to a main frequency amplitude frame segment dataset. Finally, the mapping relationship between frame segment indices and main frequency amplitude data is established, resulting in a set of main frequency amplitude frame segment correspondences.

[0030] S103: Based on each set of data in the main frequency amplitude frame segment correspondence set, map the frequency band position and frame segment index in a unified timeline, aggregate the mapped data, and obtain the periodic response frame segment index group. First, the paired content of the established frame segment index and the main frequency amplitude sequence is read one by one. The main frequency band position corresponding to each frame segment index is extracted and recorded independently. In the actual gearbox vibration acquisition, if the frequency band position of the first frame segment is set to the main frequency center of 213.4 Hz and the bandwidth ±10 Hz, then the lower limit of the frequency band of 203.4 Hz and the upper limit of the frequency band of 223.4 Hz are recorded. Then, the main frequency amplitude sequence in the first frame segment is read. If the sequence is 0.356, 0.372, 0.391, 0.384, then the above four amplitudes are sequentially associated with the frame segment index 1. This process creates a multi-point mapping between frequency bands and indices on the same time axis. Then, the frequency band position of the second frame segment is processed in the same way as its dominant frequency amplitude sequence. If the corresponding frequency band position of the second frame segment is still 203.4 to 223.4 Hz, and its dominant frequency amplitude sequence is 0.341, 0.359, 0.366, and 0.358, then the four amplitude points are recorded sequentially according to their order within the frame segment and a correspondence is established with frame segment index 2. This process is repeated for all subsequent frame segments. If an amplitude anomaly exists in the dominant frequency amplitude sequence of a certain frame segment, it needs to be judged according to the preset judgment area. The amplitude range is divided into low, medium, and high amplitude ranges. The low amplitude range is defined as 0 to 0.2, the medium amplitude range as 0.2 to 0.4, and the high amplitude range as 0.4 to 1. For each amplitude value, a range determination is performed. If the amplitude value falls into the corresponding range, it is marked as belonging to that range category. For example, if the amplitude sequence in frame 3 is 0.185, 0.223, 0.247, and 0.255, then after reading each amplitude value, its range is immediately determined. 0.185 falls into the low amplitude range, and 0.223, 0.247, and 0.255... If the data falls within the mid-amplitude range, the category label is written to the same timeline. Then, in a unified timeline, a sequential arrangement action is performed to concatenate all frequency band position points and corresponding amplitude points from frame 1 to frame N in chronological order, so that the mapped data forms a continuous sequence. Next, an aggregation action is performed on the mapped data to merge all amplitude points in the same frame into a single frame entity, and the continuous frame segments are arranged into a linear data chain according to the index order. Finally, through the above mapping and aggregation actions, all frame segments are formed into a sequentially arranged response set in a unified timeline, resulting in a periodic response frame index group.

[0031] Please see Figure 3 The specific steps of S2 are as follows: S201: Based on the frame position in the periodic response frame segment index group, extract the channel coordinates of the amplitude change point in the feature map, associate the channel index, spatial position and time point of the coordinates, and store the processing result into the sequence structure to obtain the amplitude change coordinate sequence; First, the start and end time information of each frame segment is read, and its corresponding time range is mapped to the feature map data structure. The horizontal axis index range corresponding to this time interval in the feature map is traversed, and the trend of amplitude change over time in the vertical axis direction is examined point by point. Within each frame segment, based on the amplitude abrupt change detection principle, data points with significant changes are selected and recorded: the judgment condition is set as follows: if the amplitude difference between two adjacent time points exceeds a preset threshold, it is judged as a change point. This threshold needs to be set according to the overall fluctuation amplitude of the original vibration signal. For example, when the amplitude change range is between 0 and 2, the difference threshold is set to 0.25. For instance, if the vibration amplitude of a channel is 0.65 at a certain time point and 0.92 at the next time point, the change is 0.27, exceeding the threshold, and is recorded as a change point. Its row and column coordinates in the feature map are then extracted. If the point is located in channel 13 at time point 820, then extract channel index 13 and time point 820, and associate the corresponding positional parameter information of this channel in the system structure. For example, if the position coordinates of the channel are x=200mm, y=150mm, z=300mm, then the structured change point information is represented as (channel 13, x200, y150, z300, t820). This structured data is encapsulated into a recording unit and written into the change point sequence set, and the same process is continued in the current frame segment to extract the remaining change points. If the amplitude of channel 14 at time point 829 changes from 0.48 to 0.76, a change of 0.28, it also meets the judgment condition. Then, record channel index 14 and time point 829, and integrate its position coordinates x=220mm, y=150mm, z=300mm to form a new change point recording unit. In this way, all data points that meet the change conditions are processed in a loop, and the coordinate extraction and structural annotation of all valid points in the frame segment are completed in sequence. Finally, the change point information in all frames is summarized into a data sequence structure sorted by time, and the amplitude change coordinate sequence is finally obtained.

[0032] S202: Based on each set of coordinates in the amplitude change coordinate sequence, retrieve the response position of the corresponding channel in the attention distribution map, extract coordinate points with the same spatial position, and compare the time point with the layer channel index to obtain a consistent coordinate sequence; First, the channel index, spatial location, and time point information of each record within the sequence are read sequentially. These three types of data are then decomposed from the record unit into channel number ct, spatial coordinates (px, py, pz), and time point tt, and written as initial variables for the retrieval process. Next, a channel-by-channel traversal is performed in the channel dimension of the attention distribution map. ct is compared one-to-one with the layer index ci of each channel in the attention distribution map. If the values ​​of ct and ci are equal, the channel is confirmed as the current valid retrieval channel. Then, a traversal is performed on all spatial coordinates in the layer of that channel, and (… The coordinates (px, py, pz) of a point are compared with the spatial coordinates (vx, vy, vz) of each point recorded in the layer. When all three conditions (px=vx, py=vy, pz=vz) are met, the point is determined to be a point with consistent spatial location and is added to the candidate set of the current coordinates. During this process, if the spatial coordinates (px, py, pz) of a certain recording unit are (200, 150, 300), and the layer recording point with channel ci=13 in the attention distribution map has (vx, vy, vz) of (200, 150, 300), then it is directly marked. To match points and proceed to subsequent processing steps, the time axis marker (tt) of the corresponding point in the attention distribution map is compared with the time axis marker (tv) in the time dimension. If their values ​​are the same, the point is further marked as a time-complete docking point in the candidate set. If they are different, the point is discarded and the comparison proceeds to the next point. Subsequently, the same comparison step is performed on the channel index dimension, checking the consistency between ct and ci. If ci = ct, the point is added to the final docking set; if ci and ct are different, the point is removed. Throughout this process, this operation needs to be repeated sequentially for all amplitude change coordinates, and it is necessary to... The number of successful matches is recorded once during each coordinate processing step to ensure sequence integrity. For example, when performing the above retrieval on a certain recording unit (channel 13, x200, y150, z300, time 820) in the amplitude change coordinate sequence, if there is a spatial point corresponding to time 820 with coordinates (200, 150, 300) in the layer of channel 13 in the attention distribution map, then after all three conditions are met, it is written into the final docking sequence. Then the same retrieval steps are performed on the next recording unit. The entire sequence processing is completed through point-by-point and condition-by-condition comparison, and finally a docking consistent coordinate sequence is obtained.

[0033] S203: Based on the coordinate data in the docking consistent coordinate sequence, extract the associated frame segment time and layer channel index, write the position and time information of each set of data and the response area in the figure into the sequence structure, and expand the frame segment set according to the time order to obtain the list of response frame segment focus areas; First, each record is read sequentially from the sequence. The channel index ct, spatial coordinates (px, py, pz), and time point tt in the record unit are decomposed and used as the basic parameters of the current processing object. Then, in the processing flow, ct is used as the basis for channel retrieval. ct is compared item by item with the channel layer number ci in the response layer. When the values ​​of ct and ci are completely consistent, the layer is determined as the current valid response layer. Then, within this layer, it is traversed according to spatial position, and (px, py, pz) is compared with each position point (vx, vy, vz) in the layer for equality. If (px=vx, py=vy, pz=vz)... If all three conditions are met, the point is extracted and its spatial location is recorded as the matching coordinate point. Next, the corresponding time stamp (tv) of that point in the layer is read in the time dimension, and tt and tv are compared item by item. If the values ​​of tt and tv are equal, the point is added to the candidate set of the response region; otherwise, the point is discarded and the comparison of the next point is performed. After spatial matching and time comparison of all coordinate points are completed, the selected candidate points are constructed into position-time docking data units, recorded in the form of (channel ct, position (px, py, pz), time tt). Then, the action of writing a sequence structure is performed on this record unit, storing it sequentially in a time-sorted sequence list. If a channel ct appears repeatedly within a continuous time period, these recording units are constructed into an extended unit of the frame segment set according to their chronological order. For example, if there are two records in the docking coordinate sequence (channel 13, x200, y150, z300, time 820) and (channel 13, x200, y150, z300, time 821), they are considered as a continuous frame segment extended region. By calculating the difference between time points 820 and 821, if the difference is 1 and does not exceed the set continuous interval threshold of 3, the two records are included in the same frame segment focus interval, and then the next recording unit is processed. The record is (channel 13, x200, y150, z300, time 825). Since the time difference is 4, which is greater than the threshold of 3, this record is separately divided into the starting point of a new frame segment. Then, the time difference calculation is repeated for subsequent records near this time interval to determine whether to add them to the focus interval of the frame segment. The entire expansion process is carried out sequentially in chronological order, forming multiple independent frame segment focus intervals. In each interval, the corresponding channel index, spatial position, and combination of consecutive time points are retained. As all the docking and consistent coordinate sequences are processed, all frame segment focus intervals are gathered into a linear list, and finally the list of response frame segment focus areas is obtained.

[0034] Please see Figure 4 The specific steps of S3 are as follows: S301: Based on the frame segment number in the list of response frame segment focus regions, extract the corresponding frame segment layer in the convolutional channel map, locate the response coordinate position of amplitude jump in the layer, and obtain the coordinate sequence of the abrupt change region. First, each frame record in the list is read sequentially at the beginning. The frame number (fd) stored in the record is decomposed along with its associated channel index (ct), spatial coordinates (px, py, pz), and time set (tt_group). The fd is then used as the primary key parameter for entering the convolutional channel map retrieval process. Next, a layer-by-layer traversal is performed on all layers of the convolutional channel map. The layer number (ci) encountered during the traversal is compared with the fd for equality. When the values ​​of ci and fd are completely identical, the layer is marked as the target frame layer, and the next stage, amplitude transition positioning, is initiated. All frames in the target layer are then read... The amplitude sequence av(t) of the coordinate points is used, and the amplitude difference between adjacent time points is used as the basis for jump determination. The amplitude difference threshold th_v is set according to the common amplitude range of the original convolution channel map. If the image amplitude is normally distributed in the range of 0 to 1, then th_v is set to 0.2 to form an executable jump retrieval interval. Then, the adjacent time difference calculation is performed on each coordinate point in the target layer. The absolute value operation is performed on the amplitude difference of each pair of adjacent time points. If the amplitude difference is greater than th_v, then the time and the coordinate position are recorded as the jump response point. For example, at the 42nd channel of layer ci. In the channel position (x210, y160, z300), if the amplitude is 0.33 at time point 830 and 0.61 at time point 831, the difference is 0.28, which is greater than the threshold of 0.2. Therefore, this point is recorded as a mutation point. Subsequently, the three types of information of the mutation point are written into a temporary record structure, where the channel index is recorded as ct=42, the spatial position is recorded as (x210, y160, z300), and the time point is recorded as 831. This record is added to the mutation candidate sequence. Then, the above calculation is repeated for all coordinate points in the layer, and all points with amplitude differences greater than the threshold are identified. All data are written into the candidate sequence. If the amplitude change of a certain point is in a small range, it is removed. The small range needs to be divided according to the set threshold range. If the amplitude difference is in the range of 0 to 0.1, it is defined as the low amplitude difference range; if it is in the range of 0.1 to 0.2, it is defined as the medium amplitude difference range; and if it is in the range of 0.2 to 1, it is defined as the high amplitude difference range. The mutation point only comes from the records in the high amplitude difference range. As the traversal and comparison operation continues, the candidate sequence increases. After all layers are searched, all mutation points in the candidate sequence are sorted in chronological order to form a time-increasing sequence structure, thus obtaining the mutation region coordinate sequence.

[0035] S302: Based on the coordinate content in the coordinate sequence of the abrupt change region, retrieve whether the area surrounding the jump point surrounds the edge of the layer, and the boundary coordinates of the corresponding output shaft meshing gear area are compared with the positional relationship of the coordinates in the channel to obtain the coordinate set corresponding to the gear boundary. First, each transition point record in the sequence needs to be read sequentially. The spatial coordinates (px, py, pz) in the record are decomposed with the channel number ct. Then, in the retrieval stage, (px, py, pz) are compared with the edge threshold range of the corresponding layer channel map. The edge judgment threshold needs to be set according to the layer dimension. For example, if the maximum index values ​​of the layer along the x, y, and z axes are x_max=400, y_max=300, and z_max=200 respectively, then the edge threshold in the x, y, and z directions is set to 10 units. That is, if px is less than 10 or greater than 390, py is less than 10 or greater than 290, and pz is less than 390, then the edge threshold is set to 10 units in each direction. If the value is 10 or greater than 190, the jump point is determined to be near the edge, and edge marking processing is performed. The flag of the jump point is set to 1. If the point coordinates are outside the edge area, the flag is set to 0. After the edge judgment is completed, the filtering step of the boundary coordinates of the corresponding shaft meshing gear ring area is entered. Here, the gear ring structure boundary configuration table corresponding to ct needs to be read. This table stores the set of meshing gear ring boundary point coordinates under each channel. For the ct channel number, its boundary coordinate set C_ring is extracted. For each coordinate point in C_ring, the spatial position is compared with (px, py, pz). The points (vx, vy, v) in all boundary coordinates are compared. z) performs axial difference calculations with (px, py, pz) respectively. If all three-dimensional differences are within the set boundary tolerance range, the current jump point is determined to be in the gear ring boundary area. This tolerance range should be set according to the boundary component tolerance, and is set to no more than 5 units in each axis. If the jump point position is (198, 152, 298) and the gear ring boundary point is (200, 150, 300), then the three-axis differences are 2, 2, 2, all less than or equal to 5. The jump point is determined to be within the boundary area, and then added to the boundary matching set. Each jump point that passes the determination is then compared with the spatial index range of the layer structure to which ct belongs to confirm that the point is within the boundary area. Within the effective range of the channel, if the effective area of ​​the layer space index under channel ct is 100 to 350 on the x-axis, 100 to 250 on the y-axis, and 100 to 180 on the z-axis, and the current point (198, 152, 298) exceeds the maximum value of 180 on the z-axis, it is determined that although it meets the boundary matching, it is outside the layer and is not included. If another point is (200, 150, 178), it is retained because it is within the effective range on all axes. The above judgment process is repeated to perform a dual screening action of space and boundary for all transition points. Finally, all transition point sets that meet the edge proximity and boundary coordinate coverage conditions are integrated to obtain the coordinate set corresponding to the tooth ring boundary.

[0036] S303: Based on each coordinate in the coordinate set corresponding to the tooth ring boundary, extract the corresponding frame time point, read the channel number and spatial index in the layer, and combine the time, position and channel data into the same coordinate sequence according to the frame number to obtain the boundary feature structure set. First, the spatial location information (px, py, pz), time point tt, and channel number ct of each record in the set are read and passed as basic fields to the subsequent processing flow. Then, the frame segment number fd corresponding to tt is extracted from the lookup table of the frame segment time points. If tt falls within the time interval of frame segment number fd=7, fd is marked as the frame segment identifier corresponding to the current coordinate point. Next, the spatial index range of the layer corresponding to the channel number ct in the layer structure is read to confirm that the current coordinate point exists in the layer. Then, the channel number ct is searched in the layer data for the corresponding ( For points (px, py, pz) with equal positions, if the current point is (200, 150, 300) and there is a point (200, 150, 300) in the ct=27 channel layer, then record its index number ix as the channel internal index value of the current point. If no completely equal index exists, then discard the point and do not participate in subsequent combination operations. Construct structured data units for valid coordinate points, and combine the frame segment number fd, channel number ct, time tt, spatial index (px, py, pz), and layer internal index ix as attribute fields of the same data volume to form boundary features. In subsequent processing, records are categorized according to their fd values. If there are three records with fd=7 (27, 200, 150, 300, 831, ix1), (27, 198, 152, 298, 832, ix2), and (27, 196, 148, 301, 833, ix3), these three records are merged into a boundary feature set under frame segment number 7. Then, all coordinate points are processed sequentially. If the time tt of a certain coordinate point corresponds to the time interval with fd=8, its combined fields are stored in the set numbered 8. This process is repeated sequentially. The data points are grouped and organized, and then sorted by fd number from smallest to largest at the frame level so that each set is arranged in a continuous temporal order in the final output. Throughout the process, each data point needs to be validated before combination to ensure that the channel number ct is not empty, the time point tt exists in the frame time mapping relationship, the spatial position (px, py, pz) can be located at a specific index position in the layer, and the index number ix that matches the spatial position in the layer record can be successfully parsed. If any condition is not met, the record cannot participate in the combination, and finally the boundary feature structure set is obtained.

[0037] Please see Figure 5 The specific steps of S4 are as follows: S401: Based on the spatial coordinates of the response region in the convolutional channel map within the boundary feature structure set, locate the horizontal and vertical axis values ​​of the response point in the layer and match them with the coordinates of the corresponding channel number to obtain the channel response distribution coordinate set. First, the channel number ct, spatial coordinates (px, py, pz), and frame number fd of each record unit in the set are read one by one. These three types of data are used as the initial input parameters for subsequent localization operations. Then, a layer channel retrieval operation is performed in the differential convolutional channel map. The ct is compared with the layer numbers ci of all layers in the channel map for equality. When the values ​​of ci and ct are consistent, the layer is determined as the layer to be localized. Then, according to the three-dimensional spatial division structure of the layer, (px, py, pz) is compared with each spatial grid point (vx, vy, vz) inside the layer one by one. The matching action determines the specific location of the current point within the layer when all three conditions (px=vx, py=vy, pz=vz) are met. It then reads the point's index values ​​along the horizontal and vertical axes, recording the horizontal index as hx and the vertical index as hy. For example, if a point's spatial coordinates are (200, 150, 300), and the layer's spatial arrangement matches it to horizontal index hx=52 and vertical index hy=33, these values ​​are written into the point's location result set. Subsequently, a channel coordinate comparison operation is performed on this point, using ct as the channel coordinate reference. The channel number field is written into the same coordinate recording unit, forming a combined record structure of (frame segment fd, channel ct, horizontal axis hx, vertical axis hy, space (px, py, pz)). Then, the next record in the boundary feature structure set is processed. Multiple response point records are generated sequentially by repeating the above positioning, comparison, and combination operations. Throughout the process, a validity check is performed on each data point, i.e., whether (px, py, pz) falls within the valid space range corresponding to the layer ct. For example, the valid space range of layer ct is 100 to 350 on the x-axis and 100 to 350 on the y-axis. 250, z-axis 100 to 180. When (px, py, pz) = (200, 150, 178), the entire axis meets the valid range and is retained. If (px, py, pz) = (200, 150, 300) is greater than 180 in the z-axis direction, the point is removed to avoid invalid positioning. As the positioning process of all recorded points in the set is completed, all the obtained valid records are sorted according to the frame segment number fd, so that the response points under the same frame segment form a continuous coordinate sequence, while the sequences of different frame segments maintain the independence of the numbering, and finally the channel response distribution coordinate group is obtained.

[0038] S402: Based on the horizontal and vertical axis values ​​in the channel response distribution coordinate group, scan for positional repetitions in the cross-channel graph group, and determine whether the response appears in the channel according to the relationship between the horizontal and vertical axis values ​​to obtain the inter-channel repetition response index group. First, the horizontal axis hx, vertical axis hy, channel number ct, and frame number fd are read from each recording unit. hx and hy are used as core parameters for subsequent cross-channel positioning operations. Then, a layer-by-layer scanning operation is performed in the cross-channel image group. The layer number ci and ct are compared for equality. If ci and ct are different, the next layer is scanned. If ci and ct are equal, the horizontal and vertical axis coordinates are compared within that layer. Within the grid-distributed coordinate point set within the layer, the horizontal axis index vx and vertical axis index vy of each grid point are read, and (hx, hy) is compared with (vx, ..., vy) respectively. The algorithm performs a step-by-step equality operation on each channel layer (hx=vx, hy=vy). If both conditions are met, a point is considered to have a response region in the layer. For example, if a recording cell has coordinates (hx=52, hy=33, ct=27), and a grid point (vx=52, vy=33) exists in the 27th channel layer of the channel group, then this point is recorded as a valid response. The same matching operation is then performed on other channel layers to find cross-channel duplicate responses. In subsequent judgment steps, the criteria for determining duplicates need to be clearly defined; the appearance of the same location point in multiple channel layers is defined as a duplicate response. Therefore, a repeatability threshold needs to be constructed. The number of times the same horizontal and vertical coordinates are read in two or more channels is used as the criterion. The threshold is set to 2, meaning that when the number of occurrences of a coordinate point is greater than or equal to 2, it is marked as a repeat response point. If a point (hx=52, hy=33) is scanned and found to have matching points in both channel 27 and channel 31, then the number of occurrences of that point is 2, which meets the threshold requirement, and it is recorded in the repeat response list. During the recording process, it is required to perform a valid range judgment within the channel for all matching points, comparing the hx and hy of the current point with the valid range of the horizontal and vertical axes of the layer, respectively. The comparison is performed across intervals. For example, if the horizontal axis of a channel layer is 0 to 80 and the vertical axis is 0 to 60, only records where (hx=52, hy=33) fall within the valid interval are retained; otherwise, they are discarded. Then, all records that meet the repeat response condition are organized into an inter-channel repeat response index structure. The data within the structure is sorted according to the frame segment number fd to keep the repeat response points under the same frame segment in a continuous storage format. Finally, the organized points are written into the final output in the structure of (frame segment fd, horizontal axis hx, vertical axis hy, channel list {ct1, ct2, ...}), resulting in the inter-channel repeat response index group.

[0039] S403: Based on the inter-channel repeat response index group, index the frame segment number and channel number in each corresponding channel, and pair them with time information to obtain the multi-channel structure attribution index table; First, read the horizontal axis number hx, vertical axis number hy, frame segment number fd, and channel list ct_list from each repeated response record. Using hx and hy as coordinate index pairs, fd as the identifier of the time period, and the channel numbers in ct_list as the search targets, perform intra-channel index positioning in the frame segment information table of each channel. Match the current channel number ct with the number ci in the channel frame segment record table item by item. When ct = ci, extract the time index interval tt_ recorded for that channel within frame segment fd under the corresponding layer. The range is then used to compare the horizontal and vertical axis coordinates (hx, hy) with the spatial position indices of all grid points in the channel layer. A point matching (hx, hy) is found. If such a point exists (vx=hx, vy=hy), the response is considered valid within the channel ct and frame fd. At this point, fd, ct, and the time point retrieved from tt_range are combined to form a corresponding record. For example, if fd=7, ct=31, tt_range=[830, 831, 832], and hx=52, hy=33, then... If a location can be found in layer ct=31, then the following records are generated: (Frame 7, Channel 31, Time 830), (Frame 7, Channel 31, Time 831), and (Frame 7, Channel 31, Time 832) as structural attribution records. All matching results are written into the structure table sequentially. If ct_list contains multiple channel numbers, the same matching and location process must be performed on each channel sequentially. Then, all attribution records are merged to form a set of structural entries. During the matching process, if a channel layer does not have a coordinate (hx, hy) point, no record is generated, and the next channel number is processed. If the tt_range of a channel ct under the frame segment fd is empty or a time record is missing, the writing process for that channel is skipped. During the integration process, each record needs to be labeled with the frame segment number fd as the upper-level attribution identifier to ensure that all generated entries have normalized attributes in the time dimension. For example, if multiple records come from fd=7, they belong to the same frame segment group. Finally, through the above channel-by-channel, multi-time point, and multi-coordinate mapping operations, the actual attribution of all repeated response points under each channel is summarized to obtain a multi-channel structural attribution index table.

[0040] Please see Figure 6 The specific steps of S5 are as follows: S501: Based on the map group number in the multi-channel structure attribution index table, locate the corresponding region of the map group in the attention map of the convolution output channel, extract the coordinate information of the region in the attention response map according to the channel number, and pair the coordinate data with the number to obtain the attention region number sequence. First, the frame number (fd), channel number (ct), and image group number (gk) are extracted from each index record as retrieval parameters for the convolution output channel attention map. gk is used as the master index value for image group localization, and is matched sequentially with the image group indices in the attention map dataset. When gk matches the current attention image group number, that image group is selected as the valid retrieval target region. Then, layer separation is performed within that image group according to the channel dimension, extracting the content of the layer with channel number ct. The response region corresponding to the channel number in this layer is used as the spatial retrieval range. The coordinate points in the response map layer are scanned in a two-dimensional coordinate matrix format, traversing each combination of horizontal axis number hx and vertical axis number hy. The layer data is searched to see if the response point identified by fd and ct appears in the spatial matrix. If a point (x=52, y=33) in the layer completely matches the record point in the index table, i.e., hx=52, hy=33, ct=27, fd=8, then... The actual position value of the coordinate point is extracted as the hit point. Then, the two-dimensional index coordinate pair (hx, hy) of the point is used as the core data to construct the attention map coordinate data structure. At the same time, fd is used as the time assignment identifier and ct is used as the channel assignment identifier. These are combined with the coordinate point to form a structural unit, recorded in the form of (fd, ct, hx, hy). This structural unit is written into the number sequence as the identifier of the attention region. If multiple coordinate points in the ct channel layer of a certain map group meet the conditions, all points that meet the coordinates are recorded one by one, and their channel number ct and frame number fd are marked in the structural unit. The same process is performed for the processing of each map group gk. Finally, the hit coordinate points in the effective channel layers of all map groups are extracted in sequence and combined to form a complete set composed of multiple numbered coordinate units. This constitutes a pairing set based on map group number, channel number, time assignment, and coordinate index, and finally, the attention region number sequence is obtained.

[0041] S502: Based on the response position corresponding to the number in the attention region numbering sequence, retrieve the label block information in the label mapping table, identify the time point number of the frame segment corresponding to the label block, match the mapping relationship between the number and the label, and obtain the label time mapping sequence. First, each record is read sequentially, including its frame number (fd), channel number (ct), and corresponding spatial coordinate index (hx, hy). fd represents the time period to which the record belongs, ct is the channel index, and (hx, hy) identifies the spatial location of the current response point in the layer data. Then, a dynamic matching query is performed on the label mapping table. This table records the label block number (tk), label name (tag_name), time range (tag_range), and its corresponding set of candidate position coordinates (tag_hx, tag_hy), representing the possible areas where the label may appear in the data. The matching process does not rely on a fixed coordinate mapping. Instead, it compares whether (hx, hy) in the record is included in the candidate coordinate set of a certain label, and further determines whether the time point (tt) corresponding to the current time (fd) falls within the time interval of the label (tag_range), thus determining whether the association between the label and the record is valid. For example, if a record has fd=7 and ct=27, corresponding to coordinates (52, 33), and the tag "Damage A" is within the candidate coordinate range marked in the mapping table (52, 33), and its time range tag_range is tt=820 to tt=834, while the time range corresponding to frame segment 7 is tt=830 to tt=832, then the tag is considered valid at that spatiotemporal coordinate. After a match is found, the tag name, the corresponding frame segment fd, the channel number ct, and the specific time point tt are combined into a mapping record and written into the tag time mapping sequence, with the structure (tag "Damage A", frame segment 7, channel 27, time 830). When multiple tags simultaneously contain the current coordinate point and their time ranges intersect with the time period corresponding to fd, all tags that meet the conditions must be recorded in parallel; conversely, if a tag's time range does not intersect with fd or its candidate coordinate set does not contain the current point, then the tag is excluded. The entire matching process must strictly adhere to precise coordinate matching, disallowing approximations or threshold methods. Furthermore, the time range determination employs closed-interval logic, meaning both start and end boundary points are considered valid time points. Finally, all successfully matched records are organized and archived in chronological order and channel number (ct), resulting in a tag-time mapping sequence.

[0042] S503: Based on the tag time mapping sequence, analyze the degree of correlation between tags and frame segment numbers, extract the content of associated tags and their descriptive information in the time series, and obtain the gearbox fault type identification result; First, each record in the sequence needs to be read sequentially. The tag name (tag_name), time point (tt), frame number (fd), and channel number (ct) are then broken down. The fd is used as the core positioning parameter in the time series, and the tag_name is used to determine the tag's source information. The tt is then used as the specific landing point record of the tag on the time axis. Next, the process of determining the correlation between the tag and the frame number begins. During this determination, the number of times the same tag appears in different frame segments is used as the basic metric for correlation. The density of the tag's appearance on the time axis is expressed as the distribution number of consecutive time points. If the same tag appears consecutively at multiple time points, its stability is reflected by the number of consecutive points (n). The number of consecutive points (n) is then compared with the number of occurrences (m). The combination is used to construct the association weight W of the labels. The setting of W needs to be completed through direct quantization, without using abstract model logic. The summation rule W=n+m can be adopted. For example, if the label "wear A" appears 5 times in the label time mapping sequence and the consecutive time period length is 3, then W=8. If the label "eccentricity B" appears 3 times and the consecutive point length is 1, then its W=4. Subsequently, it is necessary to perform interval division judgment on the W value of all labels, dividing the W value interval into a low association region (0 to 3), a medium association region (4 to 7), and a high association region (8 to 15) in order to clarify the association strength of the labels in the time series. For the above example, "wear A" falls into the high association region, and "eccentricity B" falls into the medium association region. The interval results are written into the recording unit. After entering the second stage of processing, all tags are grouped according to frame segment number (fd). A time-point sorting operation is performed on the tag set within each frame segment, arranging the tt values ​​from smallest to largest to form a time series description for a single frame segment. For example, if time points 830, 831, and 832 in frame segment 7 correspond to the tag "wear A", then the sequence {830, 831, 832} is formed and written into the structure record as tag description information. This is then compared with other tag structures appearing in the same frame segment. If the tag "eccentricity B" only exists at time point 829 in the same frame segment, it is added to the structure description of frame segment 7 as a single-point tag record. Subsequently, all tags in frame segment 7 are arranged in chronological order, forming a time-series description of the tag description for the frame segment. The single-frame tag chain is added, and then the above operation is repeated for all frames to ensure that each frame has a combination description of tag content and time sequence. In the final identification stage, the judgment action is performed according to the association weight W of each tag and its occurrence pattern in the frame chain. If a tag appears in multiple frames and its W is in the high association area, then the tag is regarded as the final fault type. For example, "wear A" appears in frames 7, 8, and 9 and W=8, then it is judged as the final fault type. If a tag appears only in a single frame and W falls into the low association area, then it is not listed as a fault type identification item. By repeating the above screening steps for all tag records, the high association tag is used as the gearbox fault identification result, and finally the gearbox fault type identification result is obtained.

[0043] In this embodiment of the invention, by mapping the main frequency response with the time series and combining the amplitude change with the coordinates of the attention region, a frame segment focusing region is constructed and associated with the layer boundary and the position of the tooth ring. A multi-channel response structure with consistent spatial position is extracted to form a mapping relationship between the frame segment and the label identifier, thereby enhancing the ability to extract differences under multiple working conditions and improving the discrimination ability of feature attribution judgment and the stability of classification results.

[0044] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A gearbox fault diagnosis method based on an attention mechanism using a convolutional neural network, characterized in that, Includes the following steps: S1: Obtain the vibration data sequence of the gearbox under steady-state conditions, divide it into frames according to the meshing period, extract the amplitude response of the main frequency position in the frame segment, and match the frame segment with the frequency band position in the time series to obtain the periodic response frame segment index group. S2: Based on the feature map content corresponding to the frame segment in the periodic response frame segment index group, extract the response region coordinates according to the time order, match the coordinates with the corresponding region positions in the attention distribution map, and obtain the response frame segment focus region list. S3: Based on the list of focused regions of the response frame segments, extract the abrupt change regions and locate the edges in the convolution channel map, and then compare the response position of the output shaft meshing gear ring with the layer coordinates and frame time to obtain the set of boundary feature structures. S4: Based on the position of the response region in the channel diagram in the boundary feature structure set, search for regions with consistent spatial positions, combine the corresponding layer positions with time points, and obtain a multi-channel structure attribution index table. S5: Based on the multi-channel structure attribution index table, track the corresponding region position and label of the attention output map, the corresponding label position and frame time, and obtain the gearbox fault type identification result.

2. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The periodic response frame index group includes frequency band location, response amplitude, and time series mapping; the response frame focus area list includes amplitude change location, attention response area, and time series label; the boundary feature structure set includes layer edge coordinates, meshing gear ring boundary, and frame time point; the multi-channel structure attribution index table includes channel map number, layer response location, and time point record; and the gearbox fault type identification result includes label identifier, label number mapping count, and label content.

3. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The meshing cycle refers to the time interval required for the gears to complete meshing during operation; The amplitude response refers to the intensity amplitude of the vibration signal at each frequency point extracted within the frequency band, which is the energy expression form of the main frequency region.

4. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The meshing gear ring response position refers to the response boundary coordinate point corresponding to the output shaft meshing gear ring, identified and located based on boundary features in the convolution channel diagram. The attention output map refers to a visual heatmap used by the neural network to indicate the focus at a predetermined location, showing the response intensity of key feature regions.

5. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Obtain the vibration data sequence of the gearbox under steady-state conditions, extract each group of continuous vibration data in chronological order, identify the gear meshing cycle signal segment, divide the time range corresponding to the cycle length, and obtain the meshing cycle time interval group. S102: Based on the time range content in the meshing cycle time interval group, the vibration data stream is divided into segments with each cycle segment as the boundary. The frequency band corresponding to the main frequency component in the time interval is extracted, and the amplitude data under the frequency band position is associated with the frame segment index to obtain the main frequency amplitude frame segment correspondence set. S103: Based on each set of data in the main frequency amplitude frame segment correspondence set, map the frequency band position and the frame segment index in a unified timeline, aggregate the data that has been mapped, and obtain the periodic response frame segment index group.

6. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The specific steps of S2 are as follows: S201: Based on the frame position in the periodic response frame segment index group, extract the channel coordinates of the amplitude change point in the feature map, associate the channel index, spatial position and time point of the coordinates, and store the processing result into the sequence structure to obtain the amplitude change coordinate sequence. S202: Based on each set of coordinates in the amplitude change coordinate sequence, retrieve the response position of the corresponding channel in the attention distribution map, extract coordinate points with the same spatial position, and compare the time point with the layer channel index to obtain a consistent coordinate sequence. S203: Based on the coordinate data in the docking consistent coordinate sequence, extract the associated frame segment time and layer channel index, write the position and time information of each set of data and the response area in the figure into the sequence structure, and expand the frame segment set according to the time order to obtain the response frame segment focus area list.

7. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The specific steps for S3 are as follows: S301: Based on the frame segment number in the list of response frame segment focus regions, extract the corresponding frame segment layer in the convolutional channel map, locate the response coordinate position of amplitude jump in the layer, and obtain the coordinate sequence of the abrupt change region. S302: Based on the coordinate content in the coordinate sequence of the mutation region, retrieve whether the area around the jump point surrounds the edge of the layer, the boundary coordinates of the corresponding output shaft meshing gear area, compare the positional relationship of the coordinates in the channel, and obtain the coordinate set corresponding to the gear boundary. S303: Based on each coordinate in the coordinate set corresponding to the toothed ring boundary, extract the corresponding frame time point, read the channel number and spatial index in the layer, and combine the time, position and channel data into the same coordinate sequence according to the frame number to obtain the boundary feature structure set.

8. The gearbox fault diagnosis method based on attention mechanism according to claim 1, wherein the specific steps of S4 are as follows: S401: Based on the spatial coordinates of the response region in the convolutional channel map within the boundary feature structure set, locate the horizontal and vertical axis values ​​of the response point in the layer, and obtain the channel response distribution coordinate set by matching the coordinates with the corresponding channel number. S402: Based on the horizontal and vertical axis values ​​in the channel response distribution coordinate group, scan for positional repetitions in the cross-channel graph group, and determine whether the response appears in the channel according to the relationship between the horizontal and vertical axis values ​​to obtain the inter-channel repetition response index group. S403: Based on the inter-channel repeat response index group, index the frame segment number and channel number in each corresponding channel, and pair them with time information to obtain a multi-channel structure attribution index table.

9. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The horizontal and vertical axis values ​​in the channel response distribution coordinate set are used to define the spatial scanning area. Taking each pair of horizontal and vertical axis values ​​in the channel response distribution coordinate set as the center position, the corresponding scanning area is expanded according to the preset horizontal and vertical range, and the response position is extracted within the area. During the acquisition of the inter-channel repeating response index group, the horizontal axis values ​​and vertical axis values ​​in the channel response distribution coordinate group are compared item by item. When the two values ​​are completely consistent, they are associated with and matched with the channel number. During the acquisition of the multi-channel structure attribution index table, the frame segment number and time information corresponding to the channel number in the inter-channel repeat response index group are extracted, and the correspondence between the channel number and the frame segment number is associated according to the order of the time information.

10. The gearbox fault diagnosis method based on attention mechanism using convolutional neural networks according to claim 1, characterized in that, The specific steps of S5 are as follows: S501: Based on the map group number in the multi-channel structure attribution index table, locate the corresponding region of the map group in the attention map of the convolution output channel, extract the coordinate information of the region in the attention response map according to the channel number, and pair the coordinate data with the number to obtain the attention region number sequence. S502: Based on the response position corresponding to the number in the attention region numbering sequence, retrieve the label block information in the label mapping table, identify the time point number of the frame segment corresponding to the label block, match the mapping relationship between the number and the label, and obtain the label time mapping sequence. S503: Based on the tag time mapping sequence, analyze the degree of correlation between the tag and the frame number, extract the content of the associated tag and the descriptive information in the time sequence, and obtain the gearbox fault type identification result.