Railway acoustic diagnosis method, device, system and computer readable storage medium
By reconstructing the sound field using a deep learning architecture based on a self-attention mechanism, and combining unsupervised anomaly classification and fine-grained classification through transfer learning, the problem of noise interference in trackside detection technology is solved, thereby improving the accuracy of fault diagnosis.
Patent Information
- Application Number
- CN202510518663.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing trackside detection technologies are not ideal for noise control, which affects the accurate separation of effective signals and leads to a decrease in the accuracy of fault diagnosis.
A deep learning architecture based on a self-attention mechanism is used to reconstruct the sound field. This is combined with unsupervised anomaly classification and fine-grained classification through transfer learning. Acoustic and spatial information are collected collaboratively by a dual-plane array to perform interference suppression and signal separation.
It improves the accuracy of fault diagnosis, effectively separates valid signals, and enhances the condition monitoring accuracy and fault prediction capability of rail transit equipment.
Smart Images

Figure CN120445649B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of acoustic diagnosis, in particular to a trackside acoustic diagnosis method, device, system and computer readable storage medium. BACKGROUND
[0002] Railway systems are the backbone of global freight and passenger transport. It is crucial to monitor the health of vehicles to prevent derailment accidents, especially the wheels and bogies, which are key components of vehicle performance and must be strictly monitored. Trackside condition monitoring technology, as a non-contact and cost-effective bearing defect detection method, not only eliminates the additional burden on the bogie, but also reduces costs and simplifies the operation process, making it more advantageous than on-board vibration or acoustic emission detection systems.
[0003] However, the existing trackside detection technology is not ideal for noise control, affecting the accurate separation of effective signals, and thus affecting the accuracy of subsequent fault diagnosis. SUMMARY
[0004] The purpose of the present application is to provide a trackside acoustic diagnosis method, device, system and computer readable storage medium. First, the sound field is reconstructed based on sound information and spatial information, and then the diagnosis model is input after interference suppression processing, which specifically suppresses noise, accurately separates effective signals, and improves the accuracy of fault diagnosis.
[0005] In a first aspect, the present application provides a trackside acoustic diagnosis method applied to a trackside acoustic diagnosis system, the method comprising:
[0006] obtaining sound information and spatial information of a target sound source; wherein the sound information is a time-domain signal;
[0007] performing short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal;
[0008] inputting the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on self-attention mechanism;
[0009] performing interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information;
[0010] inputting the target sound information into a pre-trained fault diagnosis model to output a fault category; wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification.
[0011] In some preferable embodiments of the present application, the trackside acoustic diagnosis system comprises two acoustic arrays and a camera; the acoustic arrays are oppositely arranged on both sides of the track; the steps of acquiring the sound information and the spatial information of the target sound source comprise:
[0012] The sound information of the target sound source is acquired based on the acoustic array;
[0013] The image of the target sound source is acquired based on the camera, and the image is taken as the spatial information.
[0014] In some preferable embodiments of the present application, the acoustic array comprises 16 micro-electro-mechanical sensing modules; the micro-electro-mechanical sensing module comprises 16 micro-electro-mechanical sound sensors.
[0015] In some preferable embodiments of the present application, the step of performing interference suppression processing on the fused three-dimensional feature tensor to obtain the target sound information comprises:
[0016] The fused three-dimensional feature tensor is subjected to spatial filtering to obtain a filtered three-dimensional feature tensor;
[0017] The filtered three-dimensional feature tensor is subjected to signal enhancement processing to obtain the target sound information.
[0018] In some preferable embodiments of the present application, the signal enhancement processing comprises time-frequency masking processing and spatial enhancement processing based on the delay feature of the space.
[0019] In some preferable embodiments of the present application, the unsupervised anomaly classification comprises:
[0020] Unlabeled sound information is acquired;
[0021] A high-discrimination feature representation space is constructed, and the unlabeled sound information is subjected to anomaly feature extraction and pattern induction;
[0022] The unlabeled sound information subjected to the feature extraction and the pattern induction is preliminarily screened by using a density-based unsupervised detection algorithm.
[0023] In some preferable embodiments of the present application, the transfer learning fine-grained classification comprises:
[0024] A basic network model is acquired; the bottom convolution kernel of the basic network model has a general texture and edge feature extraction capability;
[0025] The fully connected layer of the basic network model is replaced by a classifier matched with the number of defect categories to obtain an updated basic network model;
[0026] The updated basic network model is subjected to domain alignment;
[0027] The labeled sound information is used to train the field-aligned basic network model, and the basic network model meeting the preset condition is used as the transfer learning fine-grained classification model.
[0028] In a second aspect, the present application provides a rail edge acoustic diagnosis device applied to a rail edge acoustic diagnosis system, the device comprising:
[0029] A data acquisition module is configured to acquire sound information and spatial information of a target sound source, wherein the sound information is a time-domain signal.
[0030] A sound information processing module is configured to perform short-time Fourier transform on the sound information, wherein the transformed sound information is a time-frequency domain signal.
[0031] A sound field reconstruction module is configured to input the transformed sound information and the spatial information into a pre-trained sound field reconstruction model, and output a fused three-dimensional feature tensor, wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism.
[0032] A target sound information acquisition module is configured to perform interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information.
[0033] A fault output module is configured to input the target sound information into a pre-trained fault diagnosis model, and output a fault category, wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification.
[0034] In a third aspect, the present application provides a rail edge acoustic diagnosis system comprising a processor and a memory, wherein the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the rail edge acoustic diagnosis method provided in the first aspect.
[0035] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the rail edge acoustic diagnosis method provided in the first aspect.
[0036] The present application has the following beneficial effects:
[0037] The application provides a rail edge acoustic diagnosis method, device, system and computer readable storage medium, and is applied to a rail edge acoustic diagnosis system, and the method comprises the following steps: acquiring sound information and spatial information of a target sound source; wherein the sound information is a time domain signal; performing short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal; inputting the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism; performing interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information; inputting the target sound information into a pre-trained fault diagnosis model to output a fault category; wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification; firstly, the sound field is reconstructed based on the sound information and the spatial information, then the interference suppression processing is performed, and then the pre-trained diagnosis model is inputted, so that noise is suppressed, effective signals are accurately separated, and the accuracy of fault diagnosis is improved. BRIEF DESCRIPTION OF DRAWINGS
[0038] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0039] Figure 1 A flowchart of a rail edge acoustic diagnosis method provided by the present application is shown in the figure.
[0040] Figure 2 A deployment schematic diagram of a sound receiving array provided by the present application is shown in the figure.
[0041] Figure 3 A logic block diagram of a rail edge acoustic diagnosis system provided by the present application is shown in the figure.
[0042] Figure 4 A structural schematic diagram of a sensing subsystem of a rail edge acoustic diagnosis system provided by the present application is shown in the figure.
[0043] Figure 5 A structural schematic diagram of a rail edge acoustic diagnosis device provided by the present application is shown in the figure.
[0044] Figure 6 A structural schematic diagram of a rail edge acoustic diagnosis system provided by the present application is shown in the figure.
[0045] Icon: 1 - sound array; 2 - track; 3 - micro-electromechanical acoustic sensor; 310 - data acquisition module; 320 - sound information processing module; 330 - sound field reconstruction module; 340 - target sound information acquisition module; 350 - fault output module; 400 - memory; 401 - processor; 402 - bus; 403 - communication interface. DETAILED DESCRIPTION
[0046] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0047] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.
[0048] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0049] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present application is usually placed, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" and the like are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0050] In addition, the terms "horizontal", "vertical", "overhanging" and the like do not mean that the components must be absolutely horizontal or overhanging, but can be slightly inclined. For example, "horizontal" only means that its direction is relatively more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.
[0051] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0052] Some existing technologies utilize 12 microphone arrays deployed on both sides of the track to capture bearing noise signals. These are combined with infrared bearing temperature sensors, vehicle identification antenna systems, and speed sensors to detect faults. Faults are identified by analyzing the amplitude and frequency characteristics of the acoustic signals. However, during train travel, the wheelset rolling bearing signals mix with the impact noise from the wheelset tread and track, making it difficult to effectively separate these signals using such a sparse trackside array. This, in turn, affects the accuracy and performance of subsequent fault detection.
[0053] Some existing technologies use trackside microphones to acquire bearing noise signals and utilize neural networks to directly map fault characteristics, vehicle speed, and diagnostic results. This avoids complex Doppler distortion correction modeling and reduces the complexity of trackside acoustic diagnostic systems. However, these still face the problem of complex noise interference during train operation, such as impact noise between the wheelset tread and the track or ambient noise, which can affect the diagnostic accuracy of the target signal. Furthermore, fault data or prior data from the operational phase is difficult to obtain, which poses a significant challenge to subsequent fault diagnosis.
[0054] In summary, addressing the two major technical bottlenecks in online monitoring of rail transit wheelset bearings—difficult signal extraction and scarcity of annotated data—a trackside acoustic diagnosis method has been proposed. This method is applied to a trackside acoustic intelligent diagnosis system based on a large-scale acoustic sensor array. The innovations of this method are primarily reflected in the following three aspects:
[0055] First, in terms of hardware architecture, a dual-plane array collaborative acquisition solution was adopted. Two high-precision microphone arrays with 256 elements were symmetrically deployed on either side of the track to construct a three-dimensional sound field information acquisition system. Compared to traditional single-point sensors, this architecture can simultaneously acquire the time, frequency, and spatial domain characteristics of the sound pressure signal, improving its spatial resolution by two orders of magnitude and laying the data foundation for subsequent signal processing.
[0056] Secondly, at the signal processing level, neural network-based beamforming technology leverages the array's spatial beam-steering properties to accurately extract and reconstruct the acoustic signature of wheelset bearings in motion, significantly improving signal-to-noise ratios compared to traditional methods. Combined with a sound source localization algorithm, the system effectively isolates interference sources such as wheel-rail contact noise and ambient noise, ensuring the integrity of the target signal.
[0057] Finally, in terms of diagnostic algorithms, a phased intelligent detection architecture was designed: in the first phase, a self-supervised contrastive learning method was used to construct a representation space, and unsupervised anomaly detection was used to achieve initial screening of potential faults; in the second phase, a transfer learning strategy was combined with limited labeled data to complete the fine classification of fault types.
[0058] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0059] Example 1
[0060] The embodiment of the present invention provides a trackside acoustic diagnosis method, which is applied to a trackside acoustic diagnosis system. Figure 1 The flowchart of a trackside acoustic diagnostic method provided in an embodiment of the present invention is shown, and the method includes:
[0061] Step S102: Acquire acoustic information and spatial information of the target sound source; wherein the acoustic information is a time domain signal.
[0062] Specifically, the target sound source is generally a train running on the track. The sound information of the train is obtained through the sound collection system, and the spatial information represents the direction of the train's movement. In some preferred embodiments of the present invention, based on the position of the sound collection system, the sound information can also include the position information of some collection equipment.
[0063] Furthermore, in some preferred embodiments of the present invention, the trackside acoustic diagnostic system includes: two acoustic arrays 1 and a camera; wherein the acoustic arrays 1 are relatively arranged on both sides of the track 2; the step of obtaining the acoustic information and spatial information of the target sound source includes: collecting the acoustic information of the target sound source based on the acoustic arrays 1; collecting an image of the target sound source based on the camera, and using the image as spatial information.
[0064] For details, see Figure 2 The diagram shows a deployment diagram of a sound transmission array provided by an embodiment of the present invention, wherein the sound transmission arrays 1 are arranged opposite to each other on both sides of the track 2; Figure 2 The illustrated sound transmission arrays 1 are each provided with a plurality of micro-electromechanical acoustic sensors 3 , through which acoustic signals are collected.
[0065] For further information, see Figure 3The embodiment of the application shown provides a logic block diagram of a rail edge acoustic diagnosis system, acquires real-time positioning information of a train through a visual recognition unit, and automatically reads a locomotive electronic tag number by using a radio frequency identification unit (RFID).
[0066] Further, in some preferred embodiments of the application, the sound array 1 comprises: 16 micro-electromechanical sensing modules; each micro-electromechanical sensing module comprises: 16 micro-electromechanical sound sensors 3.
[0067] Specifically, each array is composed of 256 high-sensitivity micro-electromechanical sound sensors 3, which can realize full-circumferential sound wave signal collection of the train wheel bearing.
[0068] For example, referring to Figure 4 The embodiment of the application shown provides a structural diagram of a sensing subsystem of a rail edge acoustic diagnosis system, which comprises: 16 micro-electromechanical sensing modules, each of which comprises 16 micro-electromechanical sound sensors 3 and is connected to a processor 401 through a cascaded time-division multiplexing (TDM) data interface, and the processor 401 transmits the collected sound signals to a user terminal through a network. At the same time, the system is equipped with network cameras for image capture, in which the camera 1 is integrated with the sound array, the obtained image provides target spatial information for the reconstructed sound field, and then the target sound source can be locked through the sound field information; the cameras 2 and 3 are responsible for capturing the direction of the train, respectively, and assisting in monitoring the running state of the vehicle. Finally, the audio and video signals are converged on a network cable through a router and then transmitted to the user terminal.
[0069] Step S104: performing short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal.
[0070] Specifically, the computing service is served for the data storage unit, adopts a distributed data storage architecture and a multi-GPU parallel computing platform, and constructs a sound signal processing pipeline. First, the time domain signal collected by the 256-channel acoustic sensor array is subjected to short-time Fourier transform; in some preferred embodiments of the application, the short-time Fourier transform includes: a window function length of 256 ms, a Hamming window, and a frame shift of 50%.
[0071] Step S106: inputting the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism.
[0072] Specifically, by the sound field reconstruction model based on the deep learning (Transformer) architecture, sound field characterization learning is performed in a time-space-frequency three-dimensional feature space, the correlation between time, space and frequency domain features is obtained, and the features of the maneuvering target sound signal are extracted, so that the geometric space features can be better fused to realize the sound source beam directional tracking of the wheel set area.
[0073] In step S108, the fused three-dimensional feature tensor is subjected to interference suppression processing to obtain target sound information.
[0074] Specifically, the fused features are subjected to noise reduction enhancement processing.
[0075] Further, in some preferred embodiments of the present application, the step of suppressing interference processing on the fused three-dimensional feature tensor to obtain target sound information comprises: performing spatial filtering on the fused three-dimensional feature tensor to obtain a filtered three-dimensional feature tensor; and performing signal enhancement processing on the filtered three-dimensional feature tensor to obtain target sound information.
[0076] Specifically, for the target sound source direction, a neural network beamformer can be used to generate spatial filter weights in real time to realize high-precision and high-stability adaptive spatial filtering, and a deep learning optimized beamforming algorithm can be used to effectively suppress interference noise. The specific process is as follows: first, the neural network generates the coefficients of the spatial filter in real time according to the characteristics of the data, which is actually equivalent to real-time statistics of the characteristics of the sound signal, background noise and interference noise of each channel.
[0077] Further, in some preferred embodiments of the present application, the signal enhancement processing comprises: time-frequency masking processing and spatial enhancement processing based on the delay characteristics of the space.
[0078] Specifically, time-frequency masking is performed to remove noise and improve the signal-to-noise ratio of each channel, and then spatial enhancement is performed through the delay characteristics of the space, thereby realizing the suppression of interference noise. Finally, the signal-to-noise ratio of the output wheel set rolling bearing feature sound signal is improved, providing a high-quality data basis for subsequent fault diagnosis and condition monitoring.
[0079] In some preferred embodiments of the present application, since sound sources at different directions have different spatial characteristics, the microphones deployed in the space need to be compensated for delay to achieve the purpose of spatial increase.
[0080] In step S110, the target sound information is input into a pre-trained fault diagnosis model to output a fault category; wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification.
[0081] Specifically, the fault diagnosis model adopts a two-stage intelligent detection architecture to realize a progressive diagnosis process from abnormal preliminary screening to accurate classification.
[0082] Further, in some preferred embodiments of the present application, the unsupervised anomaly classification includes: obtaining unlabeled acoustic information; constructing a high-discrimination feature representation space, performing anomaly feature extraction and pattern induction on the unlabeled acoustic information; and using a density-based unsupervised detection algorithm to preliminarily screen the unlabeled acoustic information after feature extraction and pattern induction.
[0083] Specifically, the first stage adopts a self-supervised contrast learning framework, performs anomaly feature extraction and pattern induction on a large amount of unlabeled samples by constructing a high-discrimination feature representation space, and uses a density-based DBSCAN (Density-Based Spatial Clustering of Applications with Noise) unsupervised detection algorithm to complete preliminary screening of potential fault samples.
[0084] Further, in some preferred embodiments of the present application, the transfer learning fine-grained classification includes: obtaining a basic network model; wherein the bottom convolution kernel of the basic network model has general texture and edge feature extraction capability; replacing the fully connected layer of the basic network model with a classifier matching the number of defect categories to obtain an updated basic network model; performing domain alignment on the updated basic network model; and training the domain-aligned basic network model based on labeled acoustic information, and using the basic network model that meets the preset condition as a transfer learning fine-grained classification model.
[0085] Specifically, the second stage introduces a transfer learning strategy, through knowledge transfer and parameter fine-tuning mechanism, first, a pre-trained model (such as ResNet, EfficientNet) is used as a basic network, and its bottom convolution kernel has general texture and edge feature extraction capability. Second, through feature reuse, the first N layers (such as the first 3 residual blocks of ResNet50) of the pre-trained model are frozen to retain its general visual feature extraction capability; the top network (fully connected layer) is replaced with a classifier matching the number of defect categories. Then, domain alignment is performed through input layer interpolation or adaptive pooling layer to eliminate the scale difference between the source domain and the target domain, and realize cross-dimension transfer. Finally, the feature extraction capability of the pre-trained model is transferred to the downstream classification task, and combined with a small amount of precisely labeled samples to realize fine-grained recognition of multiple types of defects. This phased processing architecture effectively solves the problem of lack of labeled data in industrial detection, and significantly improves the practicality and classification accuracy of the model in actual industrial scenarios.
[0086] The rail edge acoustic diagnosis method innovatively combines voiceprint recognition and time-frequency analysis technology, can effectively identify early cracks, peeling, corrosion and other hidden damages of key components such as the inner ring, outer ring and rolling body of the rolling bearing, realizes intelligent evaluation and early warning decision support of the fault grade, and significantly improves the state monitoring precision and fault prediction ability of the rail transportation equipment.
[0087] The rail edge acoustic diagnosis method is applied to a rail edge acoustic diagnosis system, and the method comprises the following steps: acquiring sound information and spatial information of a target sound source; wherein the sound information is a time-domain signal; performing short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal; inputting the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism; performing interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information; inputting the target sound information into a pre-trained fault diagnosis model to output a fault category; wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification; first, the sound field is reconstructed based on the sound information and the spatial information, and then the interference suppression processing is performed, and then the pre-trained diagnosis model is inputted, so that the noise is specifically suppressed, the effective signal is accurately separated, and the accuracy of fault diagnosis is improved.
[0088] Embodiment two
[0089] On the basis of the above-mentioned embodiments, the rail edge acoustic diagnosis device provided by the embodiments of the present application is applied to a rail edge acoustic diagnosis system, and the structure schematic diagram of the rail edge acoustic diagnosis device provided by the embodiments of the present application is shown in Figure 5 The device comprises:
[0090] The data acquisition module 310 is configured to acquire sound information and spatial information of a target sound source; wherein the sound information is a time-domain signal.
[0091] The sound information processing module 320 is configured to perform short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal.
[0092] The sound field reconstruction module 330 is configured to input the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism.
[0093] The target sound information acquisition module 340 is configured to perform interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information.
[0094] The fault output module 350 is configured to input the target sound information into a pre-trained fault diagnosis model, and output a fault category; wherein the fault diagnosis model is configured to perform unsupervised anomaly classification and transfer learning fine-grained classification.
[0095] Further, in some preferred embodiments of the present application, the trackside acoustic diagnosis system comprises two sound arrays 1 and a camera; wherein the sound arrays 1 are oppositely arranged on both sides of the track 2; the data acquisition module 310 is configured to acquire sound information of the target sound source based on the sound arrays 1; and the camera is configured to acquire an image of the target sound source, and the image is used as spatial information.
[0096] Further, in some preferred embodiments of the present application, the sound array 1 comprises 16 micro-electro-mechanical sensing modules; and each micro-electro-mechanical sensing module comprises 16 micro-electro-mechanical sound sensors 3.
[0097] Further, in some preferred embodiments of the present application, the target sound information acquisition module 340 is configured to perform spatial filtering on the fused three-dimensional feature tensor to obtain a filtered three-dimensional feature tensor; and perform signal enhancement processing on the filtered three-dimensional feature tensor to obtain the target sound information.
[0098] Further, in some preferred embodiments of the present application, the signal enhancement processing comprises time-frequency masking processing and spatial enhancement processing based on delay features in the spatial domain.
[0099] Further, in some preferred embodiments of the present application, the unsupervised anomaly classification comprises: acquiring unlabeled sound information; constructing a high-discrimination feature representation space, and performing abnormal feature extraction and pattern induction on the unlabeled sound information; and using a density-based unsupervised detection algorithm to preliminarily screen the unlabeled sound information after the feature extraction and pattern induction.
[0100] Further, in some preferred embodiments of the present application, a basic network model is acquired; wherein the bottom convolution kernel of the basic network model has a general texture and edge feature extraction capability; a full connection layer of the basic network model is replaced by a classifier matched with the number of defect categories to obtain an updated basic network model; the updated basic network model is subjected to domain alignment; and the basic network model after the domain alignment is trained based on the labeled sound information, and the basic network model meeting a preset condition is used as a transfer learning fine-grained classification model.
[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the trackside acoustic diagnosis device described above can refer to the corresponding process in the embodiments of the trackside acoustic diagnosis method described above, and will not be described herein.
[0102] Embodiment three
[0103] The embodiment of the present application also provides a track-side acoustic diagnosis system for running the track-side acoustic diagnosis method; see Figure 6 The track-side acoustic diagnosis system provided by the embodiment of the present application has the structure as shown in the figure, and the system comprises a memory 400 and a processor 401, wherein the memory 400 is used for storing one or more computer instructions, and the one or more computer instructions are executed by the processor 401 to realize the track-side acoustic diagnosis method.
[0104] Further, Figure 6 The track-side acoustic diagnosis system also comprises a bus 402 and a communication interface 403, and the processor 401, the communication interface 403 and the memory 400 are connected through the bus 402.
[0105] The memory 400 can contain a high-speed random access memory 400 (RAM, Random Access Memory), and can also comprise a non-volatile memory 400 (non-volatile memory), for example at least one disk memory 400. The communication connection between the track-side acoustic diagnosis system network element and at least one other network element is realized through at least one communication interface 403 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network and the like can be used. The bus 402 can be an ISA bus 402, a PCI bus 402 or an EISA bus 402, etc. The bus 402 can be divided into an address bus 402, a data bus 402, a control bus 402, etc. For the convenience of representation, Figure 6 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus 402 or only one type of bus 402.
[0106] The processor 401 can be an integrated circuit chip having a signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 401 or the instruction in the form of software. The processor 401 described above can be a general processor 401, including a central processing unit 401 (Central Processing Unit, CPU for short), a network processor 401 (Network Processor, NP for short) and the like; can also be a digital signal processor 401 (Digital Signal Processor, DSP for short), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC for short), a field programmable gate array (Field-Programmable Gate Array, FPGA for short) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor 401 can be a microprocessor or the processor 401 can also be any conventional processor or the like. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware decoding processor 401 for execution, or a combination of hardware and software modules in the decoding processor 401 for execution. The software module can be located in a random access memory 400, a flash memory, a read-only memory 400, a programmable read-only memory 400 or an electrically erasable programmable memory 400, a register or the like mature storage medium in the art. The storage medium is located in the memory 400, and the processor 401 reads the information in the memory 400, and combines the hardware to complete the steps of the method of the above embodiments.
[0107] The embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores computer executable instructions, when the computer executable instructions are called and executed by the processor 401, the computer executable instructions cause the processor 401 to implement the above rail side acoustic diagnosis method, and the specific implementation can be referred to the method embodiment, and will not be repeated here.
[0108] The computer program product of the rail side acoustic diagnosis method, device and system provided by the embodiment of the present application includes a computer readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the method in the foregoing method embodiment, and the specific implementation can be referred to the method embodiment, and will not be repeated here.
[0109] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and / or device described above can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here.
[0110] In addition, in the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "connecting" should be understood in a broad sense, for example, can be fixedly connected, or can be detachably connected, or integrally connected; can be mechanically connected, or can be electrically connected; can be directly connected, or indirectly connected through an intermediate medium; can be internal communication of two elements. For those skilled in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0111] The functions described above, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the technical solutions that make essential contributions to the prior art or the part of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory 400 (ROM, Read-Only Memory), a random access memory 400 (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0112] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A rail-bound acoustic diagnostic method, characterized in that, The method is applied to a rail-side acoustic diagnosis system, and the method comprises: acquiring sound information and spatial information of a target sound source; wherein the sound information is a time-domain signal; performing short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal; inputting the transformed sound information and the spatial information into a pre-trained sound field reconstruction model to output a fused three-dimensional feature tensor; wherein the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism; performing interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information; inputting the target sound information into a pre-trained fault diagnosis model to output a fault category; wherein the fault diagnosis model is used for unsupervised anomaly classification and transfer learning fine-grained classification; wherein the unsupervised anomaly classification comprises: acquiring unlabeled sound information; constructing a high-discrimination feature representation space, performing anomaly feature extraction and pattern induction on the unlabeled sound information; and preliminarily screening the unlabeled sound information after feature extraction and pattern induction by using a density-based unsupervised detection algorithm; the transfer learning fine-grained classification comprises: acquiring a basic network model; wherein the bottom convolution kernel of the basic network model has general texture and edge feature extraction capability; replacing the fully connected layer of the basic network model with a classifier matched with the number of defect categories to obtain an updated basic network model; performing domain alignment on the updated basic network model; and training the domain-aligned basic network model based on labeled sound information, and taking the basic network model meeting a preset condition as a transfer learning fine-grained classification model.
2. The trackside acoustic diagnostic method of claim 1, wherein, The rail-side acoustic diagnosis system comprises: two sound receiving arrays and a camera; wherein the sound receiving arrays are oppositely arranged on both sides of a track; and the step of acquiring sound information and spatial information of a target sound source comprises: acquiring sound information of a target sound source based on the sound receiving arrays; acquiring an image of the target sound source based on the camera, and taking the image as the spatial information.
3. The trackside acoustic diagnostic method of claim 2, wherein, The sound receiving array comprises: 16 micro-electro-mechanical sensing modules; and each micro-electro-mechanical sensing module comprises: 16 micro-electro-mechanical sound sensors.
4. The trackside acoustic diagnostic method of claim 1, wherein, The step of performing interference suppression processing on the fused three-dimensional feature tensor to obtain target sound information comprises: performing spatial filtering on the fused three-dimensional feature tensor to obtain filtered three-dimensional feature tensors; performing signal enhancement processing on the filtered three-dimensional feature tensors to obtain the target sound information.
5. The trackside acoustic diagnostic method of claim 4, wherein, The signal enhancement processing comprises: time-frequency masking processing and spatial enhancement processing based on delay features in the spatial domain.
6. A trackside acoustic diagnostic apparatus, characterized by, The device is applied to a rail-side acoustic diagnosis system, and the device comprises: a data acquisition module configured to acquire sound information and spatial information of a target sound source; wherein the sound information is a time-domain signal; a sound information processing module configured to perform short-time Fourier transform on the sound information; wherein the transformed sound information is a time-frequency domain signal; The sound field reconstruction module is configured to input the transformed sound information and the spatial information into a pre-trained sound field reconstruction model, and output a fused three-dimensional feature tensor; the sound field reconstruction model is a deep learning architecture based on a self-attention mechanism; The target sound information acquisition module is configured to perform interference suppression processing on the fused three-dimensional feature tensor, and obtain target sound information; The fault output module is configured to input the target sound information into a pre-trained fault diagnosis model, and output a fault category; the fault diagnosis model is configured to perform unsupervised anomaly classification and transfer learning fine-grained classification; The unsupervised anomaly classification includes: obtaining unlabeled sound information; constructing a high-discrimination feature representation space, performing anomaly feature extraction and pattern induction on the unlabeled sound information; and performing preliminary screening on the unlabeled sound information after feature extraction and pattern induction by using a density-based unsupervised detection algorithm; the transfer learning fine-grained classification includes: obtaining a basic network model; the bottom convolution kernel of the basic network model has general texture and edge feature extraction capability; replacing the full connection layer of the basic network model with a classifier matched with the number of defect categories to obtain an updated basic network model; performing domain alignment on the updated basic network model; training the domain-aligned basic network model based on labeled sound information, and taking the basic network model meeting a preset condition as a transfer learning fine-grained classification model.
7. A trackside acoustic diagnostic system, characterized by The computer readable storage medium stores computer executable instructions, and the computer executable instructions, when called and executed by the processor, cause the processor to implement the rail edge acoustic diagnosis method in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, and the computer executable instructions, when called and executed by the processor, cause the processor to implement the rail edge acoustic diagnosis method in any one of claims 1 to 5.
Citation Information
Patent Citations
Acoustic monitoring system and method for railway vehicle traction motor bearing
CN113447270A
Steel rail real-time defect detection method based on lightweight cascade network
CN116092027A