Three-dimensional point cloud fused sound source separation method and device for power transformation main equipment
By combining three-dimensional point clouds and sound signals, using high-precision laser scanning and improved Transformer model, the precise separation and monitoring of multiple sound sources in the substation is achieved, solving the problem of poor sound source separation effect and improving the accuracy of equipment status monitoring.
Patent Information
- Application Number
- CN202510574904.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-08
AI Technical Summary
The existing substation sound source separation method has poor effect in multi-sound source environments, making it difficult to accurately separate and monitor the status of each substation main equipment.
High-precision three-dimensional laser scanner is used to obtain the device point cloud data, combine it with a microphone array to capture sound signals, and the sound source separation is performed through the improved Transformer model, and feature fusion is performed using particle swarm optimization and k-nearest neighbor classification algorithm, and the separation results are optimized through multiple iterations.
It significantly improves the accuracy of sound source positioning and identification, reduces noise interference, and improves the reliability of substation equipment fault diagnosis.
Smart Images

Figure CN120279932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of substation sound source separation, and in particular to a method and device for separating the sound sources of main substation equipment by integrating three-dimensional point clouds. Background Art
[0002] During the operation of modern substations, many main substation equipment, such as transformers, circuit breakers, cables, etc., will generate complex electromagnetic interference and mechanical noise when operating simultaneously. These sound signals are mixed together to form a multi-source sound environment, making it difficult to monitor the status of a single device separately.
[0003] The substation sound source separation method aims to effectively separate the sounds of each main substation equipment, ensure that each main substation equipment can be independently monitored, and improve the accuracy of anomaly detection. Traditional sound source separation methods are mostly based on pure acoustic signal processing or deep learning methods, such as spectral subtraction, non-negative matrix factorization methods, etc., and usually have problems such as lack of spatial information and sensitivity to the number of sound sources. Compared with the complex scenario of multiple devices running simultaneously in a substation, the sound source separation effect is poor. Summary of the Invention
[0004] To solve the problem of poor sound source separation effect in the existing technology, the primary object of the present invention is to provide a method for separating the sound sources of main substation equipment by integrating three-dimensional point clouds, which can significantly improve the accuracy of sound source localization and recognition, more accurately locate the sound source, and reduce the interference of noise.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for separating the sound sources of main substation equipment by integrating three-dimensional point clouds, the method comprising the following steps in sequence:
[0006] (1) Deploy a high-precision three-dimensional laser scanner in the substation, scan the equipment area, obtain high-density point cloud data and perform preprocessing, and perform equipment segmentation to obtain the three-dimensional point cloud data of different equipment;
[0007] (2) Capture the sound signals generated by each equipment and perform preprocessing to extract the sound signal features;
[0008] (3) Integrate the three-dimensional point cloud data of different equipment with the sound signal features to obtain integrated features;
[0009] (4) Improve the Transformer model to obtain a sound source separation network based on Transformer, input the integrated features into the sound source separation network based on Transformer, perform sound signal separation, and obtain the separated sound signals;
[0010] (5) Input the separated sound signal into the Transformer-based sound source separation network for multiple iterations to optimize the separation result until the Transformer-based sound source separation network converges.
[0011] In step (1), the obtaining of high-density point cloud data and its preprocessing, and the device segmentation specifically refer to: First, register the point cloud data in different coordinate systems, and use the feature point registration method to convert the point cloud data from different scanning positions to the same coordinate system to obtain the global three-dimensional point cloud data; adopt the denoising algorithm based on DBSCAN, introduce the height feature to distinguish electrical equipment and background noise; use the extraction method based on neighborhood feature aggregation to extract spatial features to achieve the rapid segmentation of electrical equipment and obtain the three-dimensional point cloud data of different devices.
[0012] Step (2) specifically includes the following steps in sequence:
[0013] (2a) Deploy a microphone array near the key equipment in the substation to collect the sound signals of each equipment and its surrounding environment;
[0014] (2b) Perform noise reduction processing on the sound signal, and use the noise reduction method to remove the noise and interference from the environment; the noise reduction method includes any one of wavelet transform, spectral subtraction, and adaptive filtering;
[0015] (2c) Extract features from the noise-reduced sound signal to obtain sound signal features, and the sound signal features include spectrum, amplitude, and phase features.
[0016] Step (3) specifically refers to: Using the particle swarm optimization algorithm and the k-nearest neighbor classification algorithm, fuse the three-dimensional point cloud data with the sound signal features to obtain fusion features.
[0017] In step (4), the improvement of the Transformer model to obtain the Transformer-based sound source separation network specifically refers to: According to the acoustic fingerprint time-frequency domain characteristics, divide the subspace of the multi-head attention of the Transformer model into three different scales, and input the query, key, and value vectors into the scaled dot-product attention layers of three different scales through linear layers respectively. Each scale of the scaled dot-product attention layer performs matrix multiplication, scaling, and normalization exponentiation on the query and key vectors respectively, and then performs matrix multiplication with the value vector to make different heads focus on features of different granularities. The three different scales include 32*32, 64*64, and 128*128.
[0018] Another object of the present invention is to provide an electronic device, including:
[0019] A processor; and
[0020] A memory stores computer program instructions, and when the computer program instructions are run by the processor, the processor executes the method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds as described above.
[0021] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are run by a processor, the processor executes the method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds as described above.
[0022] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention provides a new multi-sound source separation method. By combining three-dimensional point clouds with sound signals, the accuracy of sound source localization and recognition can be significantly improved. Second, compared with traditional sound separation methods, the present invention can more accurately locate sound sources with the help of three-dimensional point cloud data, reduce noise interference, improve the reliability of power transformation equipment fault diagnosis, and has broad application prospects, especially in the aspect of automatic monitoring and maintenance of power systems. Description of the Drawings
[0023] Figure 1 is a flowchart of the method of the present invention;
[0024] Figure 2 is a schematic diagram of denoising three-dimensional point clouds by DBSCAN;
[0025] Figure 3 is a schematic diagram of the structure of a sound source separation network based on Transformer. Detailed Embodiments
[0026] As Figure 1 shown, a method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds includes the following steps in sequence:
[0027] (1) Deploy a high-precision three-dimensional laser scanner in a substation. The three-dimensional laser scanner can efficiently and accurately obtain the spatial geometric information of equipment, including the shape, size of the equipment and its relative position with the environment. The accuracy of point cloud data directly affects the accuracy of subsequent equipment recognition and sound source localization. Therefore, using a high-precision three-dimensional laser scanner can ensure that the collected data has sufficient resolution and accuracy. Scan the equipment area, obtain high-density point cloud data and perform preprocessing, and perform equipment segmentation to obtain three-dimensional point cloud data of different equipment; the equipment area includes main transformers, gas-insulated switchgears, circuit breakers and cables;
[0028] (2) Capture the sound signals generated by each equipment and perform preprocessing to extract the sound signal features;
[0029] (3)Fuse the 3D point cloud data of different devices with the sound signal features to obtain fused features;
[0030] (4)Improve the Transformer model to obtain a Transformer-based sound source separation network. Input the fused features into the Transformer-based sound source separation network to perform sound signal separation and obtain the separated sound signals;
[0031] (5)Input the separated sound signals into the Transformer-based sound source separation network for multiple iterative optimizations of the separation results until the Transformer-based sound source separation network converges. To ensure that the final sound source separation results have high signal-to-noise ratio and clarity, the present invention uses a multiple iterative optimization process to continuously improve the quality of the separation results. In each iteration, the Transformer-based sound source separation network will be adjusted according to the current separation results to reduce the noise components and optimize the separation clarity. In each round of the iterative process, the Transformer-based sound source separation network will continuously adjust the weights and update the sound source features to finally achieve the best sound source separation effect. The optimized sound source signals can not only accurately restore the operating states of each main electrical equipment, but also effectively exclude irrelevant noise components to ensure the accuracy of the monitoring results.
[0032] In step (1), the obtaining of high-density point cloud data and preprocessing, and the performing of equipment segmentation specifically refer to: First, register the point cloud data in different coordinate systems. Use the feature point registration method to transform the point cloud data from different scanning positions into the same coordinate system to obtain the global 3D point cloud data. The feature point registration method uses the ICP algorithm; Adopt the denoising algorithm based on DBSCAN, which can effectively remove the background noise in the environment and only retain the point cloud data related to electrical equipment. As Figure 2 shown, introduce the height feature. The height feature usually refers to the position information of the data points in a certain dimension (such as the vertical axis in space). By analyzing the height distribution of the point cloud data points, the effective signals (such as electrical equipment) concentrated in a specific height range can be identified, and the data in other height regions are regarded as noise, so as to distinguish electrical equipment from background noise; Use the extraction method based on neighborhood feature aggregation to extract spatial features to achieve the rapid segmentation of electrical equipment and obtain the 3D point cloud data of different devices.
[0033] Step (2) specifically includes the following steps in sequence:
[0034] (2a) Deploy a microphone array near the key equipment in the substation to collect the sound signals of each equipment and its surrounding environment. The microphone array can obtain sound signals from different directions and distances, and perform spatial positioning through the time difference and intensity difference of the signals, providing valuable information for subsequent sound source separation. By arranging the microphone array, the sound signals emitted by multiple equipment simultaneously can be captured, providing rich raw data for the sound source separation algorithm.
[0035] (2b) Perform noise reduction processing on the sound signals, and use a noise reduction method to remove noise and interference from the environment. The noise reduction method includes any one of wavelet transform, spectral subtraction, and adaptive filtering.
[0036] (2c) Extract features from the noise-reduced sound signals to obtain sound signal features. The sound signal features include spectrum, amplitude, and phase features, which reflect the frequency components, intensity, and propagation mode of the sound, providing input data for subsequent sound source separation and recognition.
[0037] (3) Specifically refers to: using the particle swarm optimization algorithm and the k-nearest neighbor classification algorithm to fuse the three-dimensional point cloud data with the sound signal features to obtain fused features. The particle swarm optimization algorithm is used to optimize the weighted combination of the point cloud data and the sound signal features to ensure finding the best matching method between the spatial dimension and the sound signal dimension. The k-nearest neighbor classification algorithm is used to classify different sound sources according to the spatial position of the equipment (from the point cloud data) and the spectral features of the sound signal, thereby realizing the precise positioning of the sound source. By effectively fusing the two types of data (point cloud and sound signal features), the algorithm can more accurately locate the spatial position and its audio features of each sound source, laying a foundation for subsequent sound source separation and recognition.
[0038] In step (4), such as Figure 3As shown in the figure, the improvement of the Transformer model to obtain a sound source separation network based on Transformer specifically means that: due to the short-term stationarity of the voiceprint signal, according to the time-frequency domain characteristics of the voiceprint, the subspace of the multi-head attention of the Transformer model is divided into three different scales. The query Query, key Key, and value Value vectors are respectively input into three different scales of scaled dot-product attention layers through linear layers. Each scale of the scaled dot-product attention layer performs matrix multiplication, scaling, and normalization exponentiation on the query Query and key Key vectors respectively, and then performs matrix multiplication with the value Value vector, so that different heads focus on features of different granularities, thereby reducing the redundancy problem in traditional multi-head attention and enhancing feature diversity. The three different scales include 32*32, 64*64, and 128*128. The sound source separation network based on Transformer has a powerful ability to process sequence data and can capture long-term dependencies in sound signals. During the sound source separation process, 3D point cloud data is used as conditional input to guide the separation process of the sound signal. Specifically, the 3D point cloud data provides the spatial position information of the device, and can better utilize the spatial information during the separation process to avoid the problems of sound source overlap or interference. In this way, the separated sound source signal has a higher signal-to-noise ratio and is more in line with the actual physical space layout.
[0039] Another object of the present invention is to provide an electronic device, including:
[0040] a processor; and
[0041] a memory, in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the method for separating the sound source of the main power transformation equipment by fusing 3D point cloud as described above.
[0042] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the method for separating the sound source of the main power transformation equipment by fusing 3D point cloud as described above.
[0043] In summary, the present invention provides a new method for separating multiple sound sources. By combining 3D point cloud with sound signals, it can significantly improve the accuracy of sound source localization and recognition. Compared with traditional sound separation methods, the present invention can more accurately locate the sound source with the help of 3D point cloud data, reduce noise interference, improve the reliability of fault diagnosis of power transformation equipment, and has broad application prospects, especially in the aspect of automatic monitoring and maintenance of power systems.
[0044] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for separating sound sources of main substation equipment integrating 3D point clouds, characterized in that: The method includes the following steps in sequence: (1) Deploy a high-precision three-dimensional laser scanner in a substation, scan the equipment area, obtain high-density point cloud data and perform preprocessing, and perform equipment segmentation to obtain three-dimensional point cloud data of different equipment; (2) Capture the sound signals generated by each equipment and perform preprocessing to extract the sound signal features; (3) Fuse the three-dimensional point cloud data of different equipment with the sound signal features to obtain fused features; (4) Improve the Transformer model to obtain a Transformer-based sound source separation network, input the fused features into the Transformer-based sound source separation network, perform sound signal separation, and obtain the separated sound signals; (5) Input the separated sound signals into the Transformer-based sound source separation network for multiple iterative optimizations of the separation results until the Transformer-based sound source separation network converges.
2. The method for separating sound sources of main substation equipment integrating three-dimensional point clouds according to claim 1, wherein: In step (1), the obtaining of high-density point cloud data and performing preprocessing and performing equipment segmentation specifically refers to: First, register the point cloud data in different coordinate systems, use the feature point registration method to convert the point cloud data from different scanning positions to the same coordinate system to obtain global three-dimensional point cloud data; adopt a denoising algorithm based on DBSCAN, introduce height features to distinguish electrical equipment and background noise; use an extraction method based on neighborhood feature aggregation to extract spatial features to achieve rapid segmentation of electrical equipment and obtain three-dimensional point cloud data of different equipment.
3. The method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds according to claim 1, wherein: Step (2) specifically includes the following steps in sequence: (2a) Deploy a microphone array near the key equipment in the substation to collect the sound signals of each equipment and its surrounding environment; (2b) Perform noise reduction processing on the sound signals, and use a noise reduction method to remove the noise and interference from the environment; the noise reduction method includes any one of wavelet transform, spectral subtraction, and adaptive filtering; (2c) Extract features from the noise-reduced sound signals to obtain sound signal features, and the sound signal features include spectrum, amplitude, and phase features.
4. The method for separating sound sources of main substation equipment integrating three-dimensional point clouds according to claim 1, characterized in that: Step (3) specifically refers to: Using the particle swarm optimization algorithm and the k-nearest neighbor classification algorithm, fuse the three-dimensional point cloud data with the sound signal features to obtain fused features.
5. The method for separating sound sources of main power transformation equipment integrating three-dimensional point clouds according to claim 1, characterized in that: In step (4), the improvement of the Transformer model to obtain a Transformer-based sound source separation network specifically refers to: According to the time-frequency domain characteristics of voiceprint, divide the subspace of the multi-head attention of the Transformer model into three different scales, and input the query, key, and value vectors into the scaled dot-product attention layers of three different scales through linear layers respectively. The scaled dot-product attention layer performs matrix multiplication, scaling, and normalized exponentiation on the query and key vectors respectively at each scale, and then performs matrix multiplication with the value vector to make different heads focus on features of different granularities. The three different scales include 32*32, 64*64, and 128*128.
6. An electronic device, comprising: A processor; And A memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds according to any one of claims 1-5.
7. A computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the method for separating sound sources of main power transformation equipment by fusing three-dimensional point clouds according to any one of claims 1-5.