Underdetermined pig blind source signal separation method based on sparsification theory
By employing sparsity theory and an improved AP clustering algorithm, the problem of separating mixed pig audio signals was solved, enabling the effective extraction of pig audio features and promoting welfare-oriented pig farming.
Patent Information
- Application Number
- CN202211183294.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-09-27
AI Technical Summary
In pig farming, the mixing of audio signals due to the multiple pigs being kept in pens is difficult to separate. Existing technologies are unable to effectively separate the signal components from the mixed pig audio signals, which affects the identification of pig status and welfare farming.
An underdetermined blind source signal separation method based on sparsity theory is adopted. By acquiring mixed pig audio signals, an underdetermined blind source separation model is constructed, and sparsification and single source point extraction are performed to obtain an estimated mixing matrix. Finally, the source signals are obtained, and the calculation is accelerated by an improved AP clustering algorithm and singular value decomposition. The minimum lp norm is optimized to complete the separation of audio signals.
The method effectively separates the source signal components of the mixed pig vocalization signal, improves the accuracy of pig audio feature extraction, and contributes to the healthy breeding of pigs.
Smart Images

Figure CN116030837B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of signal processing, and particularly relates to an underdetermined pig blind source signal separation method based on a sparsification theory. BACKGROUND
[0002] Pig audio contains rich information, which can reflect the behavior characteristics of pigs. By identifying pig audio, the state of pigs can be analyzed, and the welfare breeding of pigs can be promoted. The pig audio invention at home and abroad mainly focuses on endpoint detection and audio recognition under different states. However, when using modern information technology to monitor and identify pig audio, due to environmental and economic conditions, breeders often keep multiple pigs together, which leads to the collection of mixed audio emitted by multiple pigs, which is not conducive to feature extraction and recognition of audio. In order to separate each source signal component from the mixed pig audio signal as much as possible and extract effective features, blind source separation has become an effective solution.
[0003] Blind source separation (BSS) is to recover unknown source signals from mixed signals captured by multiple microphones, and according to the number of source signals being less than, equal to or greater than the number of microphones, blind source separation can be divided into three cases of overdetermined, determined and underdetermined. Because some theoretical algorithms are less practical, the underdetermined blind source separation problem is still challenging in blind source separation. In recent years, domestic and foreign methods based on sparse component analysis (SCA) are generally used to solve the underdetermined blind source separation problem; Chen donoho and Saunders obtain the sparse representation of signals in an overcomplete basis by solving a large-scale linear programming problem; Bofill proposes a two-step method based on SCA, and six source signals are separated from two mixed audio signals, and the method has low algorithm complexity and is easy to obtain global convergence value; Pando.G.Georgio and Fabian Theis use the sparse component analysis method to separate sparse signals, and compare with the method in the literature, and obtain better results. The importance of the sparsity of source signals for the SCA algorithm is self-evident, and many time-frequency domain extension algorithms are proposed to enhance and fully utilize the sparsity of signals. Zhen et al. found and proved that the time-frequency point dominated by a single source point is related to a one-dimensional subspace, and the estimation of the mixing matrix is obtained by using a hierarchical clustering algorithm, and the source signals are recovered by solving a series of least squares problems; Jourjine et al. propose a degenerate mixing estimation technique, and verify that the method can realize the separation of mixed source signals on speech signals and wireless signals; Yuan Xie et al. propose an improved information theory criterion method to detect the number of sources in the underdetermined case, estimate the mixing matrix by using a fourth-order tensor blind identification method, and reconstruct the source signals by using an lp norm diversity measure method, and good separation results are obtained, and the running speed is fast; Arberet et al. detect the time-frequency region of a single source point by using a statistical model of local confidence measure, and merge the information from all time-frequency regions according to the confidence thereof by using a DEMIX clustering algorithm, so as to complete the estimation of the mixing matrix; Yu and Xin propose a time-frequency two-step method for the underdetermined blind source separation problem of linear mixing of non-sparse signals, and the effectiveness and accuracy of the method are proved by numerical experiments and analysis. The audio test subjects of the underdetermined blind source signal separation at home and abroad are generally function signals, and there are few inventions for actual application audio signals. SUMMARY
[0004] The purpose of the present application is to provide an underdetermined live pig blind source signal separation method based on the theory of sparsification to solve the problems existing in the prior art.
[0005] To achieve the above purpose, the present application provides an underdetermined live pig blind source signal separation method based on the theory of sparsification, comprising:
[0006] Acquire mixed audio signals from live pigs;
[0007] Construct an underdetermined blind source separation model;
[0008] Based on the underdetermined blind source separation model, the mixed pig audio signal is sparsified and single source points are extracted to obtain single source points.
[0009] The estimated mixing matrix is obtained based on the single source point;
[0010] The source signal is obtained based on the underdetermined blind source separation model and the estimated mixing matrix;
[0011] The audio quality of the source signal is measured.
[0012] Optionally, the process of acquiring the mixed pig audio signal includes:
[0013] The system monitors the external environment based on default audio parameter values, sets the recording duration and audio parameters, and acquires sound signals.
[0014] The sound signal is compressed and transmitted over a network using compressed sensing technology to obtain a reconstructed sound signal.
[0015] The reconstructed sound signal is compared and analyzed with the original sound signal. Audio parameters are set for the reconstructed sound signal to obtain the reconstructed sound signal with the highest audio quality. Based on the audio parameters set for the reconstructed sound signal with the highest audio quality, a single audio signal is acquired, and a mixed pig audio signal is acquired based on the amplitude attenuation matrix.
[0016] Optionally, the underdetermined blind source separation model is:
[0017]
[0018] Among them, X N (t)=[x1(t),x2(t),...,x n [(t)] represents the observed signal vector, S m (tt nm ) indicates the time delay t nm The source signal vector arriving at the sensor, n(t) represents the noise, a nm This is the amplitude attenuation matrix, representing the signal attenuation coefficient.
[0019] Optionally, the process of sparsifying and extracting single-source points from the mixed pig audio signal based on the underdetermined blind source separation model includes:
[0020] Based on the underdetermined blind source separation model, a single-source point criterion is obtained, and a mixed single-source point is obtained based on the single-source point criterion.
[0021] Set a first threshold value based on the underdetermined blind source separation model, and preliminarily screen the mixed single source point based on the first threshold value;
[0022] Detect the mean and variance of the adjacent mixed single source point in the same frequency domain, set a second threshold value, and further screen based on the second threshold value to obtain a single source point;
[0023] Calculate the l2 norm of the single source point and set a third threshold value, and remove low-energy points in the single source point based on the l2 norm and the third threshold value.
[0024] Optionally, the process of obtaining an estimated mixing matrix based on the single source point comprises:
[0025] Based on the cosine distance, a similarity matrix is obtained, singular value decomposition and low-rank processing are performed on the similarity matrix, clustering is performed based on the AP algorithm, and the estimated mixing matrix is obtained.
[0026] Optionally, in the process of clustering the single source point based on the AP clustering algorithm, a dynamic damping coefficient adaptive method is used to avoid algorithm divergence.
[0027] Optionally, the dynamic damping coefficient adaptive method comprises:
[0028] Determine whether numerical oscillation occurs in the clustering process of the AP clustering algorithm, and when numerical oscillation occurs, adjust the damping coefficient based on an adjustment rule; the adjustment rule is:
[0029] λ = λ old + 0.01 λ ∈ [0.5, 1]
[0030] Wherein, λ old represents the damping coefficient value used in the last iteration, and λ represents the damping coefficient.
[0031] Optionally, the process of obtaining a source signal based on the underdetermined blind source separation model and the estimated mixing matrix comprises:
[0032] Based on the underdetermined blind source separation model and the estimated mixing matrix, an lp norm-based reconstruction algorithm is used for estimation to obtain the source signal.
[0033] Optionally, the process of measuring the audio quality of the source signal comprises:
[0034] Based on the source signal and the mixed pig audio signal, similarity coefficients, signal-to-noise ratios, and mean square errors are calculated respectively, and the similarity coefficients, the signal-to-noise ratios, and the mean square errors are measured based on the similarity coefficients, the signal-to-noise ratios, and the mean square errors
[0035] The technical effects of the present application are:
[0036] The present application takes the audio of pigs in various states as the subject of invention, and proposes an underdetermined pig blind source signal separation method based on the sparsification theory. Firstly, the pig audio signal is subjected to short-time Fourier transform, and the audio signal is converted to the time-frequency domain with stronger sparseness. The single source point is extracted through grouping, and the improved AP clustering algorithm is used to cluster the feature matrix constructed by the single source point to estimate the mixing matrix. Finally, the underdetermined pig blind source signal separation is completed by optimizing the minimum l p norm. The simulation test is carried out by using MATLAB, and it is found through the comparison of various measurement criteria and subjective listening that the method can effectively separate the source signal components of the mixed pig sound signal, provides a new scheme for the feature extraction of the mixed pig audio, and is helpful for the healthy breeding of pigs. BRIEF DESCRIPTION OF DRAWINGS
[0037] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of this application and its description together with the drawings make it clear that the application is not limited to the embodiments described and illustrated herein. In the drawings:
[0038] Figure 1 The hardware schematic diagram for the mixed audio signal in the embodiment of the present application is shown in the figure;
[0039] Figure 2 The flowchart for the mixed audio signal in the embodiment of the present application is shown in the figure;
[0040] Figure 3 The flowchart for the underdetermined pig blind source signal separation algorithm in the embodiment of the present application is shown in the figure;
[0041] Figure 4 The pig original audio signal waveform diagram in different states in the embodiment of the present application is shown in the figure;
[0042] Figure 5 The pig observed audio signal waveform diagram in the embodiment of the present application is shown in the figure;
[0043] Figure 6 The observed signal scatter plot in the time-frequency domain in the embodiment of the present application is shown in the figure, wherein (a) is the observed signal 1, and (b) is the observed signal 2;
[0044] Figure 7 The observed signal and single source point real part scatter plot in the time-frequency domain in the embodiment of the present application is shown in the figure, wherein (a) is the observed signal, and (b) is the single source point;
[0045] Figure 8 The different damping coefficients and clustering results diagram in the embodiment of the present application is shown in the figure;
[0046] Figure 9 The improved AP clustering result diagram in the embodiment of the present application is shown in the figure;
[0047] Figure 10The reconstructed pig audio signal waveform diagram in the embodiment of the present application;
[0048] Figure 11 The reconstruction index diagram under different numbers of source signals and observation signals in the embodiment of the present application, wherein (a) is the average similarity coefficient, (b) is the average signal-to-noise ratio, and (c) is the average mean square error. DETAILED DESCRIPTION
[0049] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0050] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0051] Embodiment one
[0052] As Figures 1-11 shown, the present embodiment provides an underdetermined pig blind source signal separation method based on the sparsification theory, comprising:
[0053] Test materials and acquisition method
[0054] The pig sound frequency used in the test is mainly from Anhui Mengcheng Jinghuimeng Pig Farm, and is obtained in a relatively quiet space. Sound collection and transmission is to use the embedded microprocessor NanoPCT4 with cortex architecture as the main controller, and to independently design and realize the hardware system for sound collection and transmission. The iTalk-02 microphone, USB interface and USB device interface resources are externally connected. The audio collection module produced by Tyless Company is used as the sound data collection equipment, audio or PCM coding is supported, and raw and wav audio output formats are supported. The RAM chip is the voice signal collected through the UDP protocol for network transmission. The physical diagram of the collection and transmission platform is as shown in Figure 1 , and the hardware platform flowchart is as shown in Figure 2 , and the specific steps are as follows:
[0055] (1) power on the development board and initialize the USB microphone device;
[0056] (2) the remote controller (PC or server) creates a thread of the microphone collection system and the IP and port number of the UDP communication protocol;
[0057] (3) detect the index number of the USB microphone device to start the device;
[0058] (4) Detect whether the USB microphone device is successfully turned on, if not, return to step 3 to re-detect until successful detection;
[0059] (5) Sound monitoring is performed on a specific environment by using default audio parameter values, and if the sound meets the standard of sound collection, sound recording is determined;
[0060] (6) The remote controller sets the recording duration of the audio and the audio parameters;
[0061] (7) Based on the UDP protocol, the compressed sensing technology is used to compress the collected sound signal for network transmission;
[0062] (8) The reconstructed speech signal is compared and analyzed with the original sound signal at the receiving end, and it is analyzed whether the sound signal corresponding to the highest audio quality is the sound signal, if yes, the output is ended, if not, the audio parameters are returned to step 6 to be reset, until the best speech signal quality parameter is reached. According to the above eight steps, a complete hardware platform flow chart is designed for sound collection, transmission and reconstruction. The parameters of the reconstructed audio with the highest quality are used to collect the pig audio signal, and the mixed audio signal generated by setting the amplitude attenuation matrix is used as the observation signal.
[0063] Underdetermined blind source separation model
[0064] The linear instantaneous mixing model of the underdetermined blind source separation (UBSS) problem can be expressed as:
[0065]
[0066] In the formula, X N (t) = [x1(t), x2(t),..., x n (t)] represents an observation signal vector, S m (t-t nm ) represents a source signal vector arriving at the sensor after a time delay t nm , n(t) represents noise, and a nm is an amplitude attenuation matrix, which represents the attenuation coefficient of the signal. The influence of noise is not considered in the present application.
[0067] Audio signal sparsification and single source point extraction
[0068] The sparsity of signals means that the amplitude is zero in most time and is larger in a small part of time. According to the analysis of recoverability of mixed signals in the literature, the more sparse the signals are in the time domain or the transform domain, the higher the probability that each source signal can be correctly separated. Since the signals are more sparse in the time-frequency domain, the application adopts short-time Fourier transform to convert the mixed pig audio signals into the time-frequency domain.
[0069] The property of sparse signals determines that the possibility of two source signals taking non-zero values at the same time is very small. According to the short-time stationary property of non-stationary signals, there is a single-source point neighborhood U(t, f) with a constant frequency and adjacent time, and the points in the neighborhood are all generated by the same source signal s i If enough single-source points can be extracted from the mixed signals, the scatter diagram composed of the single-source points will clearly gather near N straight lines, and the estimation of the mixing matrix can be realized by using a clustering algorithm. Without considering noise, formula (1) can be expanded as:
[0070]
[0071] It is assumed that at time t, only the source signal s i (t) takes a large value, and formula (2) can be approximated as:
[0072]
[0073] In formula (3), x i (t,f) is the complex representation of the ith observed signal in the time-frequency domain, and Re(.) and Im(.) represent the real part and the imaginary part, respectively. According to formula (3), all the time points at which the source signal s i (t) takes a non-zero value will determine a straight line with a direction of the ith column vector of the mixing matrix A, and the ratio of the real part to the imaginary part of the single-source point is a constant value, so formula (4) can be used as a criterion for the single-source point:
[0074]
[0075] However, due to the influence of noise and calculation errors and a large number of low-energy points (points gathered near zero), the spatial distribution of the single-source points extracted by formula (4) will deviate from the column direction corresponding to the mixing matrix A, resulting in a large error in the estimated mixing matrix. To solve this problem, the application first uses formula (5) to preliminarily screen the single-source points by relaxing the constraint condition:
[0076]
[0077] In formula (5), ε1 is a set threshold. The mean value (formula (6)) and the variance var[x of the adjacent M points in the same frequency domain are detected M](t,f)(formula (7)) to further screen for single source points with variance less than threshold value ε2, and finally calculate the l2 norm of the single source points ||x(t,f) i ||2(formula (8)) to remove low-energy points with a value less than a set threshold σ.
[0078]
[0079]
[0080] ||x(t,f) i ||2<σ (8)
[0081] Mixed matrix estimation
[0082] After obtaining sufficient single source points, a clustering algorithm can be used to cluster the feature points to estimate the mixed matrix. Traditional clustering algorithms include K-means, fuzzy clustering (FC-means), etc. Although these classic methods have high accuracy and fast calculation speed, the clustering results are sensitive to the initial clustering center, and the number of clusters needs to be known, which does not match the actual blind source separation where the number of source signals is unknown. In contrast, the affinity propagation clustering algorithm does not need to specify the number of final clustering families, the cluster center points are existing data points, and the squared error of the results is small, which just makes up for this problem. AP clustering regards all sample points as potential clustering centers, selects the center points through cyclic iteration, and obtains the optimal class representative cluster. However, the clustering results are affected by the damping coefficient and the algorithm complexity is high, therefore, the present application combines singular value decomposition and proposes an AP clustering algorithm with an adaptive damping coefficient to estimate the mixed matrix.
[0083] Singular value decomposition
[0084] Since the extracted single source points are generally 20%-40% of the total sample points, this makes the matrix dimension obtained by AP clustering when constructing the similarity matrix large, requiring a large amount of memory and having high computational complexity. In order to speed up the calculation speed of the algorithm, the present application introduces singular value decomposition (SVD) for dimension reduction to reduce complexity, and the singular value decomposition of matrix A is defined as follows:
[0085] A′=UΣV T (9)
[0086] In the formula, Σ is an N-order diagonal matrix, the elements on the diagonal line are singular values σ i sorted from small to large; U and V are N-order orthogonal matrices, the column elements of U are left singular vectors, and the column elements of V are right singular vectors, which are respectively obtained by A′A′ Tand A' T The eigenvector composition of A'. Since the similar matrix constructed by the present application is a symmetric matrix, formula (9) can be written as:
[0087] A' = UΣU T (10)
[0088] After singular value decomposition, the low rank approximation operation is performed on matrix A, and the largest k singular values in Σ are retained. The value of k is set according to formula (11), and the loss rate ER is defined as:
[0089]
[0090] When ER is less than 10%, it is considered that the value of k is reasonable. The remaining singular values are set to 0, and the left and right corresponding singular vectors are combined to approximate the matrix A:
[0091] A' N×N = U N×N Σ N×N U T N×N ≈ U k×k Σ k×k U T k×k = A" (12)
[0092] After the low rank approximation processing of formula (12), A becomes a matrix A" with rank k.
[0093] AP clustering algorithm
[0094] The AP clustering algorithm takes the log likelihood as the similarity degree between sample points, and generally uses the negative Euclidean distance to calculate the similarity between sample points. However, the Euclidean distance is easily affected by the dimension, and cannot reflect the characteristics of the feature points in the direction. Therefore, the present application introduces the negative cosine distance to construct the feature similarity matrix, and the calculation formula is:
[0095]
[0096] In the formula, x i and x k are the i-th and k-th points.
[0097] The elements on the diagonal of the similarity matrix are bias parameters p, and the sample points with larger values are easy to select as cluster centers (called examples). The median of all similarity values is extracted and assigned to all elements on the main diagonal of S to ensure that each data point has an equal probability of becoming an example.
[0098] To find the proper cluster centers, the attraction matrix R(i, k) is defined to describe the degree that point k is suitable to be the cluster center of point i, and the membership matrix A(i, k) is defined to describe the degree that point i chooses point k as the cluster center. A zero matrix of proper size is initialized for R and A, and the optimal cluster centers are found by updating the messages between sample points through the membership and attraction messages. The update rules of the attraction matrix R and the membership matrix A are shown in equations (14) and (15) respectively:
[0099]
[0100]
[0101] where t represents the current iteration number, and i, k are the index values of different rows and columns. Equation (14) shows that any candidate cluster center can affect other candidate cluster centers and compete for the membership of other points. In the first iteration, since the initial value of A is zero, the update of R does not consider the influence of other points on the candidate examples. In the following iterations, when some points are effectively assigned to other examples, their membership values will decrease to negative values according to the update rule of equation (15), which will reduce the effective value of the input similarity in equation (14) and remove the corresponding candidate samples in the competition. If R(k, k) is finally negative, it means that point k is more suitable to belong to other examples than to be a cluster center itself. In equation (15), the update rule of the membership A(i, k) is the self-attraction plus the positive attraction obtained from other points. Here, only the positive (numerical value is positive) attraction is added because only the positive attraction supports point k as a cluster center. The value of the self-membership A(k, k) is the sum of the positive attraction obtained from other points. If A(k, k) is negative, it means that point k is more suitable to belong to another example than to be a cluster center itself.
[0102] Due to the numerical oscillation in the process of updating the messages, the algorithm is not easy to converge. Therefore, a damping factor (DF) is introduced to attenuate the attraction information and the membership information. The formulas (16) and (17) are used to update R and A:
[0103] R = (1 - λ) × R + λ × R old (16)
[0104] A = (1 - λ) × A + λ × A old (17)
[0105] where R old represents the attraction matrix updated last time; A old represents the membership matrix updated last time; λ ∈ [0, 1] represents the damping factor.
[0106] The algorithm is terminated by setting the maximum iteration number m, and the iteration termination number n is set, that is, the cluster center does not change after n consecutive iterations before the maximum iteration number m is reached, at this time, it is considered that the algorithm has converged, and the cluster center has been determined.
[0107] Adaptive damping coefficient method
[0108] The damping coefficient λ takes different values, which affects the global and local search ability of the algorithm, and then interferes with the convergence performance of the algorithm. When the traditional AP clustering is performed, the damping coefficient is often set to a fixed value based on prior experience, which makes the algorithm unable to dynamically adjust the search performance at different stages. Therefore, the present application provides a dynamic damping coefficient adaptive method.
[0109] A moving window with a length of l is used to compare whether the number of clusters at the current iteration and the number of clusters at the last iteration decreases or is consistent, 1 if yes, otherwise 0. Considering the instability of the initial stage of the algorithm and the occasional small shock, it is considered that more than 2 / 3 of the records show 0 when the shock occurs. At this time, the damping coefficient λ is adjusted, and the initial value of λ is set to the system default value 0.5 considering the convergence of the algorithm. When the maximum value is reached, it is no longer increased. The specific adjustment rule is:
[0110] λ = λ old + 0.01 λ ∈ [0.5, 1] (18)
[0111] In the formula, λ old represents the damping coefficient value used in the last iteration.
[0112] Pig audio source signal reconstruction
[0113] Since the mixing matrix estimated under the underdetermined condition is a non-full rank matrix, the source signal cannot be directly reconstructed by the estimated matrix. The present application adopts a method based on sparsity to reconstruct the source signal. Considering the instantaneous linear mixing model shown in formula (1), under the condition that the mixing matrix a has been estimated, the estimation problem of the sparse source signal S can be transformed into the following optimization problem:
[0114]
[0115] In the formula, is the estimation of the source signal S, and J p(.) is a certain sparsity measure function of the signal, the smaller the value, the stronger the sparsity of the signal. Among various sparsity measure functions, the l0 norm is the best one to reflect the sparsity, which refers to the number of non-zero values of the signal, but it is not practically valuable because its solution is too sensitive to noise; the reconstruction algorithm based on the l p norm is better than the optimization algorithm based on the l0 norm and the l1 norm in terms of audio quality and reliability, for a vector s, the l p norm is calculated as follows:
[0116]
[0117] In the formula, p is a set value. This section completes the reconstruction of the live pig audio based on the l p norm.
[0118] Assuming that the number of source signals and observation signals are M and N respectively, the l p norm solution has at most M non-zero values, for each sampling point, there are possible solutions, by comparing the l p norm of these possible solutions, the minimum l p norm solution can be obtained. The whole algorithm steps can be briefly described as follows:
[0119] 1) find the N×N submatrix
[0120] 2) for a certain time t, solve the possible solution of the l p norm minimization problem:
[0121]
[0122] 3) calculate formula (21) corresponding l p norm J k :
[0123]
[0124] 4) determine the minimum l p norm solution according to formula (23) and take it as the estimation of :
[0125]
[0126] 5) repeat steps 2-4 until the of all time is obtained, and the estimation of the source signal is obtained.
[0127] Measurement index
[0128] In order to measure the audio quality reconstructed by the algorithm, the present application introduces the similarity coefficient, signal-to-noise ratio and mean square error.
[0129] Similarity coefficient ξ ij Take the similarity coefficient of the separated output signal y i and the source signal s j as the measurement of blind source separation performance. Its calculation formula is:
[0130]
[0131] In the formula, ξ ij The value range is [0, 1], when ξ ij =1, it means that the i-th separated signal and the j-th source signal are completely the same, when ξ ij =0, y i and s j are independent of each other, the greater the value of ξ ij , the more similar they are.
[0132] The signal-to-noise ratio refers to the ratio of signal to noise in the system, and the present application uses the signal-to-noise ratio to describe the distortion degree of the reconstructed signal compared with the source signal, and its calculation formula is:
[0133]
[0134] The higher the value of SNR indicates the better effect.
[0135] The mean square error is the mean value of the square sum of the corresponding point error of the predicted data and the original data, and its calculation formula is:
[0136]
[0137] The value of MSE indicates the difference between the source signal and the reconstructed signal, and the smaller the value indicates the better effect.
[0138] Results and analysis
[0139] According to the algorithm, for the underdetermined pig blind source separation problem, some experiments are carried out in the MATLAB simulation environment, 3 source signals and 2 observation signals are used to illustrate the experiment process, and the performance of the algorithm is verified by comparing the measurement indexes of different number of source signals and observation signals.
[0140] Underdetermined pig blind source separation under source signal and 2 observation signals
[0141] In order to verify the generality of the experiment, about 10s of continuous pig audio signals in different states are selected, Figure 4For the waveform chart of the pre-processed live pig audio signals (hum, snort, roar) in different states, this section takes the three audio signals as the source signals, obtains two mixed audio signals as the observation signals through the artificially set amplitude attenuation matrix, and illustrates the entire test process. Figure 5 The audio observation signals after the amplitude attenuation matrix A (formula (27)) and zero padding alignment are shown.
[0142] The short-time Fourier transform is performed on the observation signals, the Hanning window is selected as the window function, the window size is set to 512, and the window overlap is 256, to obtain the complex matrix of the two observation signals and visualize them as Figure 6 M = 6, ε1 = 0.01, ε2 = 0.05, and σ = 0.5 are set. Figure 7 The real part comparison scatter plot of the observation signals before and after the extraction of the single source point can be directly observed that after the single source point screening by the method of the application, the amplitude of the signal clearly presents three straight lines in the two-dimensional plane, and the low-energy points are basically eliminated.
[0143]
[0144] The improved AP clustering algorithm is used to cluster the extracted feature single source points, the maximum iteration number is set to 500 during the test, the iteration termination number is set to 50, the adaptive rule is used to adjust the damping coefficient, the window length l is set to 6, the initial damping coefficient λ is set to 0.5, and the final clustering number before each damping coefficient adjustment is recorded, Figure 8 The change of the damping coefficient and the clustering result when the iteration number gradually increases is shown, when λ is the initial value 0.5, the clustering result is larger, and the numerical oscillation is large, with the continuous increase of the damping coefficient, the clustering number also changes continuously, when the value of λ increases to 0.67, the clustering number tends to be stable. Figure 9 The real part clustering result of the observation signal by the improved AP algorithm is shown, and it can be clearly seen that the characteristic points are clustered into one class in a straight line, and a total of three classes are clustered, which are represented by different colors.
[0145] The audio signals are separated from the mixed audio signals according to the above method, and Table 1 shows the average signal-to-noise ratio of the reconstructed audio signals under different p values, when the p value is selected as 0.8, the separated waveform is optimal, and the average signal-to-noise ratio value is maximum, therefore, this section selects p as 0.8 to complete the audio reconstruction. Figure 10It can be seen that the audio arrangement order after reconstruction is not consistent with the source signal input order, the prior art solves the ordering ambiguity problem by frequency clustering, however the focus of the present application is on the "two-step method" of underdetermined pig audio signal blind source separation, therefore the ordering problem is not discussed here; from the waveform, the reconstructed signal is generally consistent with the source signal, but there is a slight difference. In order to measure the quality of the reconstructed audio, the similarity coefficient, signal-to-noise ratio and mean square error of the source signal and the corresponding observed signal are measured (Table 2), overall, the similarity coefficient of the reconstructed three signals and the source signal is above 0.9, the signal-to-noise ratio is between 8-10 dB, and the mean square error is between 0.018-0.028, and the subjective listening effect is good.
[0146] Table 1
[0147]
[0148] Table 2
[0149]
[0150] Comparison of audio reconstruction indicators under different numbers of source and observed signals
[0151] In order to measure the performance of the algorithm, the present application selects a variety of pig state audio source signals, constructs different numbers of observed signals, performs underdetermined pig blind source signal separation, and compares them, and the measurement index results are shown as Figure 11 , where the x-axis coordinate number is "source signal number-observation signal number", and the value displayed on the y-axis is the average value of the corresponding evaluation indicators measured for all separated signals and source signals.
[0152] It can be seen from the figure that the audio quality indexes separated are different for different numbers of source signals and observation signals, and when the number of source signals is certain, the more the number of observation signals, the better the quality index measured by each method, and the more reliable the separated audio. On the similarity coefficient, the traditional method is 0.778-0.939 and 0.755-0.927 respectively, and the value measured by the method proposed in the application is 0.785-0.957; on the signal-to-noise ratio, the traditional method is 7.268-10.017 dB and 7.568-9.897 dB respectively, and the value measured by the method proposed in the application is 7.468-10.347 dB; on the average mean square error, the traditional method is 0.021-0.113 and 0.025-0.135 respectively, and the value measured by the method proposed in the application is 0.019-0.092; overall, the value of the method proposed in the application is higher on the average similarity coefficient and the average signal-to-noise ratio, and the value is lower on the average mean square error, which is better than the existing method. The traditional technology in the embodiment includes: document 30 (He X, He F, Cai W. Underdetermined BSS based on K-means and AP clustering. Circuits, Systems, and Signal Processing, 2016, 35(8): 2881-2913), document 31 (Ji Ce, Jiang Yutian. Underdetermined blind source separation algorithm based on direction amplitude ratio. Journal of Northeast University (Natural Science Edition))
[0153] The above is only the preferred specific embodiment of the application, but the protection scope of the application is not limited to this, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the application, which should be covered in the protection scope of the application. Therefore, the protection scope of the application should be subject to the protection scope of the claims.
Claims
1. A method for underdetermined blind source separation of piglet signals based on the theory of sparsification, characterized in that, The method comprises the following steps: acquiring a mixed pig audio signal; constructing an underdetermined blind source separation model; sparsifying the mixed pig audio signal based on the underdetermined blind source separation model and extracting a single source point to obtain the single source point; obtaining an estimated mixing matrix based on the single source point; obtaining a source signal based on the underdetermined blind source separation model and the estimated mixing matrix; measuring the audio quality of the source signal; the process of sparsifying the mixed pig audio signal based on the underdetermined blind source separation model and extracting a single source point comprises: obtaining a single source point criterion based on the underdetermined blind source separation model, and obtaining a mixed single source point based on the single source point criterion; setting a first threshold based on the underdetermined blind source separation model, and preliminarily screening the mixed single source point based on the first threshold; detecting the mean and variance of adjacent mixed single source points in the same frequency domain, setting a second threshold, and further screening based on the second threshold to obtain a single source point; calculating the l2 norm of the single source point and setting a third threshold, and removing low-energy points in the single source point based on the l2 norm and the third threshold; the process of obtaining an estimated mixing matrix based on the single source point comprises: obtaining a similarity matrix based on the cosine distance, singular value decomposition and low-rank processing of the similarity matrix, clustering based on the AP algorithm, and obtaining the estimated mixing matrix; in the process of clustering the single source point based on the AP clustering algorithm, a dynamic damping coefficient adaptive method is used to avoid algorithm divergence; the dynamic damping coefficient adaptive method comprises: judging whether numerical oscillation occurs in the clustering process of the AP clustering algorithm, and adjusting the damping coefficient based on an adjustment rule when numerical oscillation occurs; the adjustment rule is: wherein denotes the damping coefficient value used in the last iteration, and λ denotes the damping coefficient.
2. The underdetermined blind source separation method of live pigs based on the theory of sparsification according to claim 1, characterized in that, the process of acquiring a mixed pig audio signal comprises: listening to the external environment based on default audio parameter values, setting the recording duration and audio parameters, and obtaining a sound signal; compressing the sound signal based on the compression sensing technology to obtain a reconstructed sound signal; comparing and analyzing the reconstructed sound signal and the sound signal, setting the audio parameters of the reconstructed sound signal, obtaining the reconstructed sound signal with the highest audio quality, acquiring the mixed pig audio signal based on the audio parameters corresponding to the reconstructed sound signal with the highest audio quality, and obtaining an observation signal based on the amplitude attenuation matrix.
3. The underdetermined blind source signal separation method for pig based on the theory of sparsification according to claim 1, characterized in that, the underdetermined blind source separation model is: wherein represents an observation signal vector, represents a time delay a source signal vector arriving at the sensor, represents noise, is an amplitude attenuation matrix, representing attenuation coefficients of the signals.
4. The underdetermined blind source signal separation method for pig based on the theory of sparsification according to claim 1, characterized in that, the process of obtaining a source signal based on the underdetermined blind source separation model and the estimated mixing matrix comprises: Based on the underdetermined blind source separation model and the estimated mixing matrix, a reconstruction algorithm based on the norm is used to estimate and obtain the source signals. Based on the underdetermined blind source separation model and the estimated mixing matrix, a reconstruction algorithm based on the norm is used to estimate and obtain the source signals.
5. The underdetermined blind source signal separation method for pig based on the theory of sparsification according to claim 1, characterized in that, the process of measuring the audio quality of the source signal comprises: calculating the similarity coefficient, signal-to-noise ratio and mean square error based on the source signal and the mixed pig audio signal respectively, and measuring based on the similarity coefficient, signal-to-noise ratio and mean square error.