English phoneme m identification method based on spline interpolation wavelet neural network

Through the method based on spline interpolation wavelet neural network, the problem that traditional wavelet transform is difficult to extract phoneme signal characteristics in high-resolution wavelet space is solved, and the calculation speed and accuracy of phoneme recognition are improved, and stable phoneme signal characteristic parameters are provided.

CN120452422APending Publication Date: 2025-08-08UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510645564.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional wavelet transforms are difficult to fully characterize phoneme signal characteristics in high-resolution wavelet space, resulting in problems of slow calculation speed, low accuracy and stability.

Method used

A method based on spline interpolation wavelet neural network is adopted to construct a wavelet neural network through sampling, normalization and threshold processing, and a training feedback matrix is constructed using the mapping coefficients of sixth-order spline wavelets and ordinary wavelets to extract and recognize phoneme signal features.

Benefits of technology

It improves the calculation speed and accuracy of phoneme recognition, improves the stability of detection, and provides complete phoneme signal characteristic parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure QLYQS_4
    Figure QLYQS_4
Patent Text Reader

Abstract

The invention relates to an English phoneme m recognition method based on a spline interpolation wavelet neural network, and solves the problems that a traditional recognition method is low in calculation speed and low in precision and stability. Audio input is converted into an analyzed discrete signal set by using a sample-and-hold technology, and a standardized audio data set is obtained through normalization and threshold processing to serve as input of a neural network; a six-order spline interpolation wavelet is selected, and a criterion function of a neural network is constructed according to mapping coefficients of the interpolation wavelet and a common wavelet; taking a wavelet as an excitation function of the neural network; constructing a training feedback matrix by using the mapping coefficients of the interpolation wavelet and the common wavelet; the wavelet neural network training has global convergence, and an output layer weight is a corresponding wavelet coefficient; performing discrete Fourier transform on the wavelet coefficient to obtain a frequency spectrum of the wavelet coefficient, and identifying a feature point and a distribution characteristic thereof as a signal identification condition; and judging whether the input phonemes contain consonants m or not by comparing the feature points with the distribution characteristics of the feature points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of signal analysis, and in particular to an English phoneme m recognition method based on spline interpolation wavelet neural network. Background Art

[0002] Phonemes are the smallest units of speech, so determining the signal characteristics of different phonemes is a crucial step in speech recognition. This makes the application of wavelet analysis to determine the signal characteristics of different phonemes a crucial technical foundation for speech recognition. Research has shown that the wavelet components of different phonemes are distributed in wavelet spaces of varying resolutions, and their corresponding wavelet coefficients exhibit distinct distribution characteristics. Therefore, the wavelet spatial resolution where the phoneme energy is concentrated, and the corresponding wavelet coefficient distribution characteristics, become crucial signal characteristics for phoneme identification. However, due to the limitations of the traditional wavelet transform (CWT) algorithm, which is limited by the continuous integration algorithm and cannot extract the corresponding wavelet system within the high-resolution wavelet space, it is difficult to fully characterize the signal characteristics of phonemes. Summary of the Invention

[0003] The purpose of this invention is to propose an English phoneme m recognition method based on spline interpolation wavelet neural network, which reduces the computational complexity of audio signal processing and improves detection accuracy, thereby achieving the effect of optimizing calculation. The present invention is achieved in that: The specific steps are as follows: Step 1: Sample the audio signal and convert it into a digital signal, and normalize it; Step 1.1: Sample the input audio signal as a set of discrete signals ; Step 1.2: Take a discrete signal set The element with the maximum absolute value is recorded as : (1) Step 1.3: Set the audio elements read Perform maximum absolute value normalization processing and record it as the normalized signal set ; (2) Step 1.4: Normalize the signal set Perform threshold processing and set a custom threshold value as ;like ,but , record the obtained array as the audio data set ,in, represents the kth audio signal after threshold processing, and N represents the number of audio data; Step 2: Construct a wavelet neural network; Set the sixth-order spline wavelet , which is used as the hidden layer neuron of the wavelet neural network. As the hidden layer neuron threshold, As the input layer weight, is the output layer weight; and then the neural network output is obtained ; (3) in Indicates the amount of time, is the output layer weight The elements in , the calculation method of which is given later; Step 3: Setup Represents the wavelet space where the audio signal features need to be extracted, In the process of signal feature extraction, each is set as , let determine the input layer weight ; (4) Step 4: Determine the parameters in the formula and ; Set the recorded audio duration to ,but , (5) in As shown in formula (4); Step 5: Construct the training feedback matrix ; Step 5.1: Construct the vector ; (6) in As shown in formula (4), That is, the number of audio values in step 1.3. , , ; Step 5.2: Construct the matrix ; (7) in and As shown in formula (5), As shown in formula (6), As shown in formula (4); Step 5.3: Construct the feedback matrix ; (8) Step 6: Construct the criteria function ; Step 6.1: Construct a relation of wavelet in frequency domain; (9) in The sixth-order spline wavelet used in this invention The frequency domain expression of the Fourier transform is: Indicates frequency; Step 6.2: Use the inverse discrete Fourier transform to calculate the formula and obtain the coefficients , and construct the vector , ; Step 6.3: Setup Represents the obtained audio data ; (10) Indicated by The column vector formed by Indicates the number of audio data; Step 6.4: Set the output layer weights The matrix expression of is: (11) Here the symbol express The transpose of is the output layer weight The elements in ; Step 6.5: Set the error on the training set ; (12) here As shown in formula (7); Step 6.6: Constructing the Criterion Function ; (13) (14) Step 7: Determine the parameters in the formula through neural network training ; Step 7.1: Convert audio data and feedback matrix Substitute the following training process: (15) here In expression (10) In the The value in the training step iteration, In expression (10) In the The value in the step iteration training; Step 7.2: Substitute into formula (13) to calculate ; Step 7.3: Substitute into formula (14) to calculate ; Step 7.4: If , training stops; Step 7.5: According to formula (11), obtain the value of formula (15) when training stops of The value is the characteristic wavelet coefficient of the m sound; Step 8.1: Retain the characteristic wavelet coefficients in the third, fifth, and seventh wavelet spaces gather , , ,in , , Respectively represent , , The number of elements in Step 8.2: Perform discrete Fourier transform on the obtained wavelet coefficients. (16) (17) (18) The wavelet coefficient set in the t-th scale space is recorded as , the nth wavelet coefficient is , the data length is , , k is less than A positive integer; Step 8.3: The length calculated in step 8.4 is Array ; For arrays respectively The elements in are processed according to formula (19) and recorded as arrays Elements , called feature points: , (19) in, For arrays Middle Number, For arrays Middle Number, For arrays Middle The number here , for Array length; Step 8.4: Take the union of the feature point intervals of the phoneme m samples produced by multiple groups of human voices, and obtain the feature information of the standard English phoneme m in different scale spaces as shown in the following table; Spatial feature points are distributed in the negative semi-axis range The third space (-2.80, -2.72) ∪ (-1.12, -0.91) ∪ (-0.25, 0) The fifth space (-2.30, -1.75) ∪ (-0.81, -0.79) Seventh Space (-2.4,0) Step 9: Detect whether the unknown audio signal contains phoneme m; Step 9.1: Input a continuous audio signal to be detected and repeat the above steps; Step 9.2: Calculate the sum of all feature points in different scale spaces; , (20) represents the sum of the feature points in the t-th scale space, ; Step 9.3: Due to the symmetry of the interpolation wavelet, determine the sum of the eigenvalues of the third scale space Is it less than 0.01? If so, the error range condition is met and the next step is performed. Otherwise, the output cannot determine whether the audio signal contains phoneme m. Step 9.4: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the third space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.5: Judgment Is it less than 0.01? If so, further judgment is required; otherwise, it is determined that the audio signal does not contain phoneme m; Step 9.6: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the fifth scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.7: Judgment If it is less than 0.01, further judgment is required, otherwise it is judged that the audio signal does not contain phoneme m; Step 9.8: Note The non-zero interval of ,The negative semi-axis distribution range of the feature points in the seventh scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.9: Output the judgment result "the audio signal contains phoneme m" or "the audio signal does not contain phoneme m" according to steps 21 to 26.

[0004] The advantages of the present invention are as follows:

[0005] This invention uses a spline interpolation wavelet neural network to fully construct the signal characteristics of the phoneme m, providing important technical parameters (wavelet spatial resolution and wavelet coefficient distribution characteristics) for speech recognition containing the phoneme m. Compared with traditional audio signal processing methods, the algorithm of this invention significantly improves computational speed, accuracy, and stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 A flowchart of phoneme recognition according to the present invention; Figure 2 This is a flow chart of extracting characteristic wavelet coefficients from the neural network of the present invention. DETAILED DESCRIPTION

[0007] Step 1: Sample the audio signal and convert it into a digital signal, and normalize it; Step 1.1: Sample the input audio signal as a set of discrete signals ; Step 1.2: Take a discrete signal set The element with the maximum absolute value is recorded as : (1) Step 1.3: Set the read audio elements Perform maximum absolute value normalization processing and record it as the normalized signal set ; (2) Step 1.4: Normalize the signal set Perform threshold processing and set a custom threshold value as ;like ,but , record the obtained array as the audio data set ,in, represents the kth audio signal after threshold processing, and N represents the number of audio data; Step 2: Construct a wavelet neural network; Set the sixth-order spline wavelet , which is used as the hidden layer neuron of the wavelet neural network. As the hidden layer neuron threshold, As the input layer weight, is the output layer weight; and then the neural network output is obtained ; (3) in Indicates the amount of time, is the output layer weight The elements in , the calculation method of which is given later; Step 3: Setup Represents the wavelet space where the audio signal features need to be extracted, In the process of signal feature extraction, each is set as , let determine the input layer weight ; (4) Step 4: Determine the parameters in the formula and ; Set the recorded audio duration to ,but , (5) in As shown in formula (4); Step 5: Construct the training feedback matrix ; Step 5.1: Construct the vector ; (6) in As shown in formula (4), That is, the number of audio values in step 1.3. , , ; Step 5.2: Construct the matrix ; (7) in and As shown in formula (5), As shown in formula (6), As shown in formula (4); Step 5.3: Construct the feedback matrix ; (8) Step 6: Construct the criteria function ; Step 6.1: Construct a relation of wavelet in frequency domain; (9) in The sixth-order spline wavelet used in this invention The frequency domain expression of the Fourier transform is: Indicates frequency; Step 6.2: Use the inverse discrete Fourier transform to calculate the formula and obtain the coefficients , and construct the vector , ; Step 6.3: Setup Represents the obtained audio data ; (10) Indicated by The column vector formed by Indicates the number of audio data; Step 6.4: Set the output layer weights The matrix expression of is: (11) Here the symbol express The transpose of is the output layer weight The elements in ; Step 6.5: Set the error on the training set ; (12) here As shown in formula (7); Step 6.6: Constructing the Criterion Function ; (13) (14) Step 7: Determine the parameters in the formula through neural network training ; Step 7.1: Convert audio data and feedback matrix Substitute the following training process: (15) here In expression (10) In the The value in the training step iteration, In expression (10) In the The value in the step iteration training; Step 7.2: Substitute into formula (13) to calculate ; Step 7.3: Substitute into formula (14) to calculate ; Step 7.4: If , training stops; Step 7.5: According to formula (11), obtain the value of formula (15) when training stops of The value is the characteristic wavelet coefficient of the m sound; Step 8.1: Retain the characteristic wavelet coefficients in the third, fifth, and seventh wavelet spaces gather , , ,in , , Respectively represent , , The number of elements in Step 8.2: Perform discrete Fourier transform on the obtained wavelet coefficients. (16) (17) (18) The wavelet coefficient set in the t-th scale space is recorded as , the nth wavelet coefficient is , the data length is , , k is less than A positive integer; Step 8.3: The length calculated in step 8.4 is Array ; For arrays respectively The elements in are processed according to formula (19) and recorded as arrays Elements , called feature points: , (19) in, For arrays Middle Number, For arrays Middle Number, For arrays Middle The number here , for Array length; Step 8.4: Take the union of the feature point intervals of the phoneme m samples produced by multiple groups of human voices, and obtain the feature information of the standard English phoneme m in different scale spaces as shown in the following table; Spatial feature points are distributed in the negative semi-axis range The third space (-2.80, -2.72) ∪ (-1.12, -0.91) ∪ (-0.25, 0) The fifth space (-2.30, -1.75) ∪ (-0.81, -0.79) Seventh Space (-2.4,0) Step 9: Detect whether the unknown audio signal contains phoneme m; Step 9.1: Input a continuous audio signal to be detected and repeat the above steps; Step 9.2: Calculate the sum of all feature points in different scale spaces; , (20) represents the sum of the feature points in the t-th scale space, ; Step 9.3: Due to the symmetry of the interpolation wavelet, determine the sum of the eigenvalues of the third scale space Is it less than 0.01? If so, the error range condition is met and the next step is performed. Otherwise, the output cannot determine whether the audio signal contains phoneme m. Step 9.4: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the third space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.5: Judgment Is it less than 0.01? If so, further judgment is required; otherwise, it is determined that the audio signal does not contain phoneme m; Step 9.6: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the fifth scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.7: Judgment If it is less than 0.01, further judgment is required, otherwise it is judged that the audio signal does not contain phoneme m; Step 9.8: Note The non-zero interval of ,The negative semi-axis distribution range of the feature points in the seventh scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.9: Output the judgment result "the audio signal contains phoneme m" or "the audio signal does not contain phoneme m" according to steps 21 to 26.

Claims

1. The English phoneme m recognition method based on spline interpolation wavelet neural network is characterized by: First, audio is input and converted into a discrete signal set for analysis using a sample-and-hold technique. The collected discrete signal set is normalized and thresholded to obtain a standardized audio data set. A sixth-order spline interpolation wavelet is selected, and the mapping coefficients of the interpolation wavelet and the ordinary wavelet are used to construct the criterion function of the neural network. The wavelet is used as the excitation function of the neural network. The mapping coefficients of the interpolation wavelet and the ordinary wavelet are used to construct the training feedback matrix. The wavelet neural network training has global convergence, and the convergence result is the wavelet component of the audio signal, and the output layer weight is the corresponding wavelet coefficient. Based on the output layer weight of the wavelet neural network, the characteristics of the m-tone signal are uniquely determined. Perform a discrete Fourier transform on the wavelet coefficients to obtain the spectrum of the input signal, and identify the feature points and their distribution characteristics as the signal recognition conditions; by comparing the feature points and their distribution characteristics, determine whether the input phoneme contains the consonant m; The specific steps are as follows: Step 1: Sample the audio signal and convert it into a digital signal, and normalize it; Step 1.1: Sample the input audio signal as a set of discrete signals ; Step 1.2: Take a discrete signal set The element with the maximum absolute value is recorded as : (1) Step 1.3: Set the audio elements read Perform maximum absolute value normalization processing and record it as the normalized signal set ; (2) Step 1.4: Normalize the signal set Perform threshold processing and set a custom threshold value as ;like ,but , record the obtained array as the audio data set ,in, represents the kth audio signal after threshold processing, and N represents the number of audio data; Step 2: Construct a wavelet neural network; Set the sixth-order spline wavelet , which is used as the hidden layer neuron of the wavelet neural network. As the hidden layer neuron threshold, As the input layer weight, is the output layer weight; and then the neural network output is obtained ; (3) in Indicates the amount of time, is the output layer weight The elements in , the calculation method of which is given later; Step 3: Setup Represents the wavelet space where the audio signal features need to be extracted, In the process of signal feature extraction, each is set as , let determine the input layer weight ; (4) Step 4: Determine the parameters in the formula and ; Set the recorded audio duration to ,but , (5)、 in As shown in formula (4); Step 5: Construct the training feedback matrix ; Step 5.1: Construct the vector ; (6) in As shown in formula (4), That is, the number of audio values in step 1.

3. , , ; Step 5.2: Construct the matrix ; (7) in and As shown in formula (5), As shown in formula (6), As shown in formula (4); Step 5.3: Construct the feedback matrix ; (8) Step 6: Construct the criteria function ; Step 6.1: Construct a relation of wavelet in frequency domain; (9) in The sixth-order spline wavelet used in this invention The frequency domain expression of the Fourier transform is: Indicates frequency; Step 6.2: Use the inverse discrete Fourier transform to calculate the formula and obtain the coefficients , and construct the vector , ; Step 6.3: Setup Represents the obtained audio data ; (10) Indicated by The column vector formed by Indicates the number of audio data; Step 6.4: Set the output layer weights The matrix expression of is: (11) Here the symbol express The transpose of is the output layer weight The elements in ; Step 6.5: Set the error on the training set ; (12) here As shown in formula (7); Step 6.6: Constructing the Criterion Function ; (13) (14) Step 7: Determine the parameters in the formula through neural network training ; Step 7.1: Convert audio data and feedback matrix Substitute the following training process: (15) here In expression (10) In the The value in the training step iteration, In expression (10) In the The value in the step iteration training; Step 7.2: Substitute into formula (13) to calculate ; Step 7.3: Substitute into formula (14) to calculate ; Step 7.4: If , training stops; Step 7.5: According to formula (11), obtain the value of formula (15) when training stops of The value is the characteristic wavelet coefficient of the m sound; Step 8.1: Retain the characteristic wavelet coefficients in the third, fifth, and seventh wavelet spaces gather , , ,in , , Respectively represent , , The number of elements in Step 8.2: Perform discrete Fourier transform on the obtained wavelet coefficients. (16) (17) (18) The wavelet coefficient set in the t-th scale space is recorded as , the nth wavelet coefficient is , the data length is , , k is less than A positive integer; Step 8.3: The length calculated in step 8.4 is Array ; For arrays respectively The elements in are processed according to formula (19) and recorded as arrays Elements , called feature points: , (19) in, For arrays Middle Number, For arrays Middle Number, For arrays Middle The number here , for Array length; Step 8.4: Take the union of the feature point intervals of the phoneme m samples produced by multiple groups of human voices, and obtain the feature information of the standard English phoneme m in different scale spaces as shown in the following table; Spatial feature points are distributed in the negative semi-axis range The third space (-2.80, -2.72) ∪ (-1.12, -0.91) ∪ (-0.25, 0) The fifth space (-2.30, -1.75) ∪ (-0.81, -0.79) Seventh Space (-2.4,0) Step 9: Detect whether the unknown audio signal contains phoneme m; Step 9.1: Input a continuous audio signal to be detected and repeat the above steps; Step 9.2: Calculate the sum of all feature points in different scale spaces; , (20) represents the sum of the feature points in the t-th scale space, ; Step 9.3: Due to the symmetry of the interpolation wavelet, determine the sum of the eigenvalues of the third scale space Is it less than 0.01? If so, the error range condition is met and the next step is performed. Otherwise, the output cannot determine whether the audio signal contains phoneme m. Step 9.4: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the third space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.5: Judgment Is it less than 0.01? If so, further judgment is required; otherwise, it is determined that the audio signal does not contain phoneme m; Step 9.6: Note The non-zero element interval of , the negative semi-axis distribution range of the feature points in the fifth scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.7: Judgment If it is less than 0.01, further judgment is required, otherwise it is judged that the audio signal does not contain phoneme m; Step 9.8: Note The non-zero interval of ,The negative semi-axis distribution range of the feature points in the seventh scale space For collection ,like yes A subset of , further judgment is needed. If no A subset of , it is determined that the audio signal does not contain phoneme m; Step 9.9: Output the judgment result "the audio signal contains phoneme m" or "the audio signal does not contain phoneme m" according to steps 21 to 26.