New speech spectrum information extraction method based on generalized S transformation

By introducing generalized S transform into speech signal processing and adjusting the parameters in the Gaussian window function, the problem of the spectrum resolution cannot be adjusted is solved, efficient and flexible speech signal processing is achieved, and high-resolution spectrum diagrams are obtained.

CN119993191APending Publication Date: 2025-05-13XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311513080.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the time resolution and frequency resolution of the spectrogram cannot be flexibly adjusted, which affects the accuracy of obtaining speech gene information.

Method used

Using a method based on generalized S transform, the time and frequency resolution are adjusted by introducing adjustable parameters into the Gaussian window function to achieve flexible adjustment of spectrogram resolution.

Benefits of technology

It realizes flexible adjustment of spectral resolution, improves the efficiency and accuracy of speech signal processing, and obtains high-resolution spectral patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a novel spectral information extraction method based on generalized S-transform speech. In speech recognition, a speech spectrogram serves as an important spectrogram for displaying speech signal feature information and is widely applied. Nowadays, many spectrogram obtaining methods exist, but the spectrograms with high resolution cannot be flexibly obtained. In order to solve the problem, a generalized S transformation method is introduced to carry out voice signal time-frequency analysis. According to the algorithm, two parameters are introduced into a Gaussian window function on the basis of S transformation, voice signal time-frequency information with different time and frequency resolutions can be extracted by adjusting the adjusting parameters in generalized S transformation, and then spectrograms with different resolutions are obtained. The test result shows that the generalized S transformation is more flexible, and the resolution of the obtained spectrogram is higher. The method is of great significance to the fields of speech recognition, speech synthesis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing, and in particular to a new method for extracting spectral information based on generalized S transform. Background Art

[0002] The low resolution of the spectrogram will affect the accuracy of obtaining speech gene information, so it is necessary to obtain a high-resolution spectrogram. The resolution of the traditional method of obtaining spectrograms cannot be flexibly adjusted. With the introduction of the S transform, the generalized S transform obtained by improving the S transform has been widely used in various scenarios, but the generalized S transform has not yet been applied to speech signal processing.

[0003] S transform is an algorithm based on short-time Fourier transform. Generalized S transform improves S transform by introducing parameters that can adjust time and frequency resolution in the Gaussian window function in S transform, thus realizing the feasibility of adjustable time and frequency resolution in speech signals. Generalized S transform consists of two parts, namely speech signal and improved Gaussian window function. The speech signal is input into the model of generalized S transform, and the spectrogram with higher resolution is gradually adjusted by adjusting the adjustable parameters in Gaussian window function. Since the model contains two adjustment parameters, when adjusting the parameters, one of the parameters is often fixed first, and the other parameter is adjusted to obtain the resolution of the component corresponding to the corresponding parameter, and then the other parameter is adjusted to obtain the required high-resolution spectrogram according to the needs of the experiment.

[0004] Generalized S transform is a relatively mature method. The model of generalized S transform will change accordingly in different application scenarios. The new models obtained by improving S transform are collectively called generalized S transform.

[0005] In addition to the generalized S transform, there are other traditional methods that are applied to speech signals to obtain spectrograms, but the window function is relatively fixed and the resolution of the spectrogram cannot be flexibly adjusted. Some models take a long time to run and are less efficient, such as short-time Fourier transform and wavelet transform.

[0006] The method based on generalized S transform not only overcomes the problem of unadjustable resolution, but also has high efficiency and good resolution. This project is supported by the Natural Science Foundation of Xinjiang Uygur Autonomous Region, project number: 2022D01C61. Summary of the invention

[0007] The present invention provides a new method for extracting spectral information based on generalized S transform, which solves the problem that the time resolution and frequency resolution of the spectrogram cannot be flexibly adjusted in the prior art, as described below: A new method for extracting spectral information based on generalized S transform, the method comprises the following steps: Recorded speech data set, a short speech set recorded in a closed and quiet environment, the speech content is mainly Chinese and numbers; Preprocess the recorded speech signal, such as resampling, framing, windowing, etc.; The preprocessed speech signal is introduced into the generalized S transform model to obtain the most original S transform (i.e., the adjustment parameters in the generalized S transform are all assigned to 1 at this time) speech signal spectrogram; Fix the adjustment parameters of the frequency resolution and adjust the parameters controlling the time resolution, starting from 0.1 and adjusting in steps of 0.1 until the time resolution is optimal (the time resolution is determined by the "vertical line" parallel to the vertical axis and perpendicular to the horizontal axis); Fix the adjusted time resolution parameters, and adjust the frequency resolution parameters starting from 0.1 with a step size of 0.1 until a spectrogram with better frequency resolution is obtained while ensuring the time resolution.

[0008] Among them, the generalized S transform model mainly includes: a signal to be processed, a Gaussian window function, a time resolution parameter, and a frequency resolution parameter.

[0009] The beneficial effects of the technical solution provided by the present invention are: 1. The present invention proposes a method for introducing adjustable parameters based on S-transformation to adjust the time resolution and frequency resolution of the spectrogram and improve the flexibility of obtaining resolution; 2. The present invention changes the ratio of frequency and time in the Gaussian window function by adjusting the size of the parameters, thereby controlling the size of the window function, so that the resolution in the result graph is improved and an effective spectrogram of the speech signal is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flowchart of a method for processing speech signals using a generalized S transform;

[0011] Figure 2 It is the time domain diagram of speech signal;

[0012] Figure 3 Schematic diagram of fixing the frequency parameter and adjusting the time parameter to obtain high time resolution;

[0013] Figure 4 A schematic diagram of fixing the time parameter and adjusting the frequency parameter to obtain high frequency resolution;

[0014] Figure 5 It is the generalized S-transform high-resolution speech signal spectrogram;

[0015] Figure 6 The formula of the generalized S transform proposed by the present invention;

[0016] Figure 7 is the Gaussian window function of the generalized S transform. DETAILED DESCRIPTION

[0015] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.

[0016] Example 1 A new method for extracting spectral information based on generalized S transform, see Figure 1 , the method comprises the following steps: 101: recorded speech dataset, which is recorded in a closed and quiet environment, and the content is mainly Chinese and numbers; 102: Preprocessing the recorded speech signal, the preprocessing including: resampling, framing, windowing, etc.; 103: construct a generalized S transform model, use the Gaussian window function to depend on time and frequency, introduce adjustable parameters into the Gaussian window function, and obtain a generalized S transform. The generalized S transform combines the advantages of short-time Fourier transform and S transform, and can adjust the resolution while obtaining a spectrogram, making it possible to obtain a high-resolution spectrogram; 104: Adjust the parameters in the Gaussian window function, fix one parameter and adjust another parameter, and perform step-by-step adjustment to obtain a high-resolution spectrum;

[0017] Among them, the generalized S transform model mainly includes: Gaussian window function, time adjustment parameter, and frequency adjustment parameter.

[0018] Example 2 The scheme in Example 1 is further introduced below in combination with specific examples and calculation formulas, as described below for details: The model for constructing the generalized S transform is shown in Figure 6 . Where h(t) is the preprocessed speech signal, Figure 7 W(t,f,m,n) is the Gaussian window function. Among them, m is the parameter for adjusting the frequency resolution, which adjusts the frequency resolution in the spectrogram. Within a certain range, the larger the m, the higher the frequency resolution. n is the parameter for adjusting the time resolution, which adjusts the time resolution in the spectrogram. Within a certain range, the larger the n, the higher the time resolution.

[0019] The spectrogram contains pitch information and formant information. The formant information is mainly contained in each harmonic. Whether each harmonic can be easily identified represents the frequency resolution. The pitch period is mainly determined by the time resolution. The time interval between the "vertical lines" parallel to the vertical axis is the pitch period. The ease of distinguishing it in the spectrogram represents the time resolution of the speech signal. The factor affecting the resolution is mainly the size of the window function. The longer the window length, the higher the frequency resolution. Conversely, the lower the window length, the higher the time resolution. "Voiceprint" information can be observed in the spectrogram. The speaker can be determined through the "voiceprint" information, which is very meaningful in the field of speech recognition and speech synthesis.

[0020] When performing a generalized S transform on speech, adjusting the parameters within a certain range will correspondingly change the size of the window function, and thus change the size of the resolution. By adjusting the adjustment parameters in the generalized S transform, the time-frequency information of the speech signal with different time and frequency resolutions can be extracted.

[0021] By adjusting the parameters within the value range, a high-resolution spectrogram can be obtained. Further analysis of the spectrogram can obtain important information such as intonation, fundamental frequency period, fundamental frequency component, resonance peak, etc. in the speech signal, which brings convenience to speech signal processing.

Claims

1. A new method for extracting spectral information based on generalized S transform, characterized in that: Includes the following: Record a speech data set, preprocess the recorded speech data set, resample, add windows, and divide the frames, import the processed data set into the generalized S transform model, and obtain a high-resolution spectrogram by adjusting the parameters therein. Gaussian window function, based on the original Gaussian window function module, introduces parameters. The introduced parameters are mainly used to adjust the size of the window function, and then adjust the size of the time and frequency resolution. This method introduces adjustable parameters in the window function of the S transform to generate a generalized S transform. By adjusting the adjustment parameters in the generalized S transform, the time-frequency information of the speech signal with different time and frequency resolutions can be extracted. The spectrogram is obtained and the spectrogram is processed in the next step, such as pitch detection and resonance peak detection.

2. The method according to claim 1, characterized in that: The target is the speech signal.