An AI voice software data preprocessing system

Through adaptive pre-emphasis and frame-based windowing technology, the transient noise interference problem caused by fixed pre-emphasis coefficients in AI voice software is solved, the text accuracy of speech recognition is improved, and more efficient speech signal processing is achieved.

CN119323969BActive Publication Date: 2025-07-04JILIN JIUBAKU NETWORK TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411436795.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-07-04
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

In the prior art, AI voice software uses a fixed pre-emphasis coefficient during speech recognition, resulting in serious interference in transient noise, affecting the accuracy of text recognition.

Method used

Adaptive pre-emphasis method is adopted to dynamically adjust the pre-emphasis coefficient by calculating the transient recognition value, distinguishing between transient noise and non-transient noise, and using different pre-emphasis formulas to process, combining frame-based and windowing technology to optimize speech signal processing.

Benefits of technology

It improves the text accuracy after speech recognition, balances the pre-emphasis effect and efficiency, effectively suppresses transient noise interference, and improves the accuracy of the recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323969B_ABST
    Figure CN119323969B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of speech processing, and discloses an AI speech software data preprocessing system, including a pre-emphasis module. The pre-emphasis module is used to perform pre-emphasis on the speech signal to be pre-emphasized, and obtain the pre-emphasized speech signal. The pre-emphasis module includes a calculation unit and a pre-emphasis unit. The calculation unit is used to calculate the transient recognition value tranoicoef of the sampling point s(t) at the sampling moment t in the speech signal to be pre-emphasized t ; The pre-emphasis unit is used to perform pre-emphasis processing on s(t) based on the transient recognition value, including: judging whether the transient recognition value is greater than the comparison value of the set transient recognition value; if not, then using a pre-emphasis formula with a fixed pre-emphasis coefficient to perform pre-emphasis processing on s(t); if so, then using an adaptive pre-emphasis coefficient to perform pre-emphasis processing on s(t). The present invention has a good balance between the effect of pre-emphasis and the efficiency of pre-emphasis. When further recognition is performed based on the pre-emphasized speech signal, the accuracy of the recognized text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing, and in particular, to an AI speech software data preprocessing system. Background Art

[0002] With the development of artificial intelligence technology, there are more and more AI speech software such as ChatGPT. These software can provide more diverse and more user-friendly feedback based on the interaction information input by users. When users input interaction information, there are generally two ways: text input and voice input. Text input refers to inputting interaction information through a physical keyboard or a virtual keyboard, while voice input is to obtain the voice of the user speaking and then convert the voice into text.

[0003] During the voice input process, it generally includes steps such as preprocessing the voice signal, feature extraction, and model recognition. In the prior art, in the preprocessing stage, it is usually necessary to perform steps such as pre-emphasis, framing, and windowing on the voice signal. However, the pre-emphasis coefficient generally used is a constant value, which easily leads to the amplification of the interference of transient noise on the pre-emphasized voice signal, resulting in a low accuracy rate of the text recognized based on the voice signal after this pre-emphasis method. Transient noise refers to the noise that appears briefly in time. Summary of the Invention

[0004] The purpose of the present invention is to disclose an AI speech software data preprocessing system to solve the technical problem of how to improve the accuracy rate of the recognized text when the AI speech software obtains the interaction information input by the user through voice recognition.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] The present invention provides an AI speech software data preprocessing system, including a pre-emphasis module, which is used to perform pre-emphasis on the voice signal to be pre-emphasized to obtain the pre-emphasized voice signal;

[0007] The pre-emphasis module includes a calculation unit and a pre-emphasis unit;

[0008] The calculation unit is used to calculate the transient recognition value tranoicoef of the sampling point s(t) at the sampling moment t in the voice signal to be pre-emphasized t ;

[0009] The pre-emphasis unit is used to perform pre-emphasis processing on s(t) based on the transient recognition value, including:

[0010] Judging whether the transient recognition value is greater than the comparison value of the set transient recognition value;

[0011] If not, perform pre-emphasis processing on s(t) using a pre-emphasis formula with a fixed pre-emphasis coefficient;

[0012] If so, perform pre-emphasis processing on s(t) using the following method, including:

[0013] p(t) = sq(t) - δ(t) × sq(t - 1)

[0014] p(t) represents the result of pre-emphasis on s(t), sq(t) and sq(t - 1) respectively represent the amplitudes of the sampling points at sampling times t and t - 1, and δ(t) represents an adaptive pre-emphasis coefficient; ut represents the set of sampling times within the time interval [t - T, t], T represents the set duration, ampl i and ampl t respectively represent the normalized values of the amplitudes of the sampling points at sampling times i and t, and the value of ampl max is 1, and ampl max represents the comparison value of the set normalized value of the amplitude, φ(i) represents the pre-emphasis coefficient sampled when performing pre-emphasis on the sampling point at sampling time i, and β represents a proportional value.

[0015] Preferably, calculate the transient recognition value tranoicoef of the sampling point s(t) at sampling time t in the speech signal to be pre-emphasized t , including:

[0016]

[0017] α1 represents the amplitude weight, α2 represents the persistence weight, and ampl t represents the normalized value of the amplitude of s(t), and ampl pre represents the maximum value of the normalized values of the amplitudes of the sampling points within the time interval [t - T, t + T], N represents the total number of sampling points within the time interval [t - T, t + T], and Nh t represents the total number of sampling points within the time interval [t - T, t + T] whose normalized amplitude values are greater than z times the absolute value of the average amplitude of the speech signal to be pre-emphasized; T represents the set duration.

[0018] Preferably, the set duration is 10 ms.

[0019] Preferably, the pre-emphasis formula with a fixed pre-emphasis coefficient includes:

[0020]

[0021] represents the set pre-emphasis coefficient.

[0022] Preferably, the set pre-emphasis coefficient is 0.91.

[0023] Preferably, the amplitude weight is 0.6 and the persistence weight is 0.4;

[0024] The comparison value of the set transient recognition value is 0.7.

[0025] Preferably, the normalized value of the amplitude is calculated using the following formula:

[0026]

[0027] ampl t represents the normalized value of the amplitude of the sampling point at sampling time t, sq(t) represents the amplitude of the sampling point at sampling time t, sq min and sq max respectively represent the minimum and maximum values of the amplitudes of all sampling points in the speech signal to be pre-emphasized.

[0028] Preferably, it further includes a microphone module;

[0029] The microphone module is used to convert the user's voice into a speech signal to be pre-emphasized when the user needs to input interaction information by voice input.

[0030] Preferably, it further includes a framing module;

[0031] The framing module is used to frame the pre-emphasized speech signal to obtain multiple speech frames.

[0032] Preferably, it further includes a windowing module;

[0033] The windowing module is used to perform windowing processing on the speech frames to obtain multiple speech frames after windowing processing.

[0034] Beneficial effects:

[0035] Different from the existing voice pre-emphasis methods, when the present invention pre-emphasizes the amplitude of each sampling point respectively, it does not use a fixed pre-emphasis coefficient, because the fixed pre-emphasis coefficient cannot well cope with the influence of transient noise. Since pre-emphasis subtracts the amplitude of the previous sampling point from the amplitude of the current sampling point, it is easy to cause the ratio between the amplitude of the transient noise and the amplitude of the sampling point of the speech signal without noise to increase after pre-emphasis compared with before pre-emphasis, thus obtaining relatively larger-amplitude noise. Therefore, the present invention calculates the transient recognition value, so that when the probability that the sampling point is located in transient noise is greater, it is more inclined to use a formula with a non-fixed pre-emphasis coefficient to pre-emphasize the amplitude of the sampling point, and when the probability that the sampling point is located in transient noise is smaller, it is more inclined to use a formula with a fixed pre-emphasis coefficient to pre-emphasize the amplitude of the sampling point, avoiding using a fixed pre-emphasis coefficient for all sampling points for pre-emphasis, and achieving a better balance between the pre-emphasis effect and the pre-emphasis efficiency. It can improve the accuracy of the recognized text when further recognizing based on the pre-emphasized speech signal. Description of the Drawings

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0037] Figure 1 It is a schematic diagram of an AI voice software data preprocessing system of the present invention.

[0038] Figure 2 It is another schematic diagram of an AI voice software data preprocessing system of the present invention. Detailed Embodiments

[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.

[0040] The present invention provides an AI voice software data preprocessing system, including a pre-emphasis module, which is used to pre-emphasize the speech signal to be pre-emphasized to obtain a pre-emphasized speech signal;

[0041] The pre-emphasis module includes a calculation unit and a pre-emphasis unit;

[0042] The calculation unit is used to calculate the transient recognition value tranoicoef of the sampling point s(t) at the sampling moment t in the voice signal to be pre-emphasized. t ;

[0043] The pre-emphasis unit is used to perform pre-emphasis processing on s(t) based on the transient recognition value, including:

[0044] Judging whether the transient recognition value is greater than the comparison value of the set transient recognition value;

[0045] If not, perform pre-emphasis processing on s(t) using a pre-emphasis formula with a fixed pre-emphasis coefficient;

[0046] If so, perform pre-emphasis processing on s(t) using the following method, including:

[0047] p(t) = sq(t) - δ(t) × sq(t - 1)

[0048] p(t) represents the result of pre-emphasizing s(t), sq(t) and sq(t - 1) respectively represent the amplitudes of the sampling points at the sampling moments t and t - 1, and δ(t) represents the adaptive pre-emphasis coefficient; ut represents the set of sampling moments within the time interval [t - T, t], T represents the set duration, ampl i and ampl t respectively represent the normalized values of the amplitudes of the sampling points at the sampling moments i and t, the value of ampl max is 1, ampl max represents the comparison value of the set normalized value of the amplitude, φ(i) represents the pre-emphasis coefficient sampled when pre-emphasizing the sampling point at the sampling moment i, and β represents the proportional value.

[0049] The above solution calculates the transient recognition value, so that when the probability that the sampling point is located in transient noise is greater, it is more inclined to use a formula with a non-fixed pre-emphasis coefficient to pre-emphasize the amplitude of the sampling point, and when the probability that the sampling point is located in transient noise is smaller, it is more inclined to use a formula with a fixed pre-emphasis coefficient to pre-emphasize the amplitude of the sampling point, avoiding using a fixed pre-emphasis coefficient for all sampling points for pre-emphasis, and achieving a good balance between the pre-emphasis effect and the pre-emphasis efficiency. It can improve the accuracy of the recognized text when performing further recognition based on the pre-emphasized voice signal.

[0050] When calculating the adaptive pre-emphasis coefficient, the present invention obtains the pre-emphasis coefficient by the weighted value of the pre-emphasis coefficients of the sampling points within a certain time length before the sampling moment t. Thus, when the time gap between sampling points is larger and the gap of the normalized amplitude is larger, the influence of the pre-emphasis coefficient of the sampled point being referenced on the finally calculated adaptive pre-emphasis coefficient is greater; when the time gap between sampling points is smaller and the gap of the normalized amplitude is smaller, the influence of the pre-emphasis coefficient of the sampled point being referenced on the finally calculated adaptive pre-emphasis coefficient is smaller. Therefore, better suppression of transient noise is achieved in the pre-emphasis stage, which is beneficial to improving the accuracy of the results obtained in the subsequent recognition process.

[0051] Specifically, the time length between the sampling moments t and t - 1 is the length of the sampling time interval.

[0052] Preferably, the proportional value is 0.5.

[0053] Preferably, calculate the transient recognition value tranoicoef of the sampling point s(t) at the sampling moment t in the speech signal to be pre-emphasized t , including:

[0054]

[0055] α1 represents the amplitude weight, α2 represents the persistence weight, ampl t represents the normalized value of the amplitude of s(t), ampl pre represents the maximum value of the normalized amplitudes of the sampling points within the time interval [t - T, t + T] at the sampling moment, N represents the total number of sampling points within the time interval [t - T, t + T] at the sampling moment, Nh t represents the total number of sampling points within the time interval [t - T, t + T] at the sampling moment, among all sampling points, the normalized amplitude of which is greater than z times the absolute value of the average amplitude of the speech signal to be pre-emphasized; T represents the set duration.

[0056] In the present invention, the transient recognition value is calculated based on two types of parameters: the relationship between the normalized amplitude of the sampling points within a given time interval and the normalized amplitudes of other sampling points, and the number of sampling points whose normalized amplitudes meet the requirements within the given time interval. When the gap between the normalized amplitude of the sampling point at the moment t and the maximum value of the normalized amplitudes of the sampling points within the given time interval is smaller, and the number of sampling points whose normalized amplitudes meet the requirements within the given time interval is larger, the transient recognition value is larger. Thus, the probability that the sampling point belongs to the sampling point of transient noise can be more accurately represented from different perspectives.

[0057] Preferably, the value of z is 10.

[0058] Specifically, the amplitude of the transient noise is generally relatively high, and there will be peaks significantly larger than the normal speech signal on the time-domain graph. Therefore, by setting the value of z, it is possible to identify the sampling points that are likely to belong to the transient noise in a given time interval.

[0059] Preferably, the average amplitude of the speech signal to be pre-emphasized is the average of the normalized values of the amplitudes of each sampling point of the speech signal to be pre-emphasized.

[0060] Preferably, the set duration is 10 ms.

[0061] Specifically, the transient noise has the characteristic of short duration. Therefore, by setting a smaller duration to obtain a smaller time interval, it is possible to make the transient recognition value represent the probability of the sampling points belonging to the transient noise with higher accuracy.

[0062] Preferably, the pre-emphasis formula with a fixed pre-emphasis coefficient includes:

[0063]

[0064] Represents the set pre-emphasis coefficient.

[0065] Specifically, for the pre-emphasis coefficient with a fixed value, the same pre-emphasis processing method is used for both transient noise and non-transient noise. Therefore, it cannot effectively suppress the transient noise.

[0066] Preferably, the time interval between t and the first sampling moment is greater than or equal to T.

[0067] Preferably, when the time interval between t and the first sampling moment is less than T, the pre-emphasis formula with a fixed pre-emphasis coefficient is used for pre-emphasizing the sampling points.

[0068] Specifically, when t is small, due to the insufficient number of sampling points for reference, the pre-emphasis processing can be performed by the pre-emphasis formula with a fixed pre-emphasis coefficient. After the value of t becomes larger, the pre-emphasis formula with an adaptive pre-emphasis coefficient is used for pre-emphasis.

[0069] Preferably, the set pre-emphasis coefficient is 0.91.

[0070] Preferably, the amplitude weight is 0.6 and the persistence weight is 0.4;

[0071] The comparison value of the set transient recognition value is 0.7.

[0072] Preferably, the normalized value of the amplitude is calculated using the following formula:

[0073]

[0074] ampl t represents the normalized value of the amplitude of the sampling point at sampling time t, sq(t) represents the amplitude of the sampling point at sampling time t, sq min and sq max respectively represent the minimum and maximum values of the amplitudes of all sampling points in the speech signal to be pre-emphasized.

[0075] Preferably, the normalized value of the amplitude can also be obtained by logarithmic normalization, square root normalization, etc. Correspondingly, the comparison value of the set normalized value of the amplitude also needs to be changed.

[0076] Preferably, as Figure 2 shown, it further includes a microphone module;

[0077] The microphone module is used to convert the user's voice into a speech signal to be pre-emphasized when the user needs to input interaction information in the way of voice input.

[0078] Specifically, the microphone module can be set in a smart phone, or can be set in a notebook, or can also be an independent module outside the computing device.

[0079] Preferably, the user refers to the user of the AI voice software.

[0080] The microphone module converts the sound wave into an electrical signal to obtain an analog signal by collecting the user's voice. After that, after amplifying, filtering and other processing of the analog signal, and then sampling, quantifying and encoding the processed analog signal, the speech signal to be pre-emphasized is obtained.

[0081] Preferably, it further includes a framing module;

[0082] The framing module is used to frame the pre-emphasized speech signal to obtain multiple speech frames.

[0083] Framing (or segmenting) the speech signal is a common signal processing technique, mainly used to divide the continuous speech signal into smaller time windows of a certain length. This technique has many advantages in speech signal analysis, feature extraction, coding and other processing tasks:

[0084] Processing time invariance

[0085] The characteristics of the speech signal change over time. By framing the signal, the signal within each frame can be regarded as relatively stable in time, which makes the processing methods (such as Fourier transform) within each frame more accurate and reliable.

[0086] Improve computational efficiency

[0087] Framing processing converts a large signal into small signal blocks, which can reduce the amount of data processed each time and improve the calculation efficiency. This enables convenient application of algorithms such as frequency-domain analysis (e.g., short-time Fourier transform).

[0088] Reducing non-stationarity

[0089] Speech signals are non-stationary (i.e., their statistical characteristics change over time). By framing, each frame can be regarded as a stationary signal segment, thus simplifying the analysis and processing of the signal. For example, the short-time Fourier transform (STFT) assumes that the signal within each frame is stationary, which makes the spectral analysis more accurate.

[0090] Preferably, the pre-emphasized speech signal is framed to obtain multiple speech frames, including:

[0091] For the k-th speech frame, the calculation formula for the inter-frame distance corresponding to it is:

[0092]

[0093] tist k,k-1 represents the inter-frame distance between the k-th and the (k - 1)-th speech frames, tstd represents the duration of the speech frame, sq max represents the maximum value of the amplitudes of the sampling points in the set hlfu k-1 and sq j represents the amplitude of the sampling point j in hlfu k-1 avesq represents the average value of the amplitudes of the sampling points in the set hlfu k-1 nhlfu represents the total number of sampling points in the set hlfu k-1 Nhk represents the sampling points in the set hlfu k-1 whose transient recognition value is greater than the comparison value of the set transient recognition value during pre-emphasis, hlfu k-1 represents the set of sampling points whose sampling moments are in the time interval midt k-1 represents the sampling moment of the sampling point at the center of the k-th speech frame; η1 represents the fluctuation weight, and η2 represents the transient weight;

[0094] For the k-th speech frame, the time interval in which the sampling moments of the sampling points it contains are located is

[0095] In the present invention, the inter-frame distance is not a fixed value, but is comprehensively calculated based on two aspects: the degree of fluctuation of the amplitudes of the sampling points in the right half of the previous speech frame and the number of sampling points whose transient recognition value is greater than the comparison value of the set transient recognition value during the pre-emphasis stage. Thus, when the amplitude fluctuation is greater and the number of sampling points whose transient recognition value is greater than the comparison value of the set transient recognition value is more, a smaller inter-frame distance is adopted. By utilizing the calculation results of the pre-emphasis stage, the computational complexity can be effectively reduced, and the linkage between the two preprocessing processes is achieved. Therefore, when the amplitude fluctuation of the sampling points is relatively large and the probability of being interfered by transient noise is greater, a smaller inter-frame distance can be calculated, so that the area with greater recognition difficulty can be more substantially repeatedly covered by two adjacent speech frames, which can provide more effective information for the subsequent speech recognition process and improve the accuracy of the result of the text recognized by the subsequent speech recognition.

[0096] Specifically, the value of k is greater than or equal to 2. For the first speech frame, the time interval in which the sampling times of the sampling points it contains are located is [0, tstd].

[0097] Specifically, the sampling point at the center of the k-th speech frame refers to the sampling point whose sampling time is the median of the sampling times corresponding to all the sampling points in the k-th speech frame.

[0098] If the number of all the sampling points in the k speech frames is odd, the median is the number at the middle sampling time; if it is even, the median is the average of the two numbers at the middle sampling times.

[0099] Preferably, the duration of the speech frame is 40 ms.

[0100] Preferably, the fluctuation weight is 0.7 and the transient weight is 0.3.

[0101] Preferably, it further includes a windowing module;

[0102] The frame segmentation module is used to perform windowing processing on the speech frame to obtain a plurality of speech frames subjected to windowing processing.

[0103] Performing windowing processing on the segmented speech signal is a common technique, and the main purpose is to improve the quality of the result when performing signal analysis and processing. The following are several main advantages of windowing the segmented speech signal:

[0104] Reducing spectral leakage

[0105] When performing Fourier transform on a signal, if the start and end of the signal are not smooth, it will cause spectral leakage, making the spectrum of the signal inaccurate. The windowing technique reduces the influence of spectral leakage by introducing smooth transitions at the boundaries of each frame of the signal, thereby obtaining a clearer spectral representation.

[0106] Improve the frequency domain resolution

[0107] Different window functions (such as Hamming window, Hanning window, Blackman window, etc.) have different frequency domain characteristics. By selecting an appropriate window function, better resolution and lower main lobe width can be obtained in the frequency domain, thereby improving the accuracy of the spectrum.

[0108] Improve the signal smoothness

[0109] Windowing can make the start and end of each frame of the signal smoother, thereby improving the signal smoothness of each frame. This helps to reduce the analysis error caused by signal non-stationarity, making subsequent analysis and processing more stable and reliable.

[0110] Preferably, the window functions used by the windowing module include:

[0111] Rectangular window: A simple window function.

[0112] Hamming window: Has good spectral leakage characteristics.

[0113] Hanning window: Can effectively reduce spectral leakage.

[0114] Blackman window: Provides lower spectral leakage and a higher main lobe to side lobe ratio.

[0115] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. An AI voice software data preprocessing system, characterized in that, It includes a pre-emphasis module which is used to perform pre-emphasis on the speech signal to be pre-emphasized and obtain the pre-emphasized speech signal; The pre-emphasis module includes a calculation unit and a pre-emphasis unit; The calculation unit is used to calculate the transient recognition value of the sampling point at the sampling moment t in the voice signal to be weighted ;​ The pre-emphasis unit is used to perform pre-emphasis processing on including: Judge whether the transient recognition value is greater than the comparison value of the set transient recognition value; Otherwise, use a pre-emphasis formula with a fixed pre-emphasis coefficient for to perform pre-emphasis processing; If so, the following method is used to perform pre-emphasis processing, including: Indicates the result of pre-emphasis, and represent the amplitudes of the sampling points at sampling times t and t - 1 respectively, represents the adaptive pre-emphasis coefficient; , ut represents the set of sampling times within the time interval , T represents the set duration, and represent the normalized values of the amplitudes of the sampling points at sampling times i and t respectively, The value of is 1, represents the pre-emphasis coefficient sampled when pre-emphasizing the sampling point at sampling time i, represents the proportional value.

2. The AI voice software data preprocessing system according to claim 1, wherein, Calculate the sample point at sampling time t in the speech signal to be weighted of the transient recognition value , including: represents the amplitude weight, represents the persistence weight, represents the normalized value of the amplitude of represents that the sampling moment is within the time interval the maximum value of the normalized values of the amplitudes of the sampling points within it, N represents the total number of sampling points when the sampling moment is within the time interval within it, represents that the sampling moment is within the time interval among all the sampling points within it, the total number of sampling points whose normalized amplitude values are greater than z times the absolute value of the average amplitude of the speech signal to be weighted; T represents the set duration, and the value of z is 10.

3. An AI voice software data preprocessing system according to claim 1 or 2, characterized in that, The set duration is 10 ms.

4. An AI voice software data preprocessing system according to claim 1, characterized in that, The pre-emphasis formula with a fixed pre-emphasis coefficient includes: Indicates the set pre-emphasis coefficient.

5. An AI voice software data preprocessing system according to claim 4, characterized in that, The set pre-emphasis coefficient is 0.

91.

6. An AI voice software data preprocessing system according to claim 2, characterized in that, The amplitude weight is 0.6 and the persistence weight is 0.4; The comparison value of the set transient recognition value is 0.

7.

7. An AI voice software data preprocessing system according to claim 1 or 2, characterized in that The normalized value of the amplitude is calculated using the following formula: represents the normalized value of the amplitude of the sampling point at sampling time t, represents the amplitude of the sampling point at sampling time t, and respectively represent the minimum and maximum values of the amplitudes of all sampling points in the speech signal to be pre-emphasized.

8. An AI voice software data preprocessing system according to claim 1, characterized in that, It also includes a microphone module; The microphone module is used to convert the user's voice into a speech signal to be pre-emphasized when the user needs to input interaction information in the way of voice input.

9. An AI voice software data preprocessing system according to claim 1, characterized in that, It also includes a framing module; The framing module is used to frame the pre-emphasized speech signal to obtain multiple speech frames.

10. An AI voice software data preprocessing system according to claim 9, characterized in that, It also includes a windowing module; The framing module is used to perform windowing processing on the speech frames to obtain multiple speech frames after windowing processing.

Citation Information

Patent Citations

  • Family breadth installation and maintenance satisfaction degree reasoning method and system based on speech emotion recognition

    CN117612567A

  • Immersive digital doctor assessment method based on AI technology

    CN118522271A