Long audio voice copying method and system based on attention mechanism
By adopting a long audio speech imprinting method based on attention mechanism in speech synthesis technology, combining short- and long-term features, eliminating error interference, and encoding audio features through long-term short-term memory networks and attention mechanisms, the problem of lack of natural sense of speech and low timeliness in traditional technology is solved, and a more natural and efficient long-audio speech generation is achieved.
Patent Information
- Application Number
- CN202510265306.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional speech synthesis technology is difficult to effectively handle semantic coherence and emotional expression in long audio, resulting in the generated speech lacking natural sense and expressiveness, and failing to update in time, resulting in low timeliness and errors in long audio speech imprinting.
The long audio speech imprinting method based on the attention mechanism is adopted, and complex audio scene information is characterized by combining short-term features and long-term features. The year-on-month predictor, sequence discrete filter and threshold filter are used to eliminate error interference, and the audio features are encoded through the long-term memory network and attention mechanism to generate long-term audio speech.
Better reflect the overall characteristics and time scale characteristics of long audio content, improve the natural sense and expressiveness of speech, enhance robustness and distinction, improve recognition efficiency and stability, and effectively deal with semantic coherence problems in long audio.
Smart Images

Figure CN120108373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and more specifically to a long audio speech imitation method and system based on an attention mechanism. Background Art
[0002] With the development of artificial intelligence technology, speech synthesis technology plays an increasingly important role in the generation of long audio content. Traditional speech synthesis technology often cannot effectively handle the semantic coherence and emotional expression problems in long audio, resulting in the lack of naturalness and expressiveness of the generated speech.
[0003] At the same time, the speech synthesis results based on long audio content do not take into account the changes and updates on the time scale, which leads to the fact that speech synthesis technology is often unable to effectively avoid the problem of low timeliness, resulting in the generated speech failing to be updated in a timely manner, thus causing errors in the imitation of long audio speech.
[0004] Therefore, how to propose a long audio speech imitation method and system based on the attention mechanism, for long audio content, by combining short-time features and long-time features, to characterize complex audio scene information, and better reflect the overall characteristics and time scale characteristics of long audio content is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0005] In view of this, the present invention provides a long audio speech imitation method and system based on the attention mechanism, which characterizes complex audio scene information by combining short-term features and long-term features, and better reflects the overall characteristics and time scale characteristics of long audio content. In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A long audio speech imitation method based on attention mechanism, comprising:
[0007] Collect long audio data and preprocess the long audio data;
[0008] Construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor;
[0009] Using a sequence discreteness filter and a threshold filter respectively to judge the output value of the year-on-year and month-on-month predictor and the long audio data to eliminate error interference;
[0010] Extract short-time and long-time audio features from the long audio data after eliminating error interference;
[0011] The extracted short-term and long-term audio features are fused, and the fused audio features are encoded through the long short-term memory network and attention mechanism;
[0012] Based on the encoded audio features, a long audio speech is generated through a vocoder.
[0013] Optionally, the preprocessing of the long audio data includes: performing denoising, segmentation and normalization processing on the long audio data respectively.
[0014] Optionally, the constructing of a year-on-year and month-on-month predictor and using the year-on-year and month-on-month predictor to predict the preprocessed data to obtain an output value of the year-on-year and month-on-month predictor includes:
[0015] Use LSTM-Attention to build a time series data prediction model, and use the time series data prediction model to predict the preprocessed long audio data, and use the predicted value as the output value at the next moment;
[0016] Through the data in two directions of the time series data matrix, a year-on-year and month-on-month predictor including a year-on-year predictor and a month-on-month predictor is constructed, and the year-on-year and month-on-month predictor is used to predict the output value at the next moment to obtain the output value of the year-on-year and month-on-month predictor.
[0017] Optionally, using LSTM-Attention to build a time series data prediction model, and using the time series data prediction model to predict the preprocessed long audio data includes:
[0018] A long short-term memory network (LSTM) is used to extract the time series features of the preprocessed long audio data, and the time series features are used as output values of the long short-term memory network (LSTM);
[0019] The attention network Attention is used to calculate the correlation between the output value of the long short-term memory network LSTM and the sentiment parameter;
[0020] A predicted value is obtained according to the correlation, and the predicted value is used as the output value at the next moment.
[0021] Optionally, obtaining the output value of the year-on-year and month-on-month predictor includes:
[0022] Use time series data to construct a time series data matrix;
[0023] The data of each row in the time series data matrix is connected into a sequence, and the LSTM-Attention network for month-on-month comparison is trained to obtain a month-on-month comparison predictor; the data of each column in the time series data matrix is connected into a sequence, and the LSTM-Attention network for year-on-year comparison is trained to obtain a year-on-year comparison predictor;
[0024] The output value at the next moment is predicted using the long audio data updated n days before the current moment and the long audio data updated m hours before the current moment, respectively, to obtain the output value of the year-on-year predictor and the output value of the month-on-month predictor;
[0025] The average of the output value of the year-on-year predictor and the output value of the quarter-on-quarter predictor is used as the output value of the year-on-year and quarter-on-quarter predictor.
[0026] Optionally, obtaining the output value of the year-on-year predictor and the output value of the month-on-month predictor includes:
[0027] Extracting periodic data from preprocessed long audio data;
[0028] The year-on-year predictor is trained using periodic data at the same time on different days, and the month-on-month predictor is trained using pre-processed long audio data at different times of each day;
[0029] The periodic data at the current moment is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor. The long audio data of the first m hours after preprocessing is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor.
[0030] Optionally, the using a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data respectively to eliminate error interference includes:
[0031] According to the output value of the year-on-year and month-on-month forecaster within m hours and the preprocessed long audio data, an error sequence is obtained, and the dispersion of the error sequence is calculated;
[0032] Use the sequence discreteness filter to analyze the discreteness of the error sequence and select the sequence with the largest discreteness;
[0033] The threshold filter is used to filter out the moment with the largest absolute error in the error sequence, and the long audio data at the maximum moment is removed.
[0034] Optionally, the extracting short-term and long-term audio features includes:
[0035] Short-term audio feature extraction includes: short-term audio feature extraction within a short-term window or at the frame level. Short-term audio features include: time domain features, frequency domain features, and cepstrum features.
[0036] Long-term audio feature extraction includes: audio scene Gaussian super vector and audio scene total change factor feature extraction of the entire audio file.
[0037] Optionally, the correlation includes:
[0038]
[0039] Among them, Attention(Q,K,V) represents the relevance, Q, K, V represent the query matrix, the queried matrix and the value matrix respectively, d k represents the dimension of K, K T Represents the transposed matrix of K.
[0040] Optionally, a long audio speech imitation system based on an attention mechanism, comprising:
[0041] Preprocessing module: used to collect and preprocess long audio data;
[0042] Prediction module: used to construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor;
[0043] Error elimination module: used to use a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data respectively, and eliminate error interference;
[0044] Feature extraction module: used to extract short-term and long-term audio features from the long audio data after eliminating error interference;
[0045] Encoding module: used to fuse the extracted short-term and long-term audio features, and encode the fused audio features through the long short-term memory network and attention mechanism;
[0046] Output module: used to generate long audio speech through a vocoder based on the encoded audio features.
[0047] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a long audio speech imitation method and system based on the attention mechanism, which has the following beneficial effects:
[0048] The present invention proposes a long audio speech imitation method based on an attention mechanism, including: collecting long audio data and preprocessing the long audio data; constructing a year-on-year and month-on-month predictor, using the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor; using a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data, and eliminating error interference; extracting short-term and long-term audio features from the long audio data after eliminating error interference; fusing the extracted short-term and long-term audio features, encoding the fused audio features through a long short-term memory network and an attention mechanism; based on the encoded audio features, generating long audio speech through a vocoder. On the basis of short-term feature extraction, the present invention further combines the long-term features of the audio scene, can characterize complex audio scene information, input a classification model and its fusion model, perform classification and recognition, and output an identification label of the audio scene, which has stronger robustness and better discrimination, and can characterize the overall characteristics of the scene data to a greater extent, with high recognition efficiency and strong stability. A year-on-year and month-on-month predictor is constructed, and the year-on-year and month-on-month predictor is used to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor; the output value of the year-on-year and month-on-month predictor and the long audio data are judged by a sequence discreteness filter and a threshold filter respectively, and error interference is eliminated to better reflect the overall characteristics and time scale characteristics of the long audio content. The present invention better simulates the intonation and emotion of human speech through the attention mechanism, and effectively handles the semantic coherence problem in long audio. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0050] Figure 1 A schematic flow chart of a long audio speech imitation method based on an attention mechanism provided by the present invention.
[0051] Figure 2 A structural framework diagram of a long audio speech imitation system based on an attention mechanism provided by the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] The embodiment of the present invention discloses a long audio speech imitation method based on an attention mechanism, such as Figure 1 As shown, including:
[0054] Collect long audio data and preprocess the long audio data;
[0055] Construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor;
[0056] Using a sequence discreteness filter and a threshold filter respectively to judge the output value of the year-on-year and month-on-month predictor and the long audio data to eliminate error interference;
[0057] Extract short-time and long-time audio features from the long audio data after eliminating error interference;
[0058] The extracted short-term and long-term audio features are fused, and the fused audio features are encoded through the long short-term memory network and attention mechanism;
[0059] Based on the encoded audio features, a long audio speech is generated through a vocoder.
[0060] Furthermore, the preprocessing of the long audio data includes: performing denoising, segmentation and normalization processing on the long audio data respectively.
[0061] In a specific embodiment, the pre-processing further comprises:
[0062] The audio format, sampling rate, number of channels, etc. are converted. Common audio signal sampling rates include 8kHz, 16kHz, and 44.1kHz. The signal is then converted into a feature sequence by sampling at equal time intervals. Since the audio signal has short-term stability within 10ms to 50ms, the algorithm needs to frame the input signal (add a short time window). The frame length is generally set to 20ms, with a frame shift of 10ms or a frame length of 40ms and a frame shift of 20ms. Considering the increase in the sampling rate of the audio signal and the complexity of the audio signal content, the frame length and frame shift can be appropriately lengthened. The common pre-emphasis module for speech signals is not applicable here because the speech signal is affected by glottal excitation and oral and nasal radiation, and the high-frequency part is attenuated at 6dB / octave above about 800Hz, while the audio signal does not adapt to this principle.
[0063] Furthermore, the year-on-year and month-on-month predictor is constructed, and the output value of the year-on-year and month-on-month predictor is obtained by predicting the preprocessed data using the year-on-year and month-on-month predictor, including:
[0064] Use LSTM-Attention to build a time series data prediction model, and use the time series data prediction model to predict the preprocessed long audio data, and use the predicted value as the output value at the next moment;
[0065] Through the data in two directions of the time series data matrix, a year-on-year and month-on-month predictor including a year-on-year predictor and a month-on-month predictor is constructed, and the year-on-year and month-on-month predictor is used to predict the output value at the next moment to obtain the output value of the year-on-year and month-on-month predictor.
[0066] Furthermore, the use of LSTM-Attention to construct a time series data prediction model, and the use of the time series data prediction model to predict the preprocessed long audio data includes:
[0067] A long short-term memory network (LSTM) is used to extract the time series features of the preprocessed long audio data, and the time series features are used as output values of the long short-term memory network (LSTM);
[0068] The attention network Attention is used to calculate the correlation between the output value of the long short-term memory network LSTM and the sentiment parameter;
[0069] A predicted value is obtained according to the correlation, and the predicted value is used as the output value at the next moment.
[0070] In a specific implementation, the expression of the output value of the long short-term memory network LSTM is as follows:
[0071] y t =σ(Wh t );
[0072]
[0073] Among them, y t represents the output value of the long short-term memory network LSTM, σ represents the sigmoid activation function, W represents the weight matrix, and h t Represents the tanh activation function and output gate z o The current hidden state obtained, c t represents the cell state output of the memory unit at time t, z f represents the forget gate, z i represents the input gate, and z represents the input representation calculated at time t.
[0074] Furthermore, the output value of the year-on-year and month-on-month forecaster is obtained, including:
[0075] Use time series data to construct a time series data matrix;
[0076] The data of each row in the time series data matrix is connected into a sequence, and the LSTM-Attention network for month-on-month comparison is trained to obtain a month-on-month comparison predictor; the data of each column in the time series data matrix is connected into a sequence, and the LSTM-Attention network for year-on-year comparison is trained to obtain a year-on-year comparison predictor;
[0077] The output value at the next moment is predicted using the long audio data updated n days before the current moment and the long audio data updated m hours before the current moment, respectively, to obtain the output value of the year-on-year predictor and the output value of the month-on-month predictor;
[0078] The average of the output value of the year-on-year predictor and the output value of the quarter-on-quarter predictor is used as the output value of the year-on-year and quarter-on-quarter predictor.
[0079] Further, obtaining the output value of the year-on-year predictor and the output value of the month-on-month predictor includes:
[0080] Extracting periodic data from preprocessed long audio data;
[0081] The year-on-year predictor is trained using periodic data at the same time on different days, and the month-on-month predictor is trained using pre-processed long audio data at different times of each day;
[0082] The periodic data at the current moment is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor. The long audio data of the first m hours after preprocessing is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor.
[0083] Furthermore, the step of using a sequence discreteness filter and a threshold filter to respectively judge the output value of the year-on-year and month-on-month predictor and the long audio data to eliminate error interference includes:
[0084] According to the output value of the year-on-year and month-on-month forecaster within m hours and the preprocessed long audio data, an error sequence is obtained, and the dispersion of the error sequence is calculated;
[0085] Use the sequence discreteness filter to analyze the discreteness of the error sequence and select the sequence with the largest discreteness;
[0086] The threshold filter is used to filter out the moment with the largest absolute error in the error sequence, and the long audio data at the maximum moment is removed.
[0087] Furthermore, the short-term and long-term audio feature extraction includes:
[0088] Short-term audio feature extraction includes: short-term audio feature extraction within a short-term window or at the frame level. Short-term audio features include: time domain features, frequency domain features, and cepstrum features.
[0089] Long-term audio feature extraction includes: audio scene Gaussian super vector and audio scene total change factor feature extraction of the entire audio file.
[0090] In a specific implementation, the extraction of short-term and long-term audio features specifically includes:
[0091] For the preprocessed audio signal, the audio scene recognition system first extracts short-time features within a short-time window or at the frame level. Short-time features include time domain features such as short-time energy, fundamental frequency, zero-crossing rate; frequency domain features such as spectral center of gravity, spectral flux, spectral flatness, spectral entropy; and cepstrum features such as Mel frequency cepstrum coefficients and Gammatone filter group cepstrum coefficients. Assume that there are N audio scene files in the data set, and the short-time feature vector extracted from the nth (n=1, ..., N) audio is denoted by x. n To express.
[0092] The audio scene Gaussian supervector feature extraction includes:
[0093] Use a large amount of audio scene background data to train a background model that is unrelated to the target scene;
[0094] Then, maximum a posteriori estimation is performed for each audio scene, and the background model parameters are updated to obtain GMM models of different target scenes;
[0095] The target scene mean vector is updated to obtain
[0096] The mean vector of the target scene is calculated by using the method of calculating statistics concatenate into a high-dimensional supervector Sn, where Sn is the audio scene Gaussian supervector.
[0097] The audio scene total change factor feature extraction includes:
[0098] Construct the GMM-UBM model and use the expectation maximization algorithm to calculate the model parameters
[0099]
[0100] Extract Gaussian supervectors;
[0101] According to the GMM-UBM, Gaussian supervector and total change factor analysis model assumptions, the total change matrix T is calculated;
[0102] Calculate the total change factor w n expectations; n It is expected that the SI-vector feature vector is obtained by storage, and the SI-vector feature vector is a feature vector of the total change factor of the audio scene.
[0103] Furthermore, the correlation includes:
[0104]
[0105] Among them, Attention(Q,K,V) represents the relevance, Q, K, V represent the query matrix, the queried matrix and the value matrix respectively, d k represents the dimension of K, K T Represents the transposed matrix of K.
[0106] In a specific implementation, the step of fusing the extracted short-term and long-term audio features and encoding the fused audio features through a long short-term memory network and an attention mechanism includes:
[0107] A support vector machine (SVM) classifier is used to train the audio classification task. During the training process, the short-term and long-term feature vectors are concatenated together as input, and an additional weight generation module is trained. This module can be a simple neural network that outputs the weights of the short-term and long-term features based on some preliminary features of the input audio (such as audio duration, approximate frequency range, etc.). During the test phase, the short-term and long-term features are weighted and fused according to the weights output by this weight generation module.
[0108] The extracted audio features are encoded through LSTM to form the initial representation of the features;
[0109] Adopting the self-attention or graph attention mechanism, constructing the association graph between audio feature frames and strengthening the semantic association between audio feature frame nodes;
[0110] The output of LSTM is used as the input of the attention mechanism, and the features are weighted through the attention layer to highlight important features and suppress unimportant features;
[0111] The attention-weighted features are fed into a fully connected layer for final decision or classification and encoding.
[0112] In a specific implementation, generating a long audio speech by a vocoder based on the encoded audio features includes:
[0113] The encoded audio features are first organized into a format suitable for WaveNet input.
[0114] WaveNet is a generative model based on convolutional neural networks (CNNs). It gradually generates sample points of audio waveforms by processing input features. During the generation process, the generation probability of each sample point is affected by the previously generated sample points and input features. Its unique dilated convolution structure can effectively capture long-term dependencies, thereby generating high-quality audio waveforms. For example, in speech synthesis, it can generate natural and fluent speech, and can simulate different speaker styles and intonations.
[0115] Furthermore, the WaveNet model is optimized by optimizing parameters such as the convolution kernel size and the number of layers, or by using techniques such as grouped convolution, in order to improve the efficiency of the vocoder while ensuring audio quality.
[0116] In a specific implementation, a long audio speech imitation system based on an attention mechanism, such as Figure 2 As shown, including:
[0117] Preprocessing module: used to collect and preprocess long audio data;
[0118] Prediction module: used to construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor;
[0119] Error elimination module: used to use a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data respectively, and eliminate error interference;
[0120] Feature extraction module: used to extract short-term and long-term audio features from the long audio data after eliminating error interference;
[0121] Encoding module: used to fuse the extracted short-term and long-term audio features, and encode the fused audio features through the long short-term memory network and attention mechanism;
[0122] Output module: used to generate long audio speech through a vocoder based on the encoded audio features.
[0123] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0124] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A long audio speech imitation method based on attention mechanism, characterized in that: include: Collect long audio data and preprocess the long audio data; Construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor; Using a sequence discreteness filter and a threshold filter respectively to judge the output value of the year-on-year and month-on-month predictor and the long audio data to eliminate error interference; Extract short-time and long-time audio features from the long audio data after eliminating error interference; The extracted short-term and long-term audio features are fused, and the fused audio features are encoded through the long short-term memory network and attention mechanism; Based on the encoded audio features, a long audio speech is generated through a vocoder.
2. According to the long audio speech imitation method based on the attention mechanism of claim 1, it is characterized in that: The preprocessing of the long audio data includes: performing denoising, segmentation and normalization processing on the long audio data respectively.
3. According to the long audio speech imitation method based on the attention mechanism of claim 1, it is characterized in that: The year-on-year and month-on-month predictor is constructed, and the output value of the year-on-year and month-on-month predictor is obtained by predicting the preprocessed data using the year-on-year and month-on-month predictor. The output value of the year-on-year and month-on-month predictor includes: Use LSTM-Attention to build a time series data prediction model, and use the time series data prediction model to predict the preprocessed long audio data, and use the predicted value as the output value at the next moment; Through the data in two directions of the time series data matrix, a year-on-year and month-on-month predictor including a year-on-year predictor and a month-on-month predictor is constructed, and the year-on-year and month-on-month predictor is used to predict the output value at the next moment to obtain the output value of the year-on-year and month-on-month predictor.
4. According to the long audio speech imitation method based on the attention mechanism of claim 3, it is characterized in that: The method of using LSTM-Attention to construct a time series data prediction model and using the time series data prediction model to predict the preprocessed long audio data includes: A long short-term memory network (LSTM) is used to extract the time series features of the preprocessed long audio data, and the time series features are used as output values of the long short-term memory network (LSTM); The attention network Attention is used to calculate the correlation between the output value of the long short-term memory network LSTM and the sentiment parameter; A predicted value is obtained according to the correlation, and the predicted value is used as the output value at the next moment.
5. According to the long audio speech imitation method based on the attention mechanism of claim 3, it is characterized in that: The output value of the year-on-year and month-on-month forecaster is obtained as follows: Use time series data to construct a time series data matrix; The data of each row in the time series data matrix is connected into a sequence, and the LSTM-Attention network for month-on-month comparison is trained to obtain a month-on-month comparison predictor; the data of each column in the time series data matrix is connected into a sequence, and the LSTM-Attention network for year-on-year comparison is trained to obtain a year-on-year comparison predictor; The output value at the next moment is predicted using the long audio data updated n days before the current moment and the long audio data updated m hours before the current moment, respectively, to obtain the output value of the year-on-year predictor and the output value of the month-on-month predictor; The average of the output value of the year-on-year predictor and the output value of the quarter-on-quarter predictor is used as the output value of the year-on-year and quarter-on-quarter predictor.
6. A long audio speech imitation method based on attention mechanism according to claim 5, characterized in that: The output value of the year-on-year predictor and the output value of the month-on-month predictor are obtained by: Extracting periodic data from preprocessed long audio data; The year-on-year predictor is trained using periodic data at the same time on different days, and the month-on-month predictor is trained using pre-processed long audio data at different times of each day; The periodic data at the current moment is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor. The long audio data of the first m hours after preprocessing is used as the input of the trained year-on-year predictor, and the output value at the next moment is predicted to obtain the output of the year-on-year predictor.
7. The long audio speech imitation method based on the attention mechanism according to claim 1, characterized in that: The method of using a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data respectively to eliminate error interference includes: According to the output value of the year-on-year and month-on-month forecaster within m hours and the preprocessed long audio data, an error sequence is obtained, and the dispersion of the error sequence is calculated; Use the sequence discreteness filter to analyze the discreteness of the error sequence and select the sequence with the largest discreteness; The threshold filter is used to filter out the moment with the largest absolute error in the error sequence, and the long audio data at the maximum moment is removed.
8. The long audio speech imitation method based on the attention mechanism according to claim 1, characterized in that: The short-term and long-term audio feature extraction comprises: Short-term audio feature extraction includes: short-term audio feature extraction within a short-term window or at the frame level. Short-term audio features include: time domain features, frequency domain features, and cepstrum features. Long-term audio feature extraction includes: audio scene Gaussian super vector and audio scene total change factor feature extraction of the entire audio file.
9. The long audio speech imitation method based on the attention mechanism according to claim 4, characterized in that: The dependencies include: Among them, Attention(Q,K,V) represents the relevance, Q, K, V represent the query matrix, the queried matrix and the value matrix respectively, d k represents the dimension of K, K T Represents the transposed matrix of K.
10. A long audio speech imitation system based on attention mechanism, characterized in that: include: Preprocessing module: used to collect and preprocess long audio data; Prediction module: used to construct a year-on-year and month-on-month predictor, and use the year-on-year and month-on-month predictor to predict the preprocessed data to obtain the output value of the year-on-year and month-on-month predictor; Error elimination module: used to use a sequence discreteness filter and a threshold filter to judge the output value of the year-on-year and month-on-month predictor and the long audio data respectively, and eliminate error interference; Feature extraction module: used to extract short-term and long-term audio features from the long audio data after eliminating error interference; Encoding module: used to fuse the extracted short-term and long-term audio features, and encode the fused audio features through the long short-term memory network and attention mechanism; Output module: used to generate long audio speech through a vocoder based on the encoded audio features.