In-vehicle space gain speech synthesis method, device, equipment and storage medium

By constructing a spatial function library and concatenating text encoding features in a single-microphone scenario, synthesized speech Mel spectrum features are generated, solving the problem that single-microphone speech synthesis systems cannot create a sense of space and achieving better speech synthesis effects and perceptibility.

CN118737117BActive Publication Date: 2026-02-03MOBILITY ASIA SMART TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310313731.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2026-02-03
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

Current single-microphone speech synthesis systems cannot effectively create a sense of space, affecting the broadcast effect and the subjective experience of the listener.

Method used

By extracting spatial information parameters of the listener's or passenger's head in the vehicle space, a spatial function library is constructed, text encoding features are obtained, and these features are concatenated with spatial features and input into a decoder for decoding to generate synthesized speech Mel spectrum features. Finally, synthesized speech is output through a vocoder, realizing end-to-end parallel spatial information speech synthesis under a single microphone.

Benefits of technology

It improves the intelligibility and perceptibility of speech synthesis, creates a sense of space in single-microphone scenarios, and enhances the effect of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118737117B_ABST
    Figure CN118737117B_ABST
Patent Text Reader

Abstract

The application relates to a vehicle-mounted space gain speech synthesis method, device, equipment and storage medium. The method comprises the following steps: extracting a space information parameter of a listener or passenger head in a vehicle-mounted space, and constructing a space function library; acquiring a text coding feature; inputting the space parameter information, and extracting a space feature through the space function library; splicing the text coding feature, a position coding and the space feature to form a spliced coding feature; inputting the spliced coding feature and a space gain coding feature into a decoder to obtain a synthesized speech mel spectrum feature; and inputting the synthesized speech mel spectrum feature into a vocoder to output the synthesized speech. The method introduces a space parameter through the use of an end-to-end parallel speech synthesis to help the synthesized speech, realizes single-microphone space gain synthesis, and can create a sense of space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and in particular to a vehicle-mounted spatial gain speech synthesis method, apparatus, computer equipment, and storage medium. Background Technology

[0002] Hearing plays a vital role in human life. It allows us to perceive sounds in our surroundings, enabling us to make judgments and decisions. Besides perceiving subjective attributes of sound such as intensity, pitch, and timbre, human hearing can also determine the direction and distance of a sound source. Spatial information about sound is crucial for sound perception.

[0003] With the application of deep learning, speech synthesis technology is developing rapidly. Most current spatial stereo methods use multiple speakers to create multi-channel spatial stereo sound. However, in single-microphone scenarios, most current speech synthesis systems cannot create a sense of space, affecting the broadcast effect and the human's subjective experience. Summary of the Invention

[0004] Based on this, it is necessary to provide an in-vehicle spatial gain speech synthesis method, device, computer equipment, and storage medium to address the above-mentioned technical problems. This method can introduce spatial information into single-microphone speech synthesis, realize end-to-end parallel spatial information speech synthesis under a single microphone, and solve the technical problem that most of the synthesized speech in current single-microphone scenarios cannot create a sense of space, thus affecting the broadcast effect and the subjective experience of the user.

[0005] On the one hand, a vehicle-mounted spatial gain speech synthesis method is provided, the method comprising:

[0006] Extract spatial information parameters of the head of the listener or passenger in the vehicle space and construct a spatial function library;

[0007] Obtain text encoding features;

[0008] Input spatial parameter information and extract spatial features using the spatial function library;

[0009] The text encoding features are concatenated with the positional encoding and the spatial features to form a concatenated encoding feature;

[0010] The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features;

[0011] The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

[0012] Furthermore, the step of obtaining text encoding features includes:

[0013] Input a text sequence and convert the text sequence into a phoneme sequence;

[0014] After performing character embedding processing on the phoneme sequence and inserting positional encoding, text encoding is performed to output the text encoding feature matrix;

[0015] The text encoding feature matrix is ​​concatenated with the spatial features in the spatial function library and then fed into the variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features.

[0016] Furthermore, the step of inputting a text sequence and converting the text sequence into a phoneme sequence includes:

[0017] Input a text sequence;

[0018] The text sequence is processed using text regularization to form a regular expression;

[0019] The regular expression is converted into Chinese characters through phonetic conversion;

[0020] The Chinese characters are converted into phoneme sequences through polyphonic character classification and prosodic prediction.

[0021] Furthermore, the step of performing character embedding processing on the phoneme sequence and inserting positional encoding, followed by text encoding to output the text encoding feature matrix, includes:

[0022] Define a list of characters;

[0023] Convert the character into a single valid code to obtain a vector sequence;

[0024] The vector sequence is learned using one-dimensional convolution to obtain the character encoding;

[0025] The phoneme sequence is embedded using character encoding.

[0026] Insert positional encoding into the phoneme sequence after character embedding processing;

[0027] The input is fed into the encoder for encoding, and the output is the encoded text feature matrix.

[0028] Furthermore, the step of concatenating the text encoding feature matrix output by the encoder with the spatial features in the spatial function library and then feeding it into the variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features includes:

[0029] The text encoding feature matrix is ​​input into the duration predictor to predict duration and the duration prediction result is output.

[0030] The duration prediction result is input into the length adjuster to adjust the duration and obtain the length of the frame to be synthesized. The length of the frame to be synthesized is copied, and the same text phonemes are copied twice.

[0031] The output of the length adjuster is input into the base frequency predictor and the energy predictor respectively, and the prediction result is output. The length of the frame to be synthesized is copied according to the prediction result.

[0032] The output of the length adjuster is input into the predictor for duration prediction processing;

[0033] The text-encoded features are output by concatenating the features of the base frequency predictor, the energy predictor, and the output of the predictor.

[0034] Furthermore, the step of inputting the concatenated coding features and spatial gain coding features into the decoder for decoding to obtain the synthesized speech Mel spectrum features includes:

[0035] The decoder takes spatial information parameters from the spatial features in the spatial function library used for synthesizing speech as input into the decoder. The decoder generates a Mel spectrogram and then splits the Mel spectrogram to obtain the spatial gain spectrum.

[0036] The spatial gain spectrum is converted into gain audio using parameters or a neural vocoder to synthesize spatial gain speech.

[0037] Furthermore, the step of inputting spatial information parameters from the spatial features in the spatial function library used for synthesizing speech into the decoder, and the decoder generating a Mel spectrogram and splitting the Mel spectrogram to obtain the spatial gain spectrum includes:

[0038] The input spatial information parameters are fed into the head-related transfer function;

[0039] The spatial features from the previous step are output through the established spatial function library. The corresponding head-related transfer function coefficients in the head-related transfer function library are then found in conjunction with the spatial information parameters, and convolution filtering is performed. If the input spatial information parameters do not have a corresponding head-related transfer function in the spatial function library, they are obtained by interpolation using two similar head-related transfer functions.

[0040] The output of the previous step is input into the linear layer output spatial feature mapping sequence to obtain spatial extracted features;

[0041] The decoder is input with the initial Mel spectrum features, which are concatenated with the position-encoded features and the spatially extracted features. Multi-head attention is then input for attention calculation.

[0042] After passing through a loss layer, two one-dimensional convolutional layers, and a linear layer, the output predicted Mel spectrum is combined with spatial splicing features to form a Mel spectrogram.

[0043] Spatial gain spectrum is obtained by decomposing the spatial features in the Mel spectrogram.

[0044] Furthermore, before the step of concatenating the text encoding features with the positional encoding and the spatial features to form the concatenated encoding features, the method further includes:

[0045] Training is performed to align text with corresponding audio and spatial features.

[0046] The spatial features extracted from the spatial function library are concatenated with the text encoding features, input into the variational adapter for duration prediction, and the duration prediction result is output.

[0047] The duration prediction result is concatenated with the encoded location feature and then input into the decoder. The spectrum is concatenated with the spatial feature for alignment training to obtain a spatial gain speech synthesis model.

[0048] Furthermore, the step of using a spatial extractor to extract spatial information parameters of the listener's or passenger's head within the vehicle space through a head-related transfer function, and constructing a spatial function library, includes:

[0049] Determine the distribution area of ​​the listener's or passenger's head within the vehicle space;

[0050] Based on the coordinates of the head distribution area of ​​the listener or passenger, a center point is set and a head-related spherical coordinate system is established. Spatial information parameters from a point in the head-related spherical coordinate system to the center point are set, including the elevation angle θ and the horizontal angle. and distance r;

[0051] The head-related transfer function is used to express the overall filtering effect of human structure on sound waves in the frequency domain acoustic transfer function from the sound source to the two ears in a free field condition.

[0052] The head-related transfer function value was measured using real or simulated humans;

[0053] A spatial function library was constructed using experimental measurements, numerical calculations, and head-related transfer function modeling methods.

[0054] Furthermore, the step of measuring the head-related transfer function value using a real person or a simulated person includes:

[0055] Measurements are taken in an anechoic chamber, with real or simulated humans as the test subjects located at the origin of the coordinate system, and loudspeakers arranged on a sphere centered at the origin of the coordinate system.

[0056] By fixing the position of the object under test and changing the relative position between the speaker and the object under test, the parameters of the head-related transfer function in different spatial directions are measured.

[0057] The loudspeaker generates a measurement signal, and the microphones located at both ears pick up the binaural sound pressure signals; the head correlation transfer function value in the frequency domain is calculated according to the formula of the head correlation transfer function.

[0058] On the other hand, an in-vehicle spatial gain speech synthesis device is provided, the device comprising:

[0059] Spatial extractor, used to extract spatial features of the head of a listener or passenger in a vehicle space;

[0060] The text encoding feature acquisition module is used to acquire text encoding features;

[0061] The spatial function library management module is used to construct a spatial function library based on spatial features, and to extract spatial features through the spatial function library when inputting spatial parameter information.

[0062] An encoder is used to concatenate the text encoding features with the positional encoding and the spatial features to form a concatenated encoding feature;

[0063] A decoder is used to decode the spliced ​​coding features and spatial gain coding features to obtain the synthesized speech Mel spectrum features;

[0064] A vocoder is used to convert the Mel spectrum features of the synthesized speech into synthesized speech.

[0065] In another aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0066] Extract spatial information parameters of the head of the listener or passenger in the vehicle space and construct a spatial function library;

[0067] Obtain text encoding features;

[0068] Input spatial parameter information and extract spatial features using the spatial function library;

[0069] The text encoding features are concatenated with the positional encoding and the spatial features to form a concatenated encoding feature;

[0070] The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features;

[0071] The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

[0072] In another aspect, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0073] Extract spatial information parameters of the head of the listener or passenger in the vehicle space and construct a spatial function library;

[0074] Obtain text encoding features;

[0075] Input spatial parameter information and extract spatial features using the spatial function library;

[0076] The text encoding features are concatenated with the positional encoding and the spatial features to form a concatenated encoding feature;

[0077] The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features;

[0078] The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

[0079] The aforementioned vehicle-mounted spatial gain speech synthesis method, device, computer equipment, and storage medium introduce spatial parameters to assist in speech synthesis by using end-to-end parallel speech synthesis, thereby achieving single-microphone spatial gain synthesis. Furthermore, the spatial parameters help improve the synthesized speech effect, enhance intelligibility and perceptibility, and create a sense of space. Attached Figure Description

[0080] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0081] Figure 1 This is a flowchart illustrating an in-vehicle spatial gain speech synthesis method in one embodiment;

[0082] Figure 2 For one embodiment, the corresponding Figure 1 The diagram shows the module structure and processing principle of the vehicle-mounted spatial gain speech synthesis method.

[0083] Figure 3 This is a schematic diagram of the encoder and decoder structure in one embodiment;

[0084] Figure 4 This is a flowchart illustrating the steps of constructing a spatial function library by using a spatial extractor to extract spatial features of the head of a listener or passenger in a vehicle space through a head-related transfer function in one embodiment.

[0085] Figure 5 This is a schematic diagram of the head-related spherical coordinate system in one embodiment;

[0086] Figure 6This is a flowchart illustrating the steps involved in measuring the Head Related Transfer Function (HRTF) value using a real person or a simulated human in one embodiment.

[0087] Figure 7 This is a flowchart illustrating the steps for measuring the Head Related Transfer Function (HRTF) value using a real person or a simulated human in another embodiment.

[0088] Figure 8 This is a flowchart illustrating the steps of inputting a text sequence and converting the text sequence into a phoneme sequence in one embodiment.

[0089] Figure 9 This is a flowchart illustrating the steps of embedding the phoneme sequence into characters and inserting positional encoding in one embodiment, then inputting it into an encoder for text encoding and outputting a text encoding feature matrix.

[0090] Figure 10 This is a flowchart illustrating the steps in one embodiment whereby the text encoding feature matrix output by the encoder is concatenated with the spatial features in the spatial function library and then fed into the variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features.

[0091] Figure 11 This is a schematic diagram of the variational predictor in one embodiment;

[0092] Figure 12 This is a schematic diagram of the structure of a predictor or duration predictor in one embodiment;

[0093] Figure 13 This is a flowchart illustrating the steps in one embodiment where spatial information parameters from the spatial features in the spatial function library used for synthesizing speech are input into the decoder, and the decoder generates a Mel spectrogram and then splits the Mel spectrogram to obtain a spatial gain spectrum.

[0094] Figure 14 This is a structural block diagram of an in-vehicle spatial gain speech synthesis device in one embodiment;

[0095] Figure 15 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0096] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0097] like Figure 1 , Figure 2 , Figure 3As shown, in one embodiment, an in-vehicle spatial gain speech synthesis method is provided, including the following steps S1-S6.

[0098] Step S1: Extract spatial information parameters of the listener's or passenger's head within the vehicle space and construct a spatial function library.

[0099] Specifically, a spatial extractor is used to extract the spatial features of the listener's or passenger's head in the vehicle space through the Head-Related Transfer Function (HRTF), and a spatial function library is constructed.

[0100] like Figure 4 As shown, in this embodiment, the step of using a spatial extractor to extract the spatial features of the listener's or passenger's head in the vehicle space through a head-related transfer function and constructing a spatial function library specifically includes:

[0101] Step S11: Determine the distribution area of ​​the listener's or passenger's head within the vehicle space;

[0102] Step S12: Based on the coordinates of the head distribution area of ​​the listener or passenger, set a center point and establish a head-related spherical coordinate system. The spatial information parameters of a point in the head-related spherical coordinate system from the center point include the elevation angle θ and the horizontal angle. and distance r;

[0103] Step S13: The head-related transfer function (HRTF) is used to express the overall filtering effect of the human body structure on sound waves in the frequency domain acoustic transfer function from the sound source to both ears under free field conditions; the formula for the head-related transfer function is HRTF(r, θ, ...). f, a)=(P(r, θ, f, a)) / (Ps(f)); where P(r, θ, f, a) represents the sound pressure at the tympanic membrane, Ps(f) represents the sound pressure at the sound source, r represents the distance from the sound source to the center of the head, and θ represents the elevation angle from the sound source to the center of the head. denoted as the horizontal angle from the sound source to the center of the head, f is the sound source frequency, and a is the head size;

[0104] Step S14: Measure the head correlation transfer function (HRTF) value using a real person or a human simulator;

[0105] Step S15: Construct a spatial function library through experimental measurement, numerical calculation, and head-related transfer function modeling.

[0106] The spatial extractor uses the head-related transfer function (HRTF) to express the comprehensive filtering effect of the human body structure on sound waves in a free field, representing the frequency domain acoustic transfer function from the sound source to both ears. Specifically, it is the ratio of the sound pressure at the tympanic membrane to the sound pressure at the sound source, HRTF(r, θ). f, a)=(P(r, θ, f,a)) / (P_s(f)).

[0107] like Figure 5 As shown, Figure 5 Let HRTF be the head-related spherical coordinate system, where HRTF represents the distance r from the sound source to the center of the head, elevation angle θ, and horizontal angle. And a function of the sound source frequency f, where 'a' represents a parameter with personalized characteristics, such as head size. The input spatial information parameters for synthesized speech (elevation angle θ, horizontal angle θ) are also included. The data flow in the spatial extractor is: spatial information parameters -> spatial function -> spatial library features -> linear layer -> output spatial features.

[0108] Therefore, the spatial gain is: combining spatial information parameters (elevation angle θ, horizontal angle) The corresponding head-related transfer function coefficients are found in the HRTF library for the input spatial information parameters (and distance r), and then convolution filtering is performed. If no corresponding head-related transfer function is found in the HRTF library, the coefficients are obtained by interpolation using two similar head-related transfer functions.

[0109] like Figure 6 As shown, in this embodiment, the step of measuring the head correlation transfer function (HRTF) value using a real person or a simulated human includes:

[0110] Step S141: The measurement is carried out in an anechoic chamber, using a real person or a simulated person as the object of measurement located at the origin of the coordinate system, and the loudspeaker is arranged on a sphere with the origin of the coordinate system as the center (radius r).

[0111] Step S142: By fixing the position of the object under test, the relative position between the speaker and the object under test is changed to measure the parameters of the head-related transfer function in different spatial directions;

[0112] In step S143, the loudspeaker generates a measurement signal (pseudo-random signal), and the microphones located at both ears pick up the binaural sound pressure signals; the head correlation transfer function (HRTF) value in the frequency domain is calculated according to the formula of the head correlation transfer function.

[0113] like Figure 7 As shown, in another embodiment, the step of measuring the head correlation transfer function (HRTF) value using a real person or a simulated human includes:

[0114] Step S141': The measurement is carried out in an anechoic chamber, using a real person or a simulated person as the object of measurement located at the origin of the coordinate system, and the loudspeaker is arranged on a sphere with the origin of the coordinate system as the center (radius r).

[0115] Step S142': By fixing the position of the object under test and changing the relative position between the speaker and the object under test, the binaural impulse response (HRIR) is obtained by cross-correlation calculation between the binaural sound signal and the original maximum length sequence.

[0116] Step S143': The head correlation transfer function is obtained by Fourier transforming the binaural impulse response (HRIR), wherein the head correlation transfer function is characterized in the time domain by the binaural impulse response, and the binaural impulse response HRIR(r, θ, f, a) and the head-related transfer function HRTF(r, θ, f, a) are Fourier transform pairs:

[0117]

[0118] In this case, a far-field radius of r > 1.2 meters is considered, where the HRTF is approximately independent of r. To measure the HRTF in different spatial directions, the relative position between the speaker and the object being measured needs to be changed, and then repeated measurements are performed. This can be done by fixing the position of the object being measured and using mechanical equipment to change the orientation of the speaker; or by fixing the position of the speaker and moving a swivel chair to change the orientation of the object being measured.

[0119] Step S2, obtaining text encoding features. Specifically, this includes: inputting a text sequence and converting it into a phoneme sequence; performing character embedding processing on the phoneme sequence and inserting positional encoding, then performing text encoding to output a text encoding feature matrix; wherein the text encoding feature matrix is ​​a feature sequence, such as (512, 1,), where 512 is the text feature dimension and 1 is the positional encoding dimension; concatenating the text encoding feature matrix with spatial features from the spatial function library and then inputting it into a variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features.

[0120] Among them, text encoding features refer to the set of correspondences between words and position codes in the text, that is, one word corresponds to one position coordinate. All words are sorted in sequence to form a series of position coordinate information.

[0121] like Figure 8 As shown, in this embodiment, the step of inputting a text sequence and converting the text sequence into a phoneme sequence includes:

[0122] Step S21, input a text sequence;

[0123] Step S22: The text sequence is processed by text normalization to form a regular expression;

[0124] Step S23: The regular expression is converted into Chinese through grapheme-to-phoneme conversion;

[0125] Step S24: The Chinese text is converted into a phoneme sequence through polyphone classification and prosody prediction.

[0126] like Figure 9 As shown, in this embodiment, the step of performing character embedding processing on the phoneme sequence and inserting positional encoding, and then inputting it into the encoder for text encoding to output the text encoding feature matrix includes:

[0127] Step S31: Define a character list; using alphanumeric characters and a list of special characters and phonemes; for example, English characters, numbers, phonemes, and unknown characters, etc.

[0128] Step S32: Convert the character into a one-bit valid code to obtain a vector sequence; for unknown characters and whitespace characters, replace them with all zero vectors. If the character length exceeds the predefined maximum, ignore it.

[0129] Step S33: Use one-dimensional convolution to learn the vector sequence and obtain the character encoding;

[0130] Step S34: The phoneme sequence is processed by character embedding.

[0131] Step S35: Insert positional codes into the phoneme sequence after character embedding processing;

[0132] Step S36: Input the text into the encoder for encoding, and output the encoded text feature matrix, such as (512,), which is a 512-dimensional matrix.

[0133] One-hot encoding, also known as single-hot coding or one-bit valid coding, uses an N-bit state register to encode N states. Each state has its own independent register bit, and at any given time, only one bit is valid.

[0134] Character encoding, also known as character set encoding, encodes characters from a character set into a specific object within a designated set (e.g., bit patterns, sequences of natural numbers, octets, or electrical pulses) to facilitate text storage in computers and transmission over communication networks. Common examples include encoding the Latin alphabet into Morse code and ASCII. ASCII assigns numbers to letters, numbers, and other symbols, representing each integer using 7 bits of binary code. An additional bit is typically used to store the information in a single byte.

[0135] like Figure 2 As shown, Figure 2 For the corresponding Figure 1 The diagram shows the module structure processing principle of the vehicle-mounted spatial gain speech synthesis method.

[0136] The positional encoding is achieved using the sin / cosine function. The phoneme sequence is positionally encoded for each frame according to its time sequence, and the formula for positional encoding is as follows:

[0137] Where PE is the location code, pos is the location index, i is the dimension index of the time series, and d model For model dimensions.

[0138] like Figure 10 As shown, in this embodiment, the step of concatenating the text encoding feature matrix output by the encoder with the spatial features in the spatial function library and then inputting it into the variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features includes:

[0139] Step S41: Input the text encoding feature matrix into the duration predictor to predict duration and output the duration prediction result;

[0140] Step S42: The duration prediction result is input into the length adjuster to adjust the duration and obtain the length of the frame to be synthesized. The length of the frame to be synthesized is copied, and the same text phonemes are copied twice.

[0141] Step S43: Input the output of the length adjuster into the base frequency predictor and the energy predictor respectively, output the prediction result, and copy the length of the frame to be synthesized according to the prediction result;

[0142] Step S44: Input the output of the length adjuster into the predictor for duration prediction processing;

[0143] Step S45: The text-encoded features are output by concatenating the outputs of the base frequency predictor, the energy predictor, and the predictor.

[0144] like Figure 11, Figure 12 As shown, Figure 11 This is a schematic diagram of the variational predictor, which includes a duration predictor, a length adjuster, a fundamental frequency predictor, an energy predictor, and a predictor. Figure 12 This is a schematic diagram of the predictor or duration predictor structure. The duration predictor structure is a stacked structure of 1D convolution, ReLU activation function, layer normalization, dropout layer, 1D convolution, ReLU activation function, layer normalization, dropout layer, and linear layer. The predictor structure is the same as the duration predictor structure.

[0145] The data flow in the variational predictor is: input (256,) -> 1dConv_layers (convolutional layers) (k = 3, 256 -> 256) -> dropout (dropout layer) (0.5) -> 1dConv_layers (convolutional layers) (k = 3, 256 -> 256) -> dropout (dropout layer) (0.5) -> linear (256).

[0146] Step S3: Input spatial parameter information and extract spatial features using the spatial function library.

[0147] Step S4: The text encoding feature is concatenated with the positional encoding and the spatial feature to form a concatenated encoding feature.

[0148] Specifically, the text encoding features are concatenated with the spatial features in the spatial function library, positional encoding is inserted, and the input is fed into the encoder. The output features are concatenated with the spatial features and then fed into the variational adapter for duration prediction. After that, they are concatenated with the spatial features and positional encoding again to form the concatenated encoding features.

[0149] Step S5: Input the splicing coding features and spatial gain coding features into the decoder to obtain the synthesized speech Mel spectrum features.

[0150] Specifically, the spatial information parameters used for synthesizing speech are input into the spatial features in the spatial function library into the decoder. The decoder generates a Mel spectrogram and then splits the Mel spectrogram to obtain the spatial gain spectrum.

[0151] Step S6: Input the synthesized speech Mel spectrum features into the vocoder and output the synthesized speech.

[0152] Specifically, the spatial gain spectrum is converted into gain audio using parameters or a neural vocoder to synthesize spatial gain speech.

[0153] like Figure 13As shown, the step of inputting spatial information parameters from the spatial features in the spatial function library into the decoder for synthesizing speech, and the decoder generating a Mel spectrogram and splitting the Mel spectrogram to obtain a spatial gain spectrum includes:

[0154] Step S61: The input spatial information parameters are fed into the head-related transfer function. The spatial information parameters include the elevation angle θ and the horizontal angle. and distance r;

[0155] Step S62: Output the spatial features from the previous step using the established spatial function library, and combine the spatial information parameters to find the corresponding head-related transfer function coefficients in the head-related transfer function library, and perform convolution filtering; if the input spatial information parameters do not have a corresponding head-related transfer function in the spatial function library, then interpolate using two similar head-related transfer functions to obtain the function.

[0156] Step S63: Input the output result of the previous step into the linear layer output spatial feature mapping sequence to obtain spatial extracted features;

[0157] Step S64: Input the initial Mel spectrum features into the decoder, concatenate them with the position encoding features, concatenate them with the spatial extraction features, and input multi-head attention to perform attention calculation;

[0158] Step S65: After passing through the lost layer, two one-dimensional convolutional layers, and a linear layer, the predicted Mel spectrum is output and spatially spliced ​​features are used to form a Mel spectrogram.

[0159] Step S66: Decompose the spatial features in the Mel spectrogram to obtain the spatial gain spectrum.

[0160] like Figure 3 As shown, Figure 3 This is a schematic diagram of the encoder and decoder structures. Encoder: It primarily processes the input sequence of the RNN, using the state of the last RNN unit as the context C of the final output. Decoder: It takes the encoder's output C as input and a fixed-length vector as conditions to produce the output sequence Y = {y(1), y(2), ..., y(ny)}.

[0161] Combination Figure 1 , Figure 2 The data streams in the encoder and decoder are as follows.

[0162] Encoder: Input (1,)->embedding(256, 256)->multihead attention (head = 2,)->1dConv_layers (k = 9, 256->1024)->1dConv_layers (k = 1, 1024->256)->linear(256->256).

[0163] Decoder: Input Melp (1, 80) -> Input (1,) -> Embedding (256, 256) -> Multihead Attention (head = 2,) -> 1dConv_layers (k = 9, 256 -> 1024)

[0164] ->1dConv_layers (convolutional layers) (k=1, 1024->256)->linear(256->80)->output is Mel features (1, 80).

[0165] In (256, 256), the value 256 represents the dimension, and the same applies to the entire text.

[0166] Training process: Text and corresponding audio are aligned and trained using spatial features. The text is processed by a text front-end (text regularization, word segmentation, character-to-phoneme (g2p) conversion, prosody prediction, polyphonic character prediction, etc.) to output a phoneme sequence, which is then concatenated with positional encodings and fed into the encoder module (layer normalization -> multi-head attention -> droopout -> residual connection -> layer normalization -> 1D convolution -> activation function & dropout -> 1D convolution -> dropout -> residual connection) to output text-encoded features. A spatial function library is built using a spatial extractor for the training speaker. The spatial features are output through the spatial extractor head correlation output function, mapped to spatial output features through a linear layer, and concatenated with the above text-encoded features. This is then input into a variational adapter for duration prediction (1D convolution + ReLU activation function -> layer normalization -> 1D convolution + ReLU activation function -> layer normalization + dropout -> linear layer). The output is concatenated with the encoded positional features and input into the decoder. The spectrum is concatenated with the above spatial features for alignment and training, ultimately resulting in a spatial gain speech synthesis model.

[0167] Therefore, before the step of concatenating the text encoding features with the positional encoding and the spatial features to form the concatenated encoding features, the method further includes the following step:

[0168] Training is performed to align text with corresponding audio and spatial features.

[0169] The spatial features extracted from the spatial function library are concatenated with the text encoding features, input into the variational adapter for duration prediction, and the duration prediction result is output.

[0170] The duration prediction result is concatenated with the encoded location feature and then input into the decoder. The spectrum is concatenated with the spatial feature for alignment training to obtain a spatial gain speech synthesis model.

[0171] Application Process: Input the text to be synthesized, input spatial parameter information, output spatial features, extract spatial features through a spatial function library, normalize and map to spatial output features through linear layers, the text passes through a text front-end (text regularization, word segmentation, character-to-phoneme (g2p) conversion, prosody prediction, polyphonic character prediction, etc.) output phoneme sequence and position encoding concatenate into the encoder module (layer normalization -> multi-head attention -> droopout -> residual connection -> layer normalization -> 1D convolution -> activation function & dropout -> 1D convolution -> dropout -> residual connection) output text encoding features, input into a variational adapter for duration prediction (1D convolution + ReLU activation function -> layer normalization -> 1D convolution + ReLU activation function -> layer normalization + dropout -> linear layer) output concatenate with the encoded position features and the above spatial features, input into the decoder, output spatial gain Melp, and finally synthesize spatial gain speech through a traditional or neural vocoder.

[0172] In the above-mentioned vehicle-mounted spatial gain speech synthesis method, spatial parameters are introduced to help synthesize speech by using end-to-end parallel speech synthesis, thereby achieving single-microphone spatial gain synthesis; moreover, the spatial parameters are used to assist in improving the synthesized speech effect, enhancing intelligibility and perceptibility, and creating a sense of space.

[0173] It should be understood that, although Figures 1-13 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 1-13 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0174] In one embodiment, such as Figure 14As shown, an in-vehicle spatial gain speech synthesis device 10 is provided, including: a spatial extractor 1, a text encoding feature acquisition module 2, a spatial function library management module 3, an encoder 4, a decoder 5, and a vocoder 6. The spatial extractor 1 is used to extract spatial features of the listener's or passenger's head within the in-vehicle space; the text encoding feature acquisition module 2 is used to acquire text encoding features; the spatial function library management module 3 is used to construct a spatial function library based on spatial features, and extract spatial features through the spatial function library when inputting spatial parameter information; the encoder 4 is used to concatenate the text encoding features with positional encoding and the spatial features to form a concatenated encoding feature; the decoder 5 is used to decode the concatenated encoding feature and the spatial gain encoding feature to obtain the synthesized speech Mel spectrum features; and the vocoder 6 is used to convert the synthesized speech Mel spectrum features into synthesized speech.

[0175] In the aforementioned vehicle-mounted spatial gain speech synthesis device, spatial parameters are introduced through end-to-end parallel speech synthesis to help synthesize speech, thereby achieving single-microphone spatial gain synthesis. Moreover, the spatial parameters are used to assist in improving the synthesized speech effect, enhancing intelligibility and perceptibility, and creating a sense of space.

[0176] Specific limitations regarding the in-vehicle spatial gain speech synthesis device can be found in the limitations of the in-vehicle spatial gain speech synthesis method described above, and will not be repeated here. Each module in the aforementioned in-vehicle spatial gain speech synthesis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0177] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 14 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores in-vehicle spatial gain speech synthesis data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an in-vehicle spatial gain speech synthesis method.

[0178] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 15As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an in-vehicle spatial gain speech synthesis method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0179] Those skilled in the art will understand that Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0180] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0181] Extract spatial information parameters of the head of the listener or passenger in the vehicle space and construct a spatial function library;

[0182] Obtain text encoding features;

[0183] Input spatial parameter information and extract spatial features using the spatial function library;

[0184] The text encoding features are concatenated with the positional encoding and the spatial features to form a concatenated encoding feature;

[0185] The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features;

[0186] The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

[0187] For specific limitations on the steps implemented by the processor when executing a computer program, please refer to the limitations on the method of in-vehicle spatial gain speech synthesis mentioned above, which will not be repeated here.

[0188] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0189] Extract spatial information parameters of the head of the listener or passenger in the vehicle space and construct a spatial function library;

[0190] Obtain text encoding features;

[0191] Input spatial parameter information and extract spatial features using the spatial function library;

[0192] The text encoding features are concatenated with the positional encoding and the spatial features to form a concatenated encoding feature;

[0193] The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features;

[0194] The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

[0195] For specific limitations on the implementation steps of a computer program when executed by a processor, please refer to the limitations on the method of in-vehicle spatial gain speech synthesis mentioned above, which will not be repeated here.

[0196] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0197] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0198] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A vehicle-mounted spatial gain speech synthesis method, characterized in that, include: Spatial information parameters of the head of a listener or passenger in the vehicle space are extracted, and a spatial function library is constructed based on the spatial information parameters using a head-related transfer function modeling method. The spatial information parameters include elevation angle, horizontal angle, and distance. Obtain text encoding features; Input spatial parameter information and extract spatial features using the spatial function library; The text encoding features are concatenated with the positional encoding based on phoneme sequences and the spatial features to form a concatenated encoding feature; The splicing coding features and spatial gain coding features are input into the decoder to obtain the synthesized speech Mel spectrum features; The synthesized speech Mel spectrum features are input into a vocoder to output synthesized speech.

2. The in-vehicle spatial gain speech synthesis method according to claim 1, characterized in that, The step of obtaining text encoding features includes: Input a text sequence and convert the text sequence into a phoneme sequence; After performing character embedding processing on the phoneme sequence and inserting positional encoding, text encoding is performed to output the text encoding feature matrix; The text encoding feature matrix is ​​concatenated with the spatial features in the spatial function library and then fed into the variational predictor to predict the duration of the speech to be synthesized before outputting the text encoding features.

3. The in-vehicle spatial gain speech synthesis method according to claim 2, characterized in that, The step of inputting a text sequence and converting the text sequence into a phoneme sequence includes: Input a text sequence; The text sequence is processed using text regularization to form a regular expression; The regular expression is converted into Chinese characters through phonetic conversion; The Chinese characters are converted into phoneme sequences through polyphonic character classification and prosodic prediction.

4. The in-vehicle spatial gain speech synthesis method according to claim 2, characterized in that, The step of performing character embedding processing on the phoneme sequence and inserting positional encoding, followed by text encoding and outputting the text encoding feature matrix, includes: Define a list of characters; Convert the character into a single valid code to obtain a vector sequence; The vector sequence is learned using one-dimensional convolution to obtain the character encoding; The phoneme sequence is embedded using character encoding. Insert positional encoding into the phoneme sequence after character embedding processing; The input is fed into the encoder for encoding, and the output is the encoded text feature matrix.

5. The in-vehicle spatial gain speech synthesis method according to claim 2, characterized in that, The step of concatenating the text-encoded feature matrix with the spatial features in the spatial function library and then feeding it into the variational predictor to predict the duration of the speech to be synthesized before outputting the text-encoded features includes: The text encoding feature matrix is ​​input into the duration predictor to predict duration and the duration prediction result is output. The duration prediction result is input into the length adjuster to adjust the duration and obtain the length of the frame to be synthesized. The length of the frame to be synthesized is copied, and the same text phonemes are copied twice. The output of the length adjuster is input into the base frequency predictor and the energy predictor respectively, and the prediction result is output. The length of the frame to be synthesized is copied according to the prediction result. The output of the length adjuster is input into the predictor for duration prediction processing; The text-encoded features are output by concatenating the features of the base frequency predictor, the energy predictor, and the output of the predictor.

6. The in-vehicle spatial gain speech synthesis method according to claim 1, characterized in that, The step of inputting the concatenated coding features and spatial gain coding features into the decoder to obtain the synthesized speech Mel spectrum features includes: The decoder takes spatial information parameters from the spatial features in the spatial function library used for synthesizing speech as input into the decoder. The decoder generates a Mel spectrogram and then splits the Mel spectrogram to obtain the spatial gain spectrum. The spatial gain spectrum is converted into gain audio using parameters or a neural vocoder to synthesize spatial gain speech.

7. The in-vehicle spatial gain speech synthesis method according to claim 6, characterized in that, The steps of inputting spatial information parameters from the spatial features in the spatial function library used for synthesizing speech into the decoder, and generating a Mel spectrogram and splitting the Mel spectrogram to obtain the spatial gain spectrum include: The input spatial information parameters are fed into the head-related transfer function; The spatial features from the previous step are output through the established spatial function library. The corresponding head-related transfer function coefficients in the head-related transfer function library are then found in conjunction with the spatial information parameters, and convolution filtering is performed. If the input spatial information parameters do not have a corresponding head-related transfer function in the spatial function library, they are obtained by interpolation using two similar head-related transfer functions. The output of the previous step is input into the linear layer output spatial feature mapping sequence to obtain spatial extracted features; The decoder is input with the initial Mel spectrum features, which are concatenated with the position-encoded features and the spatially extracted features. Multi-head attention is then input for attention calculation. After passing through a loss layer, two one-dimensional convolutional layers, and a linear layer, the output predicted Mel spectrum is combined with spatial splicing features to form a Mel spectrogram. Spatial gain spectrum is obtained by decomposing the spatial features in the Mel spectrogram.

8. The in-vehicle spatial gain speech synthesis method according to claim 1, characterized in that, Before the step of concatenating the text encoding features with the positional encoding and the spatial features to form the concatenated encoding features, the method further includes: Training is performed to align text with corresponding audio and spatial features. The spatial features extracted from the spatial function library are concatenated with the text encoding features, input into the variational adapter for duration prediction, and the duration prediction result is output. The duration prediction result is concatenated with the encoded location feature and then input into the decoder. The spectrum is concatenated with the spatial feature for alignment training to obtain a spatial gain speech synthesis model.

9. The in-vehicle spatial gain speech synthesis method according to claim 1, characterized in that, The step of extracting spatial information parameters of the listener's or passenger's head within the vehicle space and constructing a spatial function library based on these spatial information parameters using a head-related transfer function modeling method includes: Determine the distribution area of ​​the listener's or passenger's head within the vehicle space; Based on the coordinates of the distribution area of ​​the listener's or passenger's head, a center point is set and a head-related spherical coordinate system is established. The spatial information parameters from a point in the head-related spherical coordinate system to the center point include the elevation angle θ, the horizontal angle φ, and the distance r. The head-related transfer function is used to express the overall filtering effect of human structure on sound waves in the frequency domain acoustic transfer function from the sound source to the two ears in a free field condition. The head-related transfer function value was measured using real or simulated humans; A spatial function library was constructed using experimental measurements, numerical calculations, and head-related transfer function modeling methods.

10. The vehicle-mounted spatial gain speech synthesis method according to claim 9, characterized in that, The step of measuring the head-related transfer function value using a real person or a simulated human includes: Measurements are taken in an anechoic chamber, with real or simulated humans as the test subjects located at the origin of the coordinate system, and loudspeakers arranged on a sphere centered at the origin of the coordinate system. By fixing the position of the object under test and changing the relative position between the speaker and the object under test, the parameters of the head-related transfer function in different spatial directions are measured. The loudspeaker generates a measurement signal, and the microphones located at both ears pick up the binaural sound pressure signals; the head correlation transfer function value in the frequency domain is calculated according to the formula of the head correlation transfer function.

11. A vehicle-mounted spatial gain speech synthesis device, characterized in that, The device includes: A spatial extractor is used to extract spatial information parameters of the head of a listener or passenger in a vehicle space, including elevation angle, horizontal angle and distance. The text encoding feature acquisition module is used to acquire text encoding features; The spatial function library management module is used to construct a spatial function library based on the spatial information parameters using the head-related transfer function modeling method, and to extract spatial features through the spatial function library when inputting spatial parameter information. An encoder is used to concatenate the text encoding features with positional encoding based on phoneme sequences and the spatial features to form concatenated encoding features; A decoder is used to decode the spliced ​​coding features and spatial gain coding features to obtain the synthesized speech Mel spectrum features; A vocoder is used to convert the Mel spectrum features of the synthesized speech into synthesized speech.

12. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Rhythm control voice synthesis method and system and electronic device

    CN111754976A

  • End-to-end speech synthesis method and device in combination with sound transfer function

    CN112967728A