Multi-timbre emotion adaptive speech synthesis method, device, equipment and medium

By performing timbre encoding and emotion rhythm mapping in smart vehicles and optimizing timbre data based on vehicle status and acoustic characteristics, the problem of emotional expression in multiple scenarios in smart cockpits is solved, emotion-adaptive speech synthesis is achieved, and user stickiness and the flexibility of speech synthesis are improved.

CN120726985AActive Publication Date: 2025-09-30SHANGHAI JIDOU TECH CO LTD

Patent Information

Application Number
CN202511212308.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-30
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing speech synthesis technology is difficult to meet the emotional expression needs in multiple scenarios in smart cockpits, resulting in reduced user stickiness and inability to adapt to the emotional expression needs of different scenarios (navigation, warnings, leisure, etc.).

Method used

By obtaining the driving assistance prompt text of the intelligent vehicle in the target driving scenario, performing timbre encoding and emotion rhythm mapping, optimizing the timbre data by combining the vehicle status and acoustic characteristics, generating emotion-adaptive synthetic speech, and realizing multi-timbre emotional expression.

Benefits of technology

Meet users' needs for emotional expression in different scenarios, improve user stickiness, and enhance the flexibility and emotional expression capabilities of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726985A_ABST
    Figure CN120726985A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-tone emotion adaptive speech synthesis method and device, equipment and a medium. The method comprises the following steps: acquiring a driving assistance prompt text; performing timbre coding on the driving auxiliary prompt text to obtain an auxiliary prompt timbre vector, and generating auxiliary prompt timbre data based on the auxiliary prompt timbre vector; performing emotional rhythm control on the auxiliary prompt tone data through an emotional rhythm mapping model and the target driving scene to obtain emotional expression tone data; according to the vehicle driving noise data, optimizing the emotion expression timbre data to obtain noise optimization timbre data; performing acoustic characteristic compensation on the noise optimization tone data, and performing spatial acoustic optimization on the noise optimization tone data to obtain acoustic optimization tone data; and performing context emotion expression adjustment on the acoustic optimization tone data to generate auxiliary prompt synthetic voice of the intelligent vehicle corresponding to the target driving scene. According to the embodiment of the invention, the user viscosity can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of speech synthesis technology, and in particular to a multi-timbre emotion-adaptive speech synthesis method, apparatus, device, and medium. Background Art

[0002] Speech synthesis technology plays a multifaceted role in the cockpit of intelligent vehicles, effectively improving driving convenience, enhancing safety, providing personalized services, and optimizing the human-machine interaction experience. For example, personalized services can be provided based on the driver's voice characteristics and habits; the driver can control the vehicle through simple voice commands, and the system's responses can be fed back to the driver in the form of voice, achieving natural and smooth human-machine interaction, making the human-machine interface more user-friendly, and significantly improving driving convenience and safety.

[0003] In the process of implementing this application, the applicant discovered that the prior art has at least the following problems:

[0004] Existing technologies, which combine numerous pre-recorded voice clips and then splice them together, produce natural sound quality but require a vast voice library, resulting in limited flexibility, difficulty expressing diverse emotions, and difficulty adapting to the diverse scenarios of smart cockpits. This makes it difficult for existing speech synthesis to generate speech with rich emotional expression. In smart cockpit scenarios, this technology cannot meet the emotional demands of diverse scenarios (such as navigation, warnings, and leisure activities), reducing user engagement. Summary of the Invention

[0005] The present application provides a multi-timbre emotion-adaptive speech synthesis method, apparatus, device and medium, which can meet the user's demand for emotional expression in different scenarios, thereby effectively improving user stickiness.

[0006] In a first aspect, the present invention provides a multi-timbre emotion-adaptive speech synthesis method, comprising:

[0007] Obtain the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scenario;

[0008] Performing timbre encoding on the driving assistance prompt text corresponding to the target driving scenario of the intelligent vehicle to obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generating auxiliary prompt timbre data corresponding to the target driving scenario of the intelligent vehicle based on the auxiliary prompt timbre vector;

[0009] Establish an emotional prosody mapping model from emotional state to acoustic parameters, and use the emotional prosody mapping model and the target driving scenario to perform emotional prosody control on the auxiliary prompt timbre data to obtain emotional expression timbre data with emotional prosody;

[0010] Determining vehicle driving noise data based on the current driving state of the intelligent vehicle in a target driving scenario, and optimizing the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data;

[0011] Performing acoustic characteristic compensation on the noise optimization timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and performing spatial acoustic optimization on the noise optimization timbre data based on the acoustic layout characteristics of the intelligent vehicle to obtain acoustic optimization timbre data;

[0012] The acoustically optimized timbre data is adjusted for contextual emotional expression based on the text content of the driving assistance prompt text to generate a synthetic voice for assistive prompts corresponding to the target driving scenario of the intelligent vehicle.

[0013] In a second aspect, the present application also provides a multi-timbre emotion-adaptive speech synthesis device, comprising:

[0014] An acquisition module is used to obtain the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scenario;

[0015] A first processing module is configured to perform timbre encoding on a driving assistance prompt text corresponding to a target driving scenario of the intelligent vehicle, obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generate auxiliary prompt timbre data corresponding to the target driving scenario of the intelligent vehicle based on the auxiliary prompt timbre vector;

[0016] The second processing module is used to establish an emotional prosody mapping model from emotional state to acoustic parameters, and perform emotional prosody control on the auxiliary prompt timbre data based on the emotional prosody mapping model and the target driving scene to obtain emotional expression timbre data with emotional prosody;

[0017] A third processing module is configured to determine vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scenario, and optimize the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data;

[0018] A compensation and optimization module, configured to perform acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and to perform spatial acoustic optimization on the noise-optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle, thereby obtaining acoustically optimized timbre data;

[0019] The adjustment module is used to adjust the acoustically optimized timbre data to perform contextual emotional expression based on the text content of the driving assistance prompt text, so as to generate an auxiliary prompt synthesized speech for the intelligent vehicle corresponding to the target driving scenario.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0021] one or more processors;

[0022] a memory for storing one or more programs,

[0023] When one or more programs are executed by one or more processors, the one or more processors implement the multi-timbre emotion-adaptive speech synthesis method of any embodiment of the present application.

[0024] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the multi-timbre emotion-adaptive speech synthesis method of any embodiment of the present application.

[0025] The embodiment of the present application proposes a multi-timbre emotion-adaptive speech synthesis method, device, equipment and medium, which obtains the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene; performs timbre encoding on the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene, obtains the auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generates the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene based on the auxiliary prompt timbre vector; establishes an emotion rhythm mapping model from emotion state to acoustic parameters, and performs emotion rhythm control on the auxiliary prompt timbre data through the emotion rhythm mapping model and the target driving scene, and obtains an emotion rhythm with Emotional expression timbre data of emotional rhythm; determining vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scenario, and optimizing the emotional expression timbre data based on the vehicle driving noise data to obtain noise optimized timbre data; performing acoustic characteristic compensation on the noise optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and performing spatial acoustic optimization on the noise optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle to obtain acoustic optimized timbre data; adjusting the acoustic optimized timbre data for contextual emotional expression based on the text content of the driving assistance prompt text to generate an auxiliary prompt synthesized speech corresponding to the target driving scenario of the intelligent vehicle. That is, in the technical solution of the present application, it is possible to adapt the synthesized speech to multiple emotional data in different scenarios, and at the same time, it is also possible to adjust the emotional expression of the context in the synthesized speech. In the prior art, by pre-recording a large number of voice clips and splicing them together, it is impossible to bring out the emotional expression of different voice segments. Therefore, compared with the prior art, the multi-timbre emotion-adaptive speech synthesis method, device, equipment and medium proposed in the embodiment of the present application can meet the user's demand for emotional expression in different scenarios, thereby effectively improving user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A flowchart of a multi-timbre emotion-adaptive speech synthesis method provided in one embodiment of the present application;

[0027] Figure 2A flowchart of a multi-timbre emotion-adaptive speech synthesis method provided by another embodiment of the present application;

[0028] Figure 3 A flowchart of a multi-timbre emotion-adaptive speech synthesis method provided in yet another embodiment of the present application;

[0029] Figure 4 A schematic diagram of the structure of a multi-timbre emotion-adaptive speech synthesis device provided in an embodiment of the present application;

[0030] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present application and are not intended to limit the present application. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions of the present application, not all of the structures.

[0032] Figure 1 This is a flow chart of a multi-timbre emotion-adaptive speech synthesis method provided in one embodiment of the present application. This method can be performed by a multi-timbre emotion-adaptive speech synthesis device or electronic device. The device or electronic device can be implemented in software and / or hardware. The device or electronic device can be integrated into any smart device with network communication function. Figure 1 As shown, the multi-timbre emotion-adaptive speech synthesis method may include the following steps:

[0033] S101. Obtain driving assistance prompt text corresponding to the smart vehicle in the target driving scenario.

[0034] In this step, the driver assistance prompt text can be a prompt word / sentence that is voice-announced in the smart cockpit. For example, in a turning scenario, the driver assistance prompt text may be "Turn right in 200 meters ahead"; in a deceleration scenario, the driver assistance prompt text may be "Traffic jam ahead, please slow down."

[0035] S102. Perform tone encoding on the driving assistance prompt text corresponding to the smart vehicle in the target driving scenario to obtain an auxiliary prompt tone vector corresponding to the driving assistance prompt text, and generate auxiliary prompt tone data corresponding to the smart vehicle in the target driving scenario based on the auxiliary prompt tone vector.

[0036] In this step, a unified multi-timbre acoustic space can be established, all timbres share a set of model parameters, and different timbre data can be generated through conditional control.

[0037] For voice vector injection: During the synthesis process, the system injects pre-extracted voice vectors (e.g., 32- or 64-dimensional) into the model as conditional signals. For example, to generate a "female announcer" voice, inject the corresponding voice vector [0.12, -0.34, 0.56...], and the model will generate speech with that voice. For the same text input, "Road construction ahead, please slow down," injecting different voice vectors can generate different voices, such as professional female or mature male. For voice ID embedding, simple ID numbers are used to represent different voices, such as ID=1 for a "standard male voice" and ID=2 for a "standard female voice." The model internally maps these IDs to voice embedding vectors, which serve as conditional control signals. In navigation scenarios, the system can change the speaker characteristics of voice prompts simply by switching the voice ID parameter.

[0038] For parametric control of timbre attributes: the timbre is split into multiple controllable attributes, and different timbres are generated by adjusting these parameters. For scenario-driven automatic timbre selection: the system automatically selects appropriate timbre parameters based on the usage scenario; for example, emergency warnings use highly clear and authoritative timbres; entertainment information uses lively and friendly timbres; and navigation instructions use professional and precise timbres.

[0039] After generating the auxiliary prompt timbre data, the speaker features can be compressed into a very small dimensional vector through efficient timbre encoding, significantly reducing the multi-timbre storage requirements, as shown in the following example.

[0040] Low-dimensional timbre embedding: Compresses traditional 128-256-dimensional timbre vectors to 16-32 dimensions. For example, through autoencoder training, high-dimensional timbre features are compressed into a low-dimensional latent space. The original speaker feature vector [0.18, 0.25, -0.32, ... (128 parameters)] is compressed to [0.41, -0.22, 0.15, ... (16 parameters). Quantized encoding: Discretizes continuous timbre vectors into integer codebook indices. For example, quantizes 32-dimensional floating-point vectors into 8-bit integer vectors [56, 13, 98, 27...], reducing the size of each timbre from 128 bytes (32-dimensional floating-point) to 32 bytes (32-dimensional 8-bit integer). Hierarchical timbre encoding decomposes timbre features into a base timbre and modifiers. For example, a timbre is represented as a base type index (e.g., "adult male" = 3) plus a differential feature vector [0.1, -0.2, 0.05...]. For N similar timbres, only one base type and N small differential vectors need to be stored. Sparse timbre representation uses a small number of non-zero elements to represent the complete timbre feature. For example, a complete timbre is represented by only [0, 0, 0.75, 0, 0, 0.21, 0, 0, 0, -0.48, 0...] (most positions are 0). Only the positions and values ​​of non-zero elements are stored, such as [(2, 0.75), (5, 0.21), (9, -0.48)]. Timbre principal component encoding: This method retains only the main components of timbre characteristics. For example, principal component analysis is used to project voiceprint features onto the 4-8 most discriminative dimensions. The representation is: [0.87, -0.32, 0.15, 0.64] (only the principal component projection values ​​are retained). Combined timbre parameter encoding: This method decomposes timbre into an interpretable combination of acoustic parameters. For example, timbre is represented as [pitch baseline = high, formant distribution = wide, breathiness = medium, timbre quality = bright]. This method has the advantages of having fewer parameters and being interpretable, making it easier to adjust and control intuitively.

[0041] When acquiring an appropriate voice tone, an efficient timbre extraction network can extract and replicate the target timbre characteristics using only minimal audio. The network architecture of this efficient timbre extraction network is based on an improved lightweight ResNet variant as the backbone network; a hybrid architecture combining a self-attention mechanism and depthwise separable convolutions; and a three-stage encoder-pooling-projector structure. The specific network architecture includes: a front-end feature extraction layer consisting of 2-3 1D convolutional layers (kernel size 3-7) to extract basic acoustic features, each followed by batch normalization and LeakyReLU activation. A middle representation layer consisting of 4-6 lightweight residual blocks, each containing depthwise separable convolutions instead of standard convolutions, optionally incorporating a self-attention mechanism to capture long-range speech feature dependencies, and employing skip connections to preserve multi-scale information. A statistical pooling layer performs statistical pooling along the temporal dimension, extracting global mean and standard deviation statistics, and reducing the number of parameters through weight sharing. A projection head consisting of 2-3 fully connected layers progressively reduces the dimensionality to the target timbre vector dimension (typically 16-64 dimensions), employing nonlinear activation functions and regularization techniques. Furthermore, parameter sharing and factorization techniques are used to reduce network parameters. Quantization-aware training is introduced to prepare for subsequent quantization deployment, achieving parameter sparsity and pruning non-critical connections. A meta-learning framework enables rapid adaptation to new timbres. The prototype network design captures core timbre characteristics, and an incremental learning mechanism is employed to rapidly update the timbre representation.

[0042] S103: Establish an emotional prosody mapping model from emotional state to acoustic parameters, and perform emotional prosody control on the auxiliary prompt timbre data through the emotional prosody mapping model and the target driving scene to obtain emotional expression timbre data with emotional prosody.

[0043] In this step, the emotional prosody mapping model can accurately map the emotional state to acoustic parameters (such as speaking rate, pitch, energy, etc.).

[0044] For example, for the basic emotion mapping model, the emotional state is happiness / excitement, and the acoustic parameters are: speaking rate: +15-30% (relative to standard speaking rate); pitch: average increase of 20-40Hz; pitch range: 40-60% expansion; energy: overall enhancement of 2-4dB, with more significant enhancement in high frequencies; and sound quality parameters: improved harmonic-to-noise ratio and increased spectral tilt. The precise mathematical mapping relationship is: F0happy = F0_neutral × (1 + 0.2 × E_intensity), where E_intensity represents the intensity of emotion. For the multidimensional emotion parameter matrix: Emotion intensity dimension mapping: define an intensity coefficient I of 0~1 for each emotion, and the speaking speed adjustment formula is: Speed=Speed_base×(1+Direction_factor×I×Max_change_ratio); for example: when sad emotion (I=0.7), speaking speed=base speaking speed×(1+(-1)×0.7×0.4)=base speaking speed×0.72; emotion mixture mapping: define the emotion mixture vector E=[E_happy,E_sad,E_angry,E_neutral], and the sum of each component is 1. Acoustic parameter calculation: P=∑(E_i×P_i), where P_i is the corresponding parameter of each pure emotion; for example: when E=[0.3,0.1,0.6,0], F0=0.3×F0happy+0.1×F0_sad+0.6×F0_angry.

[0045] S104. Determine vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scenario, and optimize the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data.

[0046] In this step, by establishing a noise-clarity mapping model, the synthesized speech characteristics are dynamically optimized according to the real-time acoustic state in the car to optimize the emotional expression timbre data.

[0047] Real-time noise perception analysis: The car's microphone array is used to collect acoustic environment data in real time, and lightweight spectrum analysis is applied to extract noise feature vectors N=[n1,n2,...,n k ], features include key indicators such as noise spectrum distribution, signal-to-noise ratio, acoustic reverberation, etc., and sliding window analysis is used to capture the dynamic changes of noise. Clarity impact prediction mechanism: Construct an acoustic masking prediction function M(f,N) to estimate the degree of speech masking in each frequency band, and use a lightweight neural network to establish a nonlinear mapping: N→impact map I. The prediction model outputs the speech feature dimension that is most susceptible to the current noise environment, and adopts a decision tree + linear regression hybrid architecture to ensure real-time computing capabilities. Adaptive parameter generation: Design a compensation parameter generation module C(I) to output an acoustic characteristic optimization parameter vector P. The parameter vector contains: frequency gain G=[g1,g2,...,gn ], time domain modulation D=[d1,d2,...,d m ], establish a joint optimization function for the semantic importance weight W and the noise masking degree M, and calculate the optimization parameter: P = argmax (clarity (S, N, P) - α · complexity (P)). Dynamic optimization strategy execution: Real-time adjustment of acoustic parameters: S' = Transform (S, P), selective enhancement of key frequency bands: Apply stronger compensation to frequency bands severely affected by noise, time domain structure optimization: Adjust the time domain envelope to ensure that key phonemes are clearly distinguishable, energy distribution reshaping: Reallocate energy from areas not affected by noise to affected areas. Closed-loop feedback optimization: Simulate auditory perception to evaluate the clarity of synthesized speech under current noise conditions, and evaluate the optimization effect through a lightweight auditory model A: S_clarity = A (S', N). Apply online learning to update mapping model parameters to adapt to newly emerging noise types, establish an optimization history cache, and accelerate parameter generation in similar noise environments. Intelligent scene adaptation: Senses vehicle status (speed, window opening / closing, air conditioning fan, etc.) to predict noise changes, pre-calculates parameters corresponding to possible noise scenarios, achieves ultra-low latency response, automatically adjusts optimization strategy focus based on driving mode and road condition information, learns personalized clarity preferences based on user feedback, and fine-tunes optimization parameters.

[0048] S105. Perform acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and perform spatial acoustic optimization on the noise-optimized timbre data according to the acoustic layout characteristics of the intelligent vehicle to obtain acoustically optimized timbre data.

[0049] In this step, a parameterized vehicle acoustic model is constructed to automatically compensate for the sound loss caused by the acoustic characteristics of different vehicle models. At the same time, the spatial performance of speech synthesis is optimized by combining the characteristics of the vehicle's interior acoustic layout, improving directivity and clarity.

[0050] Optionally, acoustic characteristic compensation is performed on the noise optimization timbre data based on the vehicle attribute characteristics of the intelligent vehicle, including:

[0051] Acoustic characteristic vectors corresponding to vehicle attribute characteristics of the intelligent vehicle are obtained, where the acoustic characteristic vectors are used to describe multiple acoustic characteristic dimensions, including spatial volume, reflectivity, and attractive material distribution. A mapping conversion model from acoustic characteristics to sound distortion is established, and the sound distortion corresponding to the acoustic characteristic vector is determined through the mapping conversion model. Acoustic characteristic compensation function is used to perform acoustic characteristic compensation on the noise-optimized timbre data based on the sound distortion corresponding to the acoustic characteristic vector.

[0052] In this step, the vehicle modeling framework is used to abstract the acoustic characteristics of different vehicle models into parameter vectors. Each parameter represents the dimension of the acoustic characteristics: spatial volume, reflectivity, distribution of sound-absorbing materials, etc., and establishes an acoustic characteristic to sound distortion mapping function D(A) to describe the impact of a specific vehicle model on speech. Multi-level acoustic representation structure: physical layer: vehicle cabin geometry, material properties, speaker position, etc., frequency response layer: frequency response curve, reverberation time, directionality index, etc., perception layer: clarity index, spatial perception, sound image positioning, etc. Compensation mechanism design: inverse filter design: H -1 (f) = 1 / D(A, f) pre-compensates for acoustic characteristics. Adaptive gain control: Independent gain adjustment is applied to different frequency bands. Dynamic range processing: The compression ratio and threshold are adjusted based on the vehicle model's acoustic characteristics. Vehicle characteristic acquisition methods include: Professional measurement: Acoustic measurement data is collected using test signals for various vehicle models. Parameter extraction: Parameters are extracted from the measurement data using a system identification algorithm. Model normalization: Standardizes the parameter space to enable cross-vehicle comparison. A real-time adaptive algorithm: Detects the current vehicle ID, loads corresponding acoustic parameters from a database, adjusts the synthesized speech spectrum characteristics using the compensation function C(f, A), monitors acoustic feedback, and fine-tunes the compensation parameters for closed-loop optimization. Lightweight implementation strategies include: Parameter compression: Compressing the complete acoustic model to less than 100 parameters. Computational optimization: Approximate algorithms are used instead of precise acoustic simulations. Table lookup acceleration: Common compensation values ​​are pre-calculated and retrieved in real time through interpolation. For example, luxury sedans: Compensate for material absorption and adjust mid-frequency gain; compact SUVs: Enhance low-frequency clarity and compensate for spatial reverberation; and convertibles: Dynamically adjust acoustic parameters to accommodate open / closed roof conditions.

[0053] Optionally, spatial acoustic optimization is performed on the noise optimization timbre data according to the acoustic layout characteristics of the intelligent vehicle to obtain acoustic optimization timbre data, including:

[0054] Based on the pre-compensation filter, the clarity of the timbre data at the driving position in the intelligent vehicle in the noise-optimized timbre data is enhanced to perform timbre-directional optimization at the driving position in the intelligent vehicle; the clarity of the timbre data at the non-driving position in the intelligent vehicle in the noise-optimized timbre data is reduced to perform timbre-directional optimization at the non-driving position in the intelligent vehicle.

[0055] In this step, the driver's position is optimized for directional sound: the transfer function of the driver's position relative to the in-car speakers is analyzed, a pre-compensation filter is applied to enhance the clarity of sound received at the driver's position, and a frequency-dependent phase delay is designed to create constructive interference of sound waves at the driver's position. A specific example: for the common "A-pillar + dashboard" speaker layout, -3dB / +6dB cross-compensation is applied; the A-pillar speaker enhances the 2-4kHz frequency band (+3dB) and suppresses low frequencies (<300Hz); the dashboard speaker enhances the 4-8kHz frequency band (+4dB) to maintain mid-frequency balance; an intelligent delay difference of approximately 0.2-0.3ms is introduced between the two sound sources to create a focused sound image at the driver's position. This can improve the speech intelligibility index (SII) at the driver's position by 25% without affecting the listening experience at the passenger position.

[0056] Multi-Zone Sound Personalization: Builds acoustic models for multiple listening zones within the vehicle, designs a spatially selective enhancement algorithm, implements "acoustic zoning," and applies beamforming technology to create a directional sound field. Specific examples include: Navigation command scenarios: primarily enhances sound clarity from the driver's seat; Entertainment content scenarios: optimizes the overall vehicle sound field distribution; during navigation from the driver's seat, the front row sound source balance is adjusted to 65:35 (left:right), creating a sound image biased toward the driver; voice commands and music playback use different spatial optimization presets. Acoustic Masking Avoidance: Identifies acoustically masked areas within the vehicle (such as seat backs and roof curved reflection points), rebalances the multi-speaker output, avoids the propagation path within these masked areas, and dynamically adjusts EQ parameters to compensate for masked frequency bands. Specific example: Recognizing the strong absorption of high frequencies (>3kHz) in SUV rear seats, pre-compensation enhances the 3-5kHz band (+4.5dB) during rear-seat voice playback, leveraging roof reflections to enhance the rear sound field, and increasing energy in the 100-300Hz frequency band in the front speaker output. At the same time, a correlation model between vehicle speed and interior noise spectrum was established, and a speed-adaptive spatial enhancement algorithm was designed to dynamically adjust the sound field characteristics with vehicle speed. Specific examples include: Low speed (<40 km / h): Standard sound field balance, with a slight improvement in high-frequency clarity; Medium speed (40-80 km / h): Wind noise compensation begins, enhancing the 1-2 kHz frequency band and shifting the sound image forward by 10%; High speed (>80 km / h): Enhanced compensation, enhancing the 2-4 kHz frequency band (+6dB), compressing the dynamic range by 30%, and shifting the sound image forward by 25%.

[0057] S106. Adjust the acoustically optimized timbre data for contextual emotional expression based on the text content of the driving assistance prompt text to generate a synthetic voice for assistive prompts corresponding to the target driving scenario for the intelligent vehicle.

[0058] In this step, emotional expression patterns are automatically adjusted based on text semantics and context to enhance emotional naturalness. For example, in the text semantic understanding layer, a lightweight BERT variant extracts text semantic vectors, keyword / phrase detection identifies emotional triggers, lightweight dependency parsing identifies sentence structure, and adjusts basic emotional patterns based on sentence types (declarative, interrogative, command, exclamation), and a pre-trained multi-category classifier maps text into an emotional category space. A context-aware mechanism maintains a sliding window to record recent interaction history and adjusts current emotional expression based on this historical information. Vehicle status information (speed, location, surrounding environment) is integrated to adjust the intensity of emotional expression based on the degree of danger in the scene. User response patterns to voice prompts are analyzed, and emotional expression is adjusted based on user reactions.

[0059] Context-sensitive parameter adjustments include semantic importance weighting and syntactic structure adaptation. For semantic importance weighting, the influence of emotional parameters is adjusted based on the text's semantic importance, W (0-1). For key information passages, Energy = Energy_neutral × (1 + W × E_factor). Non-critical passages use a weaker emotional mapping to maintain clarity. For syntactic structure adaptation, at the beginning of a sentence, the emotional parameter is gradually increased, with the mapping coefficient increasing from 0.3 to 1.0. At the end of a sentence, the F0 contour is adjusted based on the sentence type (decreasing for declarative sentences and increasing for interrogative sentences). At pauses, emotional state influences pause duration; for example, sadness prolongs pauses by 20-50%.

[0060] In addition, emotion-timbre interactive mapping can also be performed, such as timbre characteristic compensation: female timbre: the F0 change caused by emotion is 15~25% larger than that of male timbre; children's timbre: the emotional energy change is relatively small, but the pitch change is more significant, mapping formula: ΔF0=Base_ΔF0×Voice_type_factor; brand timbre consistency maintenance: define the range of emotional parameter changes allowed for brand timbre, and apply a nonlinear compression function when it exceeds the range: P=max_range×tanh(P / max_range).

[0061] Continuous Emotion Intensity Control: This feature enables continuous adjustment of emotion intensity to meet the needs of emotional expression in different scenarios. For example, when emotion intensity α = 0.3, F0 = 0.7 × F0neutral + 0.3 × F0_happy. When excitement intensity α = 0.8, the speech rate is increased by sigmoid(2 × (0.8 - 0.5)) ≈ 0.73 times the maximum change. Setting A = 0.8 (high arousal), V = 0.5 (positive valence), and D = 0.3 (moderate dominance) generates a "mildly excited" emotion. Warning messages gradually increase in intensity, with α(t) increasing linearly from 0.2 to 0.8. Navigation scenario emotion intensity adjustment: Normal navigation prompts: excitement α = 0.2 (mildly lively); approaching the destination: excitement α = 0.5 (moderately happy); route replanning: anxiety α = 0.3 (slightly hurried). When the risk of collision increases, the warning voice emotion intensity smoothly transitions from 0.3 to 0.9.

[0062] It is also possible to design a decoupling mechanism for the representation of emotion and timbre, achieving independent control of the two, decoupling emotion from timbre, and accurately preserving the speaker's identity characteristics when changing the emotional state to avoid identity information leakage. For example, orthogonal representation space design: Independent feature subspace: Decompose the acoustic representation space into mutually orthogonal timbre subspace and emotion subspace; Dimensional independent allocation: Clearly divide the dimensions of the representation space, specify different dimensions to encode different attributes, and enforce feature separation through structured design. Adversarial learning decoupling mechanism: Identity-preserving adversarial training: Introduce the timbre discriminator Disc_id and the emotion discriminator Disc_emo, Disc_id attempts to extract timbre information from the emotion representation; Cross-attribute adversarial constraint: Simultaneously train the emotion discriminator Disc_emo so that it cannot predict the emotion category from the timbre representation Sid, forming a two-way adversarial constraint to ensure the mutual independence of the two representations. Information Bottleneck Compression Mechanism: Representation Channel Constraints: Information bottlenecks are imposed on the timbre and emotion encoders to limit the amount of information passing through. The timbre encoder, Enc_id, is designed to retain only the minimum sufficient statistics related to speaker identity, while the emotion encoder, Enc_emo, retains only the minimum information required to distinguish emotions. Mutual Information Minimization: The mutual information between timbre and emotion representations is explicitly minimized, approximated and optimized through neural estimators and adversarial training. Consistency Constraints: The timbre representations of the same speaker under different emotion conditions should be highly similar, and the emotion representations of the same emotion under different speaker conditions should be highly similar.

[0063] The multi-timbre emotion-adaptive speech synthesis method proposed in the embodiment of the present application obtains the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene; performs timbre encoding on the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene to obtain the auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generates the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene based on the auxiliary prompt timbre vector; establishes an emotion rhythm mapping model from emotion state to acoustic parameters, and performs emotion rhythm control on the auxiliary prompt timbre data through the emotion rhythm mapping model and the target driving scene to obtain an emotion rhythm-based speech synthesis method. Emotional expression timbre data; determining vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scenario, and optimizing the emotional expression timbre data based on the vehicle driving noise data to obtain noise optimized timbre data; performing acoustic characteristic compensation on the noise optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and performing spatial acoustic optimization on the noise optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle to obtain acoustic optimized timbre data; adjusting the acoustic optimized timbre data for contextual emotional expression based on the text content of the driving assistance prompt text to generate an auxiliary prompt synthesized voice corresponding to the target driving scenario of the intelligent vehicle. That is, in the technical solution of the present application, a variety of emotional data can be adapted to the synthesized voice in different scenarios, and at the same time, the emotional expression of the context in the synthesized voice can also be adjusted. In the prior art, by pre-recording a large number of voice clips and splicing them together, it is impossible to bring out the emotional expression of different voice segments. Therefore, compared with the prior art, the multi-timbre emotional adaptive voice synthesis method, device, equipment and medium proposed in the embodiment of the present application can meet the user's demand for emotional expression in different scenarios, thereby effectively improving user stickiness.

[0064] Figure 2 This is a flow chart of a multi-timbre emotion-adaptive speech synthesis method provided by another embodiment of the present application. Based on the above technical solution, it is further optimized and expanded, and can be combined with the above optional implementation methods. Figure 2 As shown, the multi-timbre emotion-adaptive speech synthesis method may include the following steps:

[0065] S201: Input the auxiliary prompt timbre vector into a preset network model, and determine initial timbre data based on the output of the preset network model.

[0066] In this step, a preset network model such as a small-parameter, efficient model is used. The small-parameter, efficient model architecture can solve the problem of computing resource limitations in the smart cockpit environment in speech synthesis while ensuring the quality of speech synthesis.

[0067] The hierarchical lightweight design of the small-parameter and efficient model specifically includes: hybrid neural architecture design: integrating the improved non-autoregressive structure and efficient encoder, introducing depthwise separable convolution instead of standard convolution, and significantly reducing the number of model parameters; non-autoregressive structure and efficient encoder can improve synthesis speed, reduce the number of parameters, maintain synthesis quality, enhance parallel computing capabilities, improve long sentence stability, etc.; parameter sharing multi-task encoder: adopting the key representation layer sharing strategy to greatly reduce the feature extraction parameters and improve the cross-task representation capability; the key representation layer sharing strategy is a technical method adopted in the parameter sharing multi-task encoder, which specifically includes: feature extraction layer sharing, selective sharing mechanism, hierarchical representation mapping, task-specific branch design, and cross-task knowledge transfer; non-autoregressive decoding framework: abandoning autoregressive dependence, realizing high-speed parallel decoding, and designing an adaptive length control mechanism to solve the rhythm prediction problem; the adaptive length control mechanism mainly solves the rhythm prediction problem in non-autoregressive speech synthesis. Efficient acoustic parameter conversion: Design a parameter-efficient acoustic feature generator to achieve high-fidelity audio reconstruction with minimal parameters; the role of acoustic feature generators in speech synthesis systems: audio reconstruction core, high-fidelity restoration, feature compression and expression, cross-representation mapping of sound color and emotion, computing resource optimization, cross-representation mapping, etc.

[0068] A hierarchical distillation framework is designed to efficiently transfer large model knowledge to small parameter models in a multi-stage manner. Dynamic sparse training technology based on parameter importance retains key connections while significantly reducing redundant calculations. A mixed precision quantization strategy adopts a differentiated quantization bit width allocation method and implements different quantization strategies for different layers to further compress the model size.

[0069] S202: Split the initial timbre data into a plurality of controllable timbre attributes, and configure attribute parameters for the plurality of controllable timbre attributes to obtain target timbre data with different timbre attributes.

[0070] In this step, the timbre is broken down into controllable timbre attributes (such as pitch, roughness, and breathiness). These parameters are adjusted to create different timbres. For example, setting pitch high, breathiness strong, and roughness low creates a "young, feminine" timbre; setting pitch low, breathiness weak, and roughness high creates a "calm, masculine" timbre.

[0071] S203. Limit the target timbre data by timbre condition parameters based on the scenario category to which the target driving scenario belongs, and perform timbre mixing on the target timbre data after the timbre condition parameter limitation to obtain the auxiliary prompt timbre data corresponding to the smart vehicle in the target driving scenario.

[0072] In this step, the spectral feature adaptive normalization method of small sample audio design is used to reduce the noise interference caused by differences in environmental and recording conditions, so as to limit the timbre condition parameters of the target timbre data.

[0073] For hybrid timbre generation: a hybrid timbre is generated by weighted combination of multiple basic timbre vectors; for example, the timbre vectors of "professional male voice" and "gentle female voice" are mixed in a ratio of 7:3 to create a new timbre that is authoritative yet approachable.

[0074] The multi-timbre emotion-adaptive speech synthesis method proposed in the embodiment of the present application inputs the auxiliary prompt timbre vector into a preset network model, determines the initial timbre data based on the output of the preset network model, splits the initial timbre data into multiple controllable timbre attributes, and configures the attribute parameters of the multiple controllable timbre attributes to obtain target timbre data with different timbre attributes, limits the timbre condition parameters of the target timbre data based on the scene category to which the target driving scene belongs, and performs timbre mixing on the target timbre data after the timbre condition parameter limitation to obtain the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene, thereby obtaining the auxiliary prompt timbre data adapted to the target driving scene.

[0075] Figure 3 This is a flow chart of a multi-timbre emotion-adaptive speech synthesis method provided in another embodiment of the present application. It is further optimized and expanded based on the above technical solution and can be combined with the above optional implementation methods. Figure 3 As shown, the multi-timbre emotion-adaptive speech synthesis method may include the following steps:

[0076] S301: Detect the voice data capacity corresponding to the auxiliary prompt synthesized voice.

[0077] S302: If the voice data capacity corresponding to the auxiliary prompt synthesized voice exceeds a preset capacity threshold, the auxiliary prompt synthesized voice is voice-segmented based on the driving assistance prompt text to obtain multiple auxiliary prompt voice segments.

[0078] In this step, if the auxiliary prompt synthesized voice whose voice data capacity exceeds the preset capacity threshold is "There is a complex roundabout ahead, please take the third exit to enter Qingshan Road", it is determined that the voice content complexity of the auxiliary prompt synthesized voice is relatively high, and it is split into three auxiliary prompt voice segments: "There is a complex roundabout ahead" + "Please take the third exit" + "Turn into Qingshan Road".

[0079] S303. Generate a speech response strength corresponding to each auxiliary prompt speech segment based on the speech context logic between multiple auxiliary prompt speech segments, and adjust the response strength of the auxiliary prompt synthesized speech based on the speech response strength corresponding to each auxiliary prompt speech segment.

[0080] In this step, in combination with the above, the number "third" is specially emphasized by voice intensity enhancement, and a turn signal sound is added to assist, while key information such as "third" and "green mountain" is repeated as appropriate.

[0081] The multi-timbre emotion-adaptive speech synthesis method proposed in the embodiment of the present application detects the speech data capacity corresponding to the auxiliary prompt synthesized speech and segments the complex speech data to emphasize the key information, making the complex navigation speech easier to understand and reducing the cognitive burden on the driver.

[0082] Furthermore, a mechanism for assessing the semantic importance of text can be established, employing optimized expression strategies for safety-related information. For example, a multi-dimensional semantic analysis framework and hierarchical semantic parsing can be used to construct a multi-level parsing structure encompassing vocabulary, phrases, and sentences. Key information identification algorithms utilize lightweight Transformer variants to extract core content elements. Information entropy assessment can be used to calculate the information content and irreplaceability of text fragments. Importance scoring models and multi-feature vectorization can be used to map semantic units to an importance feature space. A classifier cascade architecture can be used to combine rule models and neural network models for two-stage evaluation. Dynamic context-dependent weight adjustment can be used to adjust the importance of local semantic units based on the complete context. Scenario-based importance mapping and safety keyword identification can be used to establish a special vocabulary and priority level for driving safety. Navigation key information extraction can be used to design specialized evaluation rules for key navigation elements such as direction, distance, and landmarks. Functional expression recognition can be used to identify important functional expressions representing system operations and control instructions. Semantic importance quantification mechanism, multi-level importance division: divide text content into four levels: key, important, ordinary, and auxiliary; probability distribution representation: assign importance probability distribution to each semantic unit instead of a single score; differentiated attenuation model: different types of important information use different contextual attenuation models.

[0083] Optionally, also include:

[0084] An auxiliary prompt synthesized voice corresponding to a target driving scenario is played in the intelligent vehicle; if the warning prompt level of the auxiliary prompt synthesized voice is higher than the preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesized voice, the scene key voice in the auxiliary prompt synthesized voice is extracted; the voice intensity of the scene key voice in the auxiliary prompt synthesized voice is enhanced, and the voice intensity of the auxiliary prompt synthesized voice is adjusted based on the scene key voice after the voice intensity enhancement; the auxiliary prompt synthesized voice is continuously played in the intelligent vehicle.

[0085] In this step, the timbre feature enhancement mechanism includes data layer enhancement, feature layer enhancement, model layer enhancement, and inference layer enhancement. In the data layer enhancement, random mask enhancement: randomly masking some areas in the time-frequency domain forces the network to learn more robust timbre representations rather than relying on potentially unstable local features; synthetic sample augmentation: generating additional synthetic samples based on limited samples to expand the training data, such as creating variants by adding controllable noise or speed changes to the original samples. In the feature layer enhancement, spatial constraint regularization: introducing distance constraints in the timbre embedding space to ensure that the representations of different samples of the same speaker are closely clustered; contrastive noise filtering: designing a contrastive learning framework to identify and filter out feature changes that are not related to timbre identity, leaving stable identity information; timbre feature decomposition: decomposing timbre features into content-independent and content-related parts, focusing on extracting and enhancing the content-independent part. In the model layer enhancements, multi-perspective representation fusion is implemented: timbre features are extracted from different perspectives (time, frequency, and cepstral domains), and multi-perspective information is integrated through an attention mechanism. Hierarchical feature aggregation utilizes both shallow and deep network features, capturing fundamental timbre features at the shallow layer and abstract identity features at the deep layer. Dynamic confidence weighting assigns dynamic confidence weights to different feature dimensions, prioritizing high-confidence features during aggregation. In the inference layer enhancements, integrated inference strategies are implemented: features are extracted from multiple segments within a small sample, and more stable representations are obtained through weighted averaging or median filtering. Uncertainty modeling establishes an uncertainty model for timbre representation, dynamically adjusting feature influences based on the degree of uncertainty during synthesis. A self-calibration mechanism introduces a reference timbre library to automatically calibrate the extracted timbre features to reduce bias.

[0086] Therefore, by enhancing the voice intensity of the scene key voice in the auxiliary prompt synthesized voice, and adjusting the voice intensity of the auxiliary prompt synthesized voice based on the scene key voice after voice intensity enhancement; continuously playing the auxiliary prompt synthesized voice in the smart vehicle to continuously prompt the driver can effectively avoid dangerous driving.

[0087] Optionally, also include:

[0088] The core speech segments of the auxiliary prompt synthesized speech corresponding to the target driving scene of the intelligent vehicle are extracted to obtain multiple core speech segments in the auxiliary prompt synthesized speech, and the speech is reorganized according to the contextual relationship between the multiple core speech segments to obtain the auxiliary prompt core speech; the auxiliary prompt core speech is backed up to perform voice broadcast when the intelligent vehicle moves to the next target driving scene.

[0089] In this step, the synthesized voice is played in the same scene to save voice synthesis resources, or a basic voice backup solution is set for the core functions to ensure basic voice services even if problems occur in the main system.

[0090] Core functions include safety warnings, critical navigation instructions, system status notifications, and emergency communications. Safety warnings include collision warnings ("Vehicle ahead is approaching, please brake urgently"); system fault warnings ("Brake system abnormality, please stop safely"); vehicle status warnings ("Tire pressure is too low, please check as soon as possible"); and emergency avoidance warnings ("Obstacle ahead, please steer immediately"). Critical navigation instructions include upcoming turn warnings ("Turn right at the intersection ahead"); highway exit warnings ("Exit the highway in 500 meters"); lane change instructions ("Please change to the left lane"); and route replanning notifications ("Replanning route"). System status notifications include autonomous driving mode switching ("Exiting autonomous driving mode"); key system activation / deactivation ("Emergency braking system activated"); and critical fuel / battery low warnings ("Battery less than 10%, please charge as soon as possible"). Emergency communications include emergency call status ("Calling emergency assistance"); safety assistance confirmation ("Do you need to contact rescue services?"); and automatic accident reporting ("The accident location has been automatically reported").

[0091] In addition, an acoustic feature enhancement module is designed to target the acoustic characteristics of the smart cockpit, selectively enhancing information in key frequency bands. For example, a mapping of the human ear's sensitivity to different frequencies is established, applying higher gain to the key frequency band for speech recognition (1kHz-4kHz), and dynamically adjusting the weight distribution of different frequency bands. Typical in-vehicle noise spectrum characteristics are predicted, speech features orthogonal to the noise bands are targeted for enhancement, and spectrum valley filling and peak protection strategies are implemented. Key areas of speech content are identified in real time, with enhanced processing of information-rich areas such as consonant boundaries. An attention mechanism is applied to highlight key elements of speech recognition. Contrast enhancement is also implemented, enhancing vowel-consonant and voiced-unvoiced contrasts. Prosodic variations are enhanced to ensure clear emotional expression.

[0092] By introducing a specialized adversarial training framework, the clarity of synthesized speech in noisy in-car environments is improved. A noise-adaptive self-attention layer is introduced in the generator, dynamically adjusting attention weights based on noise conditions to selectively enhance speech features relevant to the current noise characteristics. A multi-scale spectral discriminator simultaneously evaluates speech quality at multiple time-frequency resolutions, designing specific loss weights to emphasize regions overlapping with in-car noise bands and applying a perceptually weighted spectral loss to align with the masking characteristics of the human ear. Vehicle-specific noise modeling establishes a noise characteristic database for different vehicle models, using vehicle-specific noise conditions in training to achieve targeted optimization of synthesized speech for the acoustic environment of that specific vehicle model.

[0093] By developing an emotional transition and blending mechanism, a smooth transition from calm to alert is achieved, with gradual control of emotional parameters: pitch, speed, and timbre are adjusted sentence by sentence; context-driven emotional fusion: emotional intensity is automatically adjusted based on vehicle speed, and the emotional tone of the voice is adjusted according to weather conditions; multimodal emotional synergy: voice emotion changes in coordination with interior lighting, and volume and timbre are adjusted in conjunction with driving mode. Furthermore, the system can adapt to user preferences, such as learning user preferences for timbre and emotional expression through interaction history; automatically adjusting timbre and emotional parameters based on user preferences to provide a personalized experience; and supporting cloning the driver's voice as the system timbre, creating a unique and personalized experience.

[0094] Figure 4 This is a schematic diagram of the structure of the multi-timbre emotion-adaptive speech synthesis device provided in the embodiment of the present application. Figure 4 As shown, the multi-timbre emotion-adaptive speech synthesis device includes:

[0095] An acquisition module 401 is used to acquire a driving assistance prompt text corresponding to the intelligent vehicle in a target driving scenario;

[0096] The first processing module 402 is configured to perform timbre encoding on the driving assistance prompt text corresponding to the target driving scenario of the intelligent vehicle, obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generate auxiliary prompt timbre data corresponding to the target driving scenario of the intelligent vehicle based on the auxiliary prompt timbre vector;

[0097] The second processing module 403 is used to establish an emotional prosody mapping model from emotional state to acoustic parameters, and perform emotional prosody control on the auxiliary prompt timbre data based on the emotional prosody mapping model and the target driving scene to obtain emotional expression timbre data with emotional prosody;

[0098] The third processing module 404 is configured to determine vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scenario, and optimize the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data;

[0099] The compensation and optimization module 405 is used to perform acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and to perform spatial acoustic optimization on the noise-optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle to obtain acoustically optimized timbre data;

[0100] The adjustment module 406 is used to adjust the acoustic optimization timbre data based on the text content of the driving assistance prompt text to perform contextual emotional expression to generate an auxiliary prompt synthesized speech for the intelligent vehicle corresponding to the target driving scenario.

[0101] Optionally, the first processing module 402 is specifically configured to:

[0102] The auxiliary prompt timbre vector is input into a preset network model, and the initial timbre data is determined based on the output of the preset network model; the initial timbre data is split into multiple controllable timbre attributes, and the attribute parameters of the multiple controllable timbre attributes are configured to obtain target timbre data with different timbre attributes; the timbre condition parameters of the target timbre data are limited based on the scene category to which the target driving scene belongs, and the timbre of the target timbre data after the timbre condition parameters are limited is mixed to obtain the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene.

[0103] Optionally, the compensation and optimization module 405 is specifically configured to:

[0104] Acoustic characteristic vectors corresponding to vehicle attribute characteristics of the intelligent vehicle are obtained, where the acoustic characteristic vectors are used to describe multiple acoustic characteristic dimensions, including spatial volume, reflectivity, and attractive material distribution. A mapping conversion model from acoustic characteristics to sound distortion is established, and the sound distortion corresponding to the acoustic characteristic vector is determined through the mapping conversion model. Acoustic characteristic compensation function is used to perform acoustic characteristic compensation on the noise-optimized timbre data based on the sound distortion corresponding to the acoustic characteristic vector.

[0105] Optionally, the compensation and optimization module 405 is specifically configured to:

[0106] Based on the pre-compensation filter, the clarity of the timbre data at the driving position in the intelligent vehicle in the noise-optimized timbre data is enhanced to perform timbre-directional optimization at the driving position in the intelligent vehicle; the clarity of the timbre data at the non-driving position in the intelligent vehicle in the noise-optimized timbre data is reduced to perform timbre-directional optimization at the non-driving position in the intelligent vehicle.

[0107] Optionally, it also includes: a fourth processing module.

[0108] The fourth processing module is used to detect the voice data capacity corresponding to the auxiliary prompt synthesized voice; if the voice data capacity corresponding to the auxiliary prompt synthesized voice exceeds a preset capacity threshold, the auxiliary prompt synthesized voice is voice-splitting based on the driving assistance prompt text to obtain multiple auxiliary prompt voice segments; based on the voice context logic between the multiple auxiliary prompt voice segments, the voice response strength corresponding to each auxiliary prompt voice segment is generated, and the response strength of the auxiliary prompt synthesized voice is adjusted based on the voice response strength corresponding to each auxiliary prompt voice segment.

[0109] Optionally, it also includes: a fifth processing module.

[0110] The fifth processing module is used to play the auxiliary prompt synthesized voice corresponding to the target driving scenario in the intelligent vehicle; if the warning prompt level of the auxiliary prompt synthesized voice is higher than the preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesized voice, then the scene key voice in the auxiliary prompt synthesized voice is extracted; the scene key voice in the auxiliary prompt synthesized voice is enhanced in voice intensity, and the voice intensity of the auxiliary prompt synthesized voice is adjusted based on the scene key voice after voice intensity enhancement; the auxiliary prompt synthesized voice is continuously played in the intelligent vehicle.

[0111] Optionally, it also includes: a sixth processing module.

[0112] The sixth processing module is used to extract the core speech segments of the auxiliary prompt synthesized speech corresponding to the target driving scene of the intelligent vehicle, obtain multiple core speech segments in the auxiliary prompt synthesized speech, and reorganize the speech according to the contextual relationship between the multiple core speech segments to obtain the auxiliary prompt core speech; back up the auxiliary prompt core speech so as to perform voice broadcast when the intelligent vehicle moves to the next target driving scene.

[0113] The above-mentioned multi-timbre emotion-adaptive speech synthesis device can execute the method provided in any embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the multi-timbre emotion-adaptive speech synthesis method provided in any embodiment of the present application.

[0114] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application is shown. Figure 5 The electronic device 12 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0115] like Figure 5 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).

[0116] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0117] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0118] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 5 Not shown, usually called a "hard drive"). Although Figure 5 Although not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), as well as an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present application.

[0119] A program / utility 40 having a set (at least one) of program modules may be stored, for example, in memory 28. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. The program modules generally implement the functions and / or methods of the embodiments described herein.

[0120] The electronic device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the electronic device 12 via the bus 18. It should be understood that although Figure 5 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0121] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the multi-timbre emotion-adaptive speech synthesis method provided in the embodiment of the present application.

[0122] An embodiment of the present application also provides a computer storage medium.

[0123] The computer-readable storage medium of the embodiments of the present application may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.

[0124] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0125] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0126] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0127] The embodiment of the present application also provides a computer program product.

[0128] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer program products, which can include one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0129] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the appended claims.

Claims

1. A multi-timbre emotion-adaptive speech synthesis method, characterized in that: The method comprises: Obtain the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scenario; Performing timbre encoding on the driving assistance prompt text corresponding to the smart vehicle in the target driving scenario to obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generating auxiliary prompt timbre data corresponding to the smart vehicle in the target driving scenario based on the auxiliary prompt timbre vector; Establishing an emotional prosody mapping model from emotional state to acoustic parameters, and performing emotional prosody control on the auxiliary prompt timbre data using the emotional prosody mapping model and the target driving scene to obtain emotional expression timbre data with emotional prosody; Determining vehicle driving noise data based on a current driving state of the intelligent vehicle in the target driving scenario, and optimizing the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data; Performing acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute characteristics of the smart vehicle, and performing spatial acoustic optimization on the noise-optimized timbre data according to the acoustic layout characteristics of the smart vehicle to obtain acoustically optimized timbre data; The acoustically optimized timbre data is adjusted for contextual emotional expression based on the text content of the driving assistance prompt text to generate an auxiliary prompt synthesized voice for the smart vehicle corresponding to the target driving scenario.

2. The method according to claim 1, characterized in that The generating, based on the auxiliary prompt tone color vector, the auxiliary prompt tone color data corresponding to the intelligent vehicle in the target driving scenario includes: Inputting the auxiliary prompt timbre vector into a preset network model, and determining initial timbre data based on an output of the preset network model; Splitting the initial timbre data into a plurality of controllable timbre attributes, and configuring attribute parameters for the plurality of controllable timbre attributes to obtain target timbre data with different timbre attributes; The target timbre data is limited by timbre condition parameters based on the scene category to which the target driving scene belongs, and the target timbre data after being limited by the timbre condition parameters is timbre mixed to obtain the auxiliary prompt timbre data corresponding to the target driving scene of the intelligent vehicle.

3. The method according to claim 1, characterized in that The performing acoustic characteristic compensation on the noise optimization timbre data based on the vehicle attribute characteristics of the intelligent vehicle includes: Obtaining an acoustic characteristic vector corresponding to a vehicle attribute feature of the intelligent vehicle, wherein the acoustic characteristic vector is used to describe multiple acoustic characteristic dimensions, the multiple acoustic characteristic dimensions including: spatial volume, reflectivity, and attractive material distribution; Establishing a mapping conversion model from acoustic characteristics to sound distortion, and determining the sound distortion corresponding to the acoustic characteristic vector through the mapping conversion model; The noise-optimized timbre data is subjected to acoustic characteristic compensation based on the sound distortion corresponding to the acoustic characteristic vector through a characteristic compensation function.

4. The method according to claim 3, characterized in that The performing spatial acoustic optimization on the noise-optimized timbre data according to the acoustic layout characteristics of the intelligent vehicle to obtain the acoustically optimized timbre data includes: Based on the pre-compensation filter, clarity enhancement is performed on the timbre data at the driving position in the smart vehicle in the noise-optimized timbre data, so as to perform timbre-directional optimization at the driving position in the smart vehicle; The clarity of the timbre data at the non-driving position in the smart vehicle in the noise-optimized timbre data is reduced to perform timbre-directional optimization at the non-driving position in the smart vehicle.

5. The method according to claim 1, wherein The method further comprises: Detecting the voice data capacity corresponding to the auxiliary prompt synthesized voice; If the voice data capacity corresponding to the auxiliary prompt synthesized voice exceeds a preset capacity threshold, performing voice segmentation on the auxiliary prompt synthesized voice based on the driving assistance prompt text to obtain a plurality of auxiliary prompt voice segments; Based on the voice context logic between the multiple auxiliary prompt voice segments, the voice response strength corresponding to each auxiliary prompt voice segment is generated, and the response strength of the auxiliary prompt synthesized voice is adjusted based on the voice response strength corresponding to each auxiliary prompt voice segment.

6. The method according to claim 1, characterized in that The method further comprises: Playing an auxiliary prompt synthesized voice corresponding to the target driving scenario in the smart vehicle; If the warning prompt level of the auxiliary prompt synthesized voice is higher than the preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesized voice, extracting the scene key voice in the auxiliary prompt synthesized voice; Performing voice intensity enhancement on the scene key voice in the auxiliary prompt synthesized voice, and adjusting the voice intensity of the auxiliary prompt synthesized voice based on the scene key voice after the voice intensity enhancement; The auxiliary prompt synthesized voice is continuously played in the smart vehicle.

7. The method according to claim 1, characterized in that The method further comprises: Extracting core speech segments from the auxiliary prompt synthesized speech of the intelligent vehicle corresponding to the target driving scenario to obtain multiple core speech segments in the auxiliary prompt synthesized speech, and reorganizing speech based on contextual relationships between the multiple core speech segments to obtain an auxiliary prompt core speech; The auxiliary prompt core voice is backed up so as to perform voice broadcast when the smart vehicle moves to the next target driving scene.

8. A multi-timbre emotion-adaptive speech synthesis device, characterized in that: The device comprises: An acquisition module is used to obtain the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scenario; a first processing module, configured to perform timbre encoding on the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scenario, obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generate auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scenario based on the auxiliary prompt timbre vector; a second processing module, configured to establish an emotional prosody mapping model from emotional state to acoustic parameters, and perform emotional prosody control on the auxiliary prompt timbre data using the emotional prosody mapping model and the target driving scenario to obtain emotional expression timbre data with emotional prosody; a third processing module, configured to determine vehicle driving noise data based on a current driving state of the intelligent vehicle in the target driving scenario, and optimize the emotion expression timbre data according to the vehicle driving noise data to obtain noise optimized timbre data; a compensation and optimization module, configured to perform acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and to perform spatial acoustic optimization on the noise-optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle, thereby obtaining acoustically optimized timbre data; An adjustment module is used to adjust the acoustically optimized timbre data for contextual emotional expression based on the text content of the driving assistance prompt text, so as to generate an auxiliary prompt synthesized speech for the smart vehicle corresponding to the target driving scenario.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-timbre emotion-adaptive speech synthesis method according to any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multi-timbre emotion-adaptive speech synthesis method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Speech playing method and device, and electronic equipment

    CN111627417A

  • Speech synthesis method and device and device for speech synthesis

    CN113409765A

  • Voice synthesis method and device thereof, electronic equipment, storage medium and program product

    CN114005428A

  • Speech synthesis method, speech synthesis device, electronic equipment and storage medium

    CN118629388A

  • Method and system for optimizing the speech intelligibility in a passenger compartment of a vehicle

    EP2814266A1

Cited By

  • Voice generation method and device based on tone parameterization control, equipment and storage medium

    CN121122235A