Method, device and equipment for multi-voice emotion adaptive speech synthesis and medium
By acquiring driver assistance prompt text from intelligent vehicles, performing timbre encoding and emotional prosody mapping, and optimizing timbre data in combination with vehicle status and acoustic features, the flexibility of emotional expression in intelligent cockpits has been solved, enabling emotionally adaptive speech synthesis in multiple scenarios and improving user engagement.
Patent Information
- Application Number
- CN202511212308.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing speech synthesis technology struggles to meet the emotional expression needs across multiple scenarios in smart cockpits. It lacks flexibility, cannot adapt to the emotional expression requirements of different scenarios, and reduces user engagement.
By acquiring the driving assistance prompts from intelligent vehicles in the target driving scenario, performing timbre encoding and emotional prosody mapping, and combining vehicle status and acoustic features to optimize the timbre data, emotionally adaptive synthesized speech is generated.
It achieves adaptation of various emotional data and contextual emotional expression in different scenarios, enhances user engagement, and meets users' needs for emotional expression in different scenarios.
Smart Images

Figure CN120726985B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of speech synthesis, and in particular to a multi-voice timbre emotion adaptive speech synthesis method, device, equipment and medium. BACKGROUND
[0002] Speech synthesis technology plays an important role in the cockpit of an intelligent vehicle in many aspects, which can effectively improve driving convenience, enhance safety, provide personalized services, and optimize human-computer interaction experience, etc. For example, personalized services can be provided according to the voice characteristics and habits of the driver; the driver can complete the control of the vehicle through simple voice instructions, and the response of the system can be fed back to the driver in the form of voice, realizing natural and smooth human-computer interaction, making the human-computer interaction interface more friendly, and greatly improving the convenience and safety of driving.
[0003] In the process of implementing the present application, the applicant found that there are at least the following problems in the prior art:
[0004] In the prior art, a large number of voice segments are pre-recorded and spliced, although the sound quality is natural, but a large voice library is needed, the flexibility is poor, it is difficult to express diversified emotions, and it is difficult to adapt to the multi-scene demand of the intelligent cockpit. The existing speech synthesis is difficult to generate speech with rich emotional changes, and cannot meet the demand of emotional expression in different scenes (navigation, warning, leisure, etc.) in the intelligent cockpit scene, reducing user stickiness. SUMMARY
[0005] The present application provides a multi-voice timbre emotion adaptive speech synthesis method, device, equipment and medium, which can meet the demand of emotional expression of users in different scenes, thereby effectively improving user stickiness.
[0006] In a first aspect, the embodiments of the present application provide a multi-voice timbre emotion adaptive speech synthesis method, comprising:
[0007] obtaining a driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene;
[0008] performing timbre coding on the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene to obtain an assistance prompt timbre vector corresponding to the driving assistance prompt text, and generating assistance prompt timbre data corresponding to the intelligent vehicle in the target driving scene based on the assistance prompt timbre vector;
[0009] establishing an emotion prosody mapping model from emotion state to acoustic parameter, and performing emotion prosody control on the assistance prompt timbre data through the emotion prosody mapping model and the target driving scene to obtain emotion expression timbre data with emotion prosody;
[0010] determine vehicle running noise data based on the current running state of the intelligent vehicle in the target running scene, and optimize the emotional expression timbre data based on the vehicle running noise data to obtain noise-optimized timbre data;
[0011] compensate the noise-optimized timbre data based on the vehicle attribute features of the intelligent vehicle, and optimize the noise-optimized timbre data based on the acoustic layout features of the intelligent vehicle to obtain acoustic-optimized timbre data;
[0012] adjust the acoustic-optimized timbre data based on the text content of the driving assistance prompt text to generate the synthesized voice of the intelligent vehicle corresponding to the target running scene.
[0013] In a second aspect, the embodiments of the present application also provide a multi-timbre emotion-adaptive speech synthesis device, comprising:
[0014] The acquisition module is configured to acquire the driving assistance prompt text corresponding to the target running scene of the intelligent vehicle.
[0015] The first processing module is configured to perform timbre coding on the driving assistance prompt text corresponding to the target running scene of the intelligent vehicle to obtain an assistance prompt timbre vector corresponding to the driving assistance prompt text, and generate assistance prompt timbre data corresponding to the target running scene of the intelligent vehicle based on the assistance prompt timbre vector.
[0016] The second processing module is configured to establish an emotion prosody mapping model of emotion state to acoustic parameter, and perform emotion prosody control on the assistance prompt timbre data through the emotion prosody mapping model and the target running scene to obtain emotional expression timbre data with emotion prosody.
[0017] The third processing module is configured to determine vehicle running noise data based on the current running state of the intelligent vehicle in the target running scene, and optimize the emotional expression timbre data based on the vehicle running noise data to obtain noise-optimized timbre data.
[0018] The compensation and optimization module is configured to compensate the noise-optimized timbre data based on the vehicle attribute features of the intelligent vehicle, and optimize the noise-optimized timbre data based on the acoustic layout features of the intelligent vehicle to obtain acoustic-optimized timbre data.
[0019] The adjustment module is configured to adjust the acoustic-optimized timbre data based on the text content of the driving assistance prompt text to generate the synthesized voice of the intelligent vehicle corresponding to the target running scene.
[0020] In a third aspect, the embodiments of the present application provide an electronic device, comprising:
[0021] one or more processors;
[0022] a memory for storing one or more programs,
[0023] When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-voice timbre emotion adaptive speech synthesis method of any embodiment of the present application.
[0024] In a fourth aspect, the embodiments of the present application provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-voice timbre emotion adaptive speech synthesis method of any embodiment of the present application.
[0025] The embodiments of the present application provide a multi-voice timbre emotion adaptive speech synthesis method, device, equipment and medium. The driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene is obtained. The driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene is timbre coded to obtain the auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene is generated based on the auxiliary prompt timbre vector. An emotion prosody mapping model from an emotion state to an acoustic parameter is established, and the emotion prosody control is performed on the auxiliary prompt timbre data through the emotion prosody mapping model and the target driving scene to obtain emotion expression timbre data with emotion prosody. The vehicle driving noise data is determined based on the current driving state of the intelligent vehicle in the target driving scene, and the emotion expression timbre data is optimized according to the vehicle driving noise data to obtain noise optimized timbre data. The noise optimized timbre data is compensated for acoustic characteristics based on the vehicle attribute features of the intelligent vehicle, and the noise optimized timbre data is spatially acoustically optimized according to the acoustic layout features of the intelligent vehicle to obtain acoustically optimized timbre data. The acoustically optimized timbre data is adjusted for context emotion expression based on the text content of the driving assistance prompt text to generate the auxiliary prompt synthesized speech corresponding to the target driving scene of the intelligent vehicle. That is, in the technical solution of the present application, the synthesized speech can be adapted to various emotion data in different scenes, and the context emotion expression in the synthesized speech can also be adjusted. In the prior art, a large number of voice segments are recorded in advance and spliced, and the emotion expression of different voice segments cannot be presented. Therefore, compared with the prior art, the multi-voice timbre emotion adaptive speech synthesis method, device, equipment and medium provided by the embodiments of the present application can meet the needs of users for emotional expression in different scenes, thereby effectively improving user stickiness. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 A flowchart of a multi-voice timbre emotion adaptive speech synthesis method provided by an embodiment of the present application is shown in the figure;
[0027] Figure 2A flowchart of a multi-voice timbre emotion-adaptive speech synthesis method provided for another embodiment of the application is shown in the figure;
[0028] Figure 3 A flowchart of a multi-voice timbre emotion-adaptive speech synthesis method provided for another embodiment of the application is shown in the figure;
[0029] Figure 4 A structural diagram of a multi-voice timbre emotion-adaptive speech synthesis device provided for an embodiment of the application is shown in the figure;
[0030] Figure 5 A structural diagram of an electronic device provided for an embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0031] The application will be further described below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely intended for the purpose of interpretation of the application and not for the limitation of the application. In addition, it should be noted that only parts related to the application are shown in the drawings for the purpose of description.
[0032] Figure 1 A flowchart of a multi-voice timbre emotion-adaptive speech synthesis method provided for an embodiment of the application is shown in the figure, which can be executed by a multi-voice timbre emotion-adaptive speech synthesis device or an electronic device. The device or the electronic device can be realized by software and / or hardware, and can be integrated in any smart device with network communication function. As shown in the figure, Figure 1 The multi-voice timbre emotion-adaptive speech synthesis method can include the following steps:
[0033] S101, obtaining a driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle.
[0034] In this step, the driving assistance prompt text can be a prompt word / sentence for voice broadcast in the intelligent cockpit. For example, in a turning scene, the driving assistance prompt text is "turn right 200 meters ahead"; in a deceleration scene, the driving assistance prompt text is "there is a traffic jam ahead, please slow down".
[0035] S102, encoding the driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle in terms of timbre, obtaining an assistance prompt timbre vector corresponding to the driving assistance prompt text, and generating assistance prompt timbre data corresponding to the target driving scene of the intelligent vehicle based on the assistance prompt timbre vector.
[0036] In this step, a unified multi-voice timbre acoustic space can be established, and all timbres share a set of model parameters, and different timbre data can be generated through conditional control.
[0037] For timbre vector injection: during synthesis, the system injects pre-extracted timbre vectors (e.g., 32-dimensional or 64-dimensional) as conditioning signals into the model; for example: to generate a "female announcer" timbre, inject the corresponding timbre vector [0.12, -0.34, 0.56,...], and the model can generate the timbre voice; for the same text input "road construction ahead, please slow down and drive slowly", inject different timbre vectors to generate different timbers such as professional women, mature men, etc.; for timbre ID embedding: use a simple ID number to represent different timbers, such as ID = 1 representing "standard male voice" and ID = 2 representing "standard female voice"; the model internally maps the ID to a timbre embedding vector as a conditioning control signal. In the navigation scenario, the system only needs to switch the timbre ID parameter to change the speaker characteristics of the voice prompt.
[0038] For timbre attribute parameterized control: the timbre is divided into multiple controllable attributes, and different timbres are generated by adjusting these parameters. For scene-driven automatic timbre selection: the system automatically selects the appropriate timbre conditioning parameters according to the use scenario; for example: use a highly clear and authoritative timbre for emergency warnings; use a lively and friendly timbre for entertainment information; use a professional and accurate timbre for navigation instructions.
[0039] After generating auxiliary prompt timbre data, the speaker characteristics can also be compressed to a very small dimensional vector through efficient timbre encoding, significantly reducing the storage requirements for multi-timbre, as shown in the following example.
[0040] Low-dimensional timbre embedding vectors: compress traditional 128-256 dimensional timbre vectors to 16-32 dimensions; e.g., by autoencoder training, compress high-dimensional timbre features to a low-dimensional latent space, original speaker feature vector [0.18, 0.25, -0.32,... (128 parameters)] is compressed to [0.41, -0.22, 0.15,... (16 parameters). Quantized and encoded representation: discretize continuous timbre vectors into integer codebook indices; e.g., quantize a 32-dimensional floating-point vector to an 8-bit integer vector [56, 13, 98, 27,...], which reduces each timbre from 128 bytes (32-dimensional floating-point) to 32 bytes (32-dimensional 8-bit integer). Hierarchical timbre coding: decompose timbre features into base timbre + modification features; e.g., timbre is represented as a base type index (e.g., "adult male" = 3) plus a differential feature vector [0.1, -0.2, 0.05,...]; where N similar timbres only need to store one base type and N small differential vectors. Sparse timbre representation: express complete timbre features using a small number of non-zero elements; e.g., a complete timbre is represented by [0, 0, 0.75, 0, 0, 0.21, 0, 0, 0, -0.48, 0,...] (most positions are 0); only store non-zero element positions and values, such as [(2, 0.75), (5, 0.21), (9, -0.48)]. Timbre principal component coding: only retain the principal components of timbre features; e.g., project voiceprint features to the most discriminative 4-8 dimensions by principal component analysis; representation: [0.87, -0.32, 0.15, 0.64] (only retain principal component projection values). Combined timbre parameter coding: decompose timbre into a combination of interpretable acoustic parameters; e.g., timbre is represented as [pitch baseline = high, formant distribution = wide, breathiness = medium, timbre = bright]; with the advantages of fewer parameters and interpretability, easy to intuitively adjust and control.
[0041] In acquiring an adaptive voice tone, a high-efficiency voice tone extraction network can be used to extract and copy target voice tone features with only a small amount of audio. The high-efficiency voice tone extraction network is based on the following network architecture: an improved lightweight ResNet variant as the backbone network; a hybrid architecture combining self-attention mechanisms and depth separable convolutions; and a three-stage structure of encoder-pooling-projector. The specific network structure includes: a front-end feature extraction layer: 2-3 layers of 1D convolution layers (convolution kernel size 3-7) to extract basic acoustic features, each layer followed by batch normalization and LeakyReLU activation function; an intermediate representation layer: 4-6 lightweight residual blocks, each block containing a depth separable convolution instead of a standard convolution, selectively introducing a self-attention mechanism to capture long-range speech feature dependencies, and using a skip connection to preserve multi-scale information; a statistical pooling layer: statistical pooling is performed on the time dimension to extract global mean and standard deviation statistical features, and parameter reduction is achieved through weight sharing; a projection head: 2-3 layers of fully connected layers, gradually reducing the dimension to the target voice tone vector dimension (usually 16-64 dimensions), using a nonlinear activation function and regularization techniques. Furthermore, parameter sharing and factorization techniques are used to reduce network parameters, introduce quantization-aware training to prepare for subsequent quantization deployment, achieve parameter sparsification, and prune non-critical connections. The meta-learning framework supports rapid adaptation to new voice tones, the prototype network design captures the core features of voice tones, and an incremental learning mechanism is introduced to quickly update voice tone representations.
[0042] In S103, a sentiment prosody mapping model is established, and the auxiliary prompt tone data is controlled in sentiment prosody through the sentiment prosody mapping model and the target driving scene, to obtain sentiment expression tone data with sentiment prosody.
[0043] In this step, the sentiment prosody mapping model can accurately map the relationship between the sentiment state and the acoustic parameters (such as speech rate, pitch, energy, etc.).
[0044] For example, for the basic emotion mapping model, the emotional state is happy / excited, and the acoustic parameter is speech rate: +15~30% (relative to the standard speech rate), pitch: average value increased by 20~40Hz, pitch range: expanded by 40~60%, energy: overall enhanced by 2~4dB, high frequency band enhanced more obviously, voice quality parameter: harmonic-to-noise ratio increased, spectral tilt increased. The precise mathematical mapping relationship is: F0happy=F0_neutral×(1+0.2×E_intensity), wherein E_intensity represents the emotional intensity. For the multi-dimensional emotional parameter matrix: emotional intensity dimension mapping: define an intensity coefficient I of 0~1 for each emotion, speech rate adjustment formula: Speed=Speed_base×(1+Direction_factor×I×Max_change_ratio); for example: when the sad emotion (I=0.7), the speech rate = base speech rate × (1+(-1)×0.7×0.4) = base speech rate × 0.72; emotional mixed mapping: define an emotional mixed vector E=[E_happy,E_sad,E_angry,E_neutral], the sum of each component is 1, acoustic parameter calculation: P=∑(E_i×P_i), wherein P_i is the parameter corresponding to each pure emotion; for example: E=[0.3,0.1,0.6,0], F0=0.3×F0happy+0.1×F0_sad+0.6×F0_angry.
[0045] In S104, vehicle driving noise data is determined based on the current driving state of the intelligent vehicle in the target driving scene, and the emotional expression timbre data is optimized according to the vehicle driving noise data, to obtain noise-optimized timbre data.
[0046] In this step, the emotional expression timbre data is optimized by establishing a noise-clarity mapping model and dynamically optimizing the synthesized speech characteristics according to the real-time acoustic state in the vehicle.
[0047] Real-time noise perception analysis: real-time acoustic environment data is collected by using an in-vehicle microphone array, and a lightweight spectrum analysis is applied to extract a noise feature vector N=[n1, n2,..., n k ] including noise spectrum distribution, signal-to-noise ratio, acoustic reverberation and other key indicators, and a sliding window analysis is used to capture the dynamic changes of the noise. Clarity influence prediction mechanism: an acoustic masking prediction function M(f, N) is constructed to estimate the masking degree of each frequency band of speech, a lightweight neural network is used to establish a nonlinear mapping: N→influence map I, the prediction model outputs the speech feature dimension most susceptible to the current noise environment, and a decision tree+linear regression hybrid architecture is adopted to ensure real-time computing capability. Adaptive parameter generation: a compensation parameter generation module C(I) is designed to output an acoustic characteristic optimization parameter vector P, and the parameter vector includes: frequency gain G=[g1, g2,..., gn ], time domain modulation D = [d1, d2,..., d m ], establish a joint optimization function of semantic importance weight W and noise masking degree M, calculate the optimization parameter: P = argmax (clarity (S, N, P) - a · complexity (P)). Dynamic optimization strategy execution: acoustic parameter real-time adjustment: S' = Transform (S, P), key band selective enhancement: apply stronger compensation to the frequency band severely affected by noise, time domain structure optimization: adjust the time domain envelope to ensure that the key phonemes are clear and distinguishable, energy distribution remodeling: redistribute energy from areas not affected by noise to affected areas. Closed-loop feedback optimization: simulate auditory perception to evaluate the clarity of synthesized speech in the current noise, evaluate the optimization effect through a lightweight auditory model A: S_clarity = A (S', N), apply online learning to update the mapping model parameters to adapt to new noise types, establish an optimization history cache to speed up parameter generation in similar noise environments. Intelligent scene adaptation: perceive vehicle state (vehicle speed, window opening and closing, air conditioner fan, etc.), predict noise changes, pre-calculate the corresponding parameters of possible noise scenes, achieve ultra-low delay response, automatically adjust the optimization strategy focus combined with driving mode and road condition information, learn individual clarity preferences through user feedback, and fine-tune optimization parameters.
[0048] S105, based on the vehicle attribute features of the intelligent vehicle, acoustically compensating the noise-optimized timbre data, and based on the acoustic layout features of the intelligent vehicle, spatially optimizing the noise-optimized timbre data to obtain acoustically optimized timbre data.
[0049] In this step, by constructing a parameterized vehicle acoustic model, the sound loss caused by the acoustic characteristics of different vehicle models is automatically compensated. At the same time, the spatial performance of speech synthesis is optimized combined with the acoustic layout characteristics of the vehicle, improving directivity and clarity.
[0050] Optionally, the acoustically compensating the noise-optimized timbre data based on the vehicle attribute features of the intelligent vehicle comprises:
[0051] Obtaining an acoustic characteristic vector corresponding to the vehicle attribute features of the intelligent vehicle, the acoustic characteristic vector being used to describe a plurality of acoustic characteristic dimensions, the plurality of acoustic characteristic dimensions comprising: spatial volume, reflectivity, and distribution of absorbing material; establishing a mapping and conversion model of acoustic characteristics to sound distortion, and determining a sound distortion degree corresponding to the acoustic characteristic vector through the mapping and conversion model; and acoustically compensating the noise-optimized timbre data based on the sound distortion degree corresponding to the acoustic characteristic vector through a characteristic compensation function.
[0052] In this step, the vehicle acoustic characteristic modeling framework: abstract different vehicle acoustic characteristics into a parameter vector, each parameter representing an acoustic characteristic dimension: spatial volume, reflectivity, sound-absorbing material distribution, etc., to establish an acoustic characteristic to sound distortion mapping function D(A) that describes the impact of a specific vehicle on speech. Multi-level acoustic representation structure: physical layer: cabin geometry, material properties, speaker position, etc., frequency response layer: frequency response curve, reverberation time, directivity index, etc., perception layer: intelligibility index, spatial perception, sound image localization, etc. Compensation mechanism design: inverse filter design: H -1 (f)=1 / D(A,f) pre-compensate acoustic characteristics, adaptive gain control: apply independent gain adjustment for different frequency bands, dynamic range processing: adjust compression ratio and threshold according to vehicle acoustic characteristics; vehicle characteristic acquisition method: professional measurement: use test signals to collect acoustic measurement data in each vehicle, parameter extraction: apply system identification algorithm to extract parameters from measurement data, model normalization: standardize parameter space to enable cross-vehicle comparison. Real-time adaptive algorithm: detect the current vehicle ID, load the corresponding acoustic parameters from the database, adjust the synthesized speech spectral characteristics through the compensation function C(f,A), monitor acoustic feedback, fine-tune compensation parameters, and achieve closed-loop optimization. Lightweight implementation strategy: parameter compression: compress the complete acoustic model to <100 parameters. Computational optimization: use approximation algorithms instead of precise acoustic simulation. Look-up table acceleration: pre-compute common compensation values and obtain them in real-time through interpolation. For example, luxury sedan: compensate for sound-absorbing characteristics of materials and adjust mid-frequency gain; compact SUV: enhance low-frequency intelligibility and compensate for spatial reverberation; open-top sports car: dynamically adjust acoustic parameters to cope with open / closed top conditions.
[0053] Optionally, the noise-optimized timbre data is spatially acoustically optimized according to the acoustic layout characteristics of the intelligent vehicle to obtain acoustically optimized timbre data, including:
[0054] Based on the pre-compensation filter, the intelligibility of the timbre data at the driving position in the intelligent vehicle in the noise-optimized timbre data is enhanced for directional optimization of the timbre at the driving position in the intelligent vehicle; the intelligibility of the timbre data at the non-driving position in the intelligent vehicle in the noise-optimized timbre data is reduced for directional optimization of the timbre at the non-driving position in the intelligent vehicle.
[0055] In this step, the driving position orientation is optimized: the transfer function of the driving position relative to the in-vehicle loudspeaker is analyzed, a pre-compensation filter is applied to enhance the clarity of the sound received at the driving position, a frequency-dependent phase delay is designed, and constructive interference of sound waves at the driving position is created. Specific example: for the common "A-pillar + dashboard" loudspeaker layout, a -3dB / +6dB crossover compensation is applied; the A-pillar loudspeaker is enhanced in the 2-4kHz frequency band (+3dB), and low frequencies (<300Hz) are suppressed; the dashboard loudspeaker is enhanced in the 4-8kHz frequency band (+4dB), and the mid-frequency balance is maintained; an intelligent delay difference of about 0.2-0.3ms is introduced between the two sound sources, forming a sound image focus at the driving position. This can improve the speech intelligibility index (SII) at the driving position by 25% without affecting the listening experience at the passenger position.
[0056] Multi-zone sound personalization: build an acoustic model of multiple listening areas in the vehicle, design a spatially selective enhancement algorithm, achieve "acoustic zoning", and apply beamforming technology to create a directional sound field. Specific example: navigation instruction scenario: mainly enhance the sound clarity at the driving position; entertainment content scenario: balanced optimization of sound field distribution throughout the vehicle; when navigating at the driving position, the front row sound source balance is adjusted to 65:35 (left:right), forming a sound image biased towards the driver; different spatial optimization presets are used for voice instructions and music playback. Acoustic shadow zone avoidance: identify acoustic shadow zones in the vehicle (such as behind the seat back, roof curve reflection points), rebalance multi-loudspeaker output, avoid shadow zone propagation paths, and dynamically adjust EQ parameters to compensate for the shielded frequency band. Specific example: identify the strong absorption characteristics of the rear seats of an SUV for high frequencies (>3kHz), and when playing voice at the back row, pre-compensate and enhance the 3-5kHz frequency band (+4.5dB), use roof reflection to enhance the sound field at the back row, and increase the energy of specific frequency bands in the 100-300Hz range for front row loudspeaker output. At the same time, a correlation model between vehicle speed and in-vehicle noise spectrum is established, and a speed-adaptive spatial enhancement algorithm is designed to dynamically adjust the sound field characteristics with speed. Specific example: low speed (<40km / h): standard sound field balance, slightly improve high frequency clarity; medium speed (40-80km / h): start compensating for wind noise, enhance 1-2kHz frequency band, sound image moves forward by 10%; high speed (>80km / h): intensive compensation, enhance 2-4kHz frequency band (+6dB), compress dynamic range by 30%, sound image moves forward by 25%.
[0057] S106, based on the text content of the driving assistance prompt text, the acoustic optimization timbre data is adjusted in context emotional expression to generate a synthesized voice of the intelligent vehicle corresponding to the target driving scene.
[0058] In this step, the emotional expression mode is automatically adjusted according to the text semantics and context to enhance the naturalness of emotion. For example, the text semantic understanding layer: a lightweight BERT variant extracts text semantic vectors, and a keyword / phrase detection identifies emotional triggers; a lightweight dependency syntax analysis identifies sentence structure, and adjusts the basic emotional mode according to the sentence type (statement, question, command, exclamation); a pre-trained multi-class classifier maps the text to the emotional category space. The context perception mechanism: maintains a sliding window to record recent interaction history, adjusts the current emotional expression based on historical information; integrates vehicle state information (speed, location, surrounding environment), adjusts the emotional expression intensity according to the scene danger level; analyzes the user's response mode to voice prompts, and adjusts the emotional expression mode according to the user's reaction.
[0059] Context-related parameter adjustment includes semantic importance weighting and syntactic structure adaptation. For semantic importance weighting, the emotional parameter influence is adjusted according to the text semantic importance W (0~1), and the key information paragraph: Energy=Energy_neutral×(1+W×E_factor), the non-key paragraph uses weaker emotional mapping, and maintains clarity. For syntactic structure adaptation, the beginning of the sentence: the emotional parameters gradually increase, and the mapping coefficient increases from 0.3 to 1.0, the end of the sentence: adjust the F0 contour according to the sentence type (statement sentence down, question sentence up), pause: emotional state affects pause duration, such as sadness emotion prolongs pause 20~50%.
[0060] In addition, emotional-voice interaction mapping can also be performed, such as voice characteristic compensation: female voice: the F0 change caused by emotion is 15~25% larger than male voice; child voice: emotional energy change is relatively small, but pitch change is more significant, mapping formula: ΔF0=Base_ΔF0×Voice_type_factor; brand voice consistency maintenance: define the range of emotional parameter changes allowed by the brand voice, and apply a nonlinear compression function when the range is exceeded: P=max_range×tanh(P / max_range).
[0061] Emotion intensity continuous control: Realize the continuous adjustment ability of emotion intensity, meet the emotional expression needs in different scenes. Specific examples: emotion intensity a = 0.3, F0 = 0.7 * F0 neutral+ 0.3 * F0_happy; Excited emotion intensity a = 0.8, speech speed increases sigmoid(2 * (0.8-0.5))≈0.73 times the maximum change; Set A = 0.8 (high arousal), V = 0.5 (positive valence), D = 0.3 (moderate dominance) to generate "mild excitement" emotion; Warning information gradually increases, a(t) linearly increases from 0.2 to 0.8; Navigation scene emotion intensity adjustment: ordinary navigation prompt: excited emotion a = 0.2 (slightly lively), about to reach the destination: excited emotion a = 0.5 (moderate pleasant), route re-planning: anxious emotion a = 0.3 (slightly urgent); When the collision risk increases, the warning voice emotion intensity smoothly transitions from 0.3 to 0.9.
[0062] A representation decoupling mechanism for emotion and timbre can also be designed to achieve independent control of the two, realize emotion and timbre decoupling, and accurately retain speaker identity features when changing emotion state, avoiding identity information leakage. For example, orthogonal representation space design: independent feature subspace: decompose the acoustic representation space into mutually orthogonal timbre subspace and emotion subspace; Dimension independent allocation: explicitly divide the dimensions of the representation space, assign different dimensions to encode different attributes, and force feature separation through structured design. Adversarial learning decoupling mechanism: identity preserving adversarial training: introduce timbre discriminator Disc_id and emotion discriminator Disc_emo, Disc_id attempts to extract timbre information from emotion representation; Cross attribute adversarial constraint: simultaneously train emotion discriminator Disc_emo cannot predict emotion category from timbre representation Sid, forming a bidirectional adversarial constraint to ensure the mutual independence of the two representations. Information bottleneck compression mechanism: representation channel constraint: set information bottleneck for timbre and emotion encoders to limit the amount of information passing through, timbre encoder Enc_id is designed to only retain the minimum sufficient statistics related to speaker identity, and emotion encoder Enc_emo only retains the minimum information required for emotion discrimination. Mutual information minimization: explicitly minimize the mutual information between timbre representation and emotion representation, and approximate calculate and optimize mutual information through neural estimator and adversarial training. Consistency constraint: the timbre representation of the same speaker under different emotion conditions should be highly similar, and the emotion representation of the same emotion under different speaker conditions should be highly similar.
[0063] The multi-voice emotion adaptive speech synthesis method provided by the embodiment of the application comprises the following steps: obtaining driving assistance prompt text corresponding to a target driving scene of an intelligent vehicle; performing voice coding on the driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle to obtain an auxiliary prompt voice vector corresponding to the driving assistance prompt text, and generating auxiliary prompt voice data corresponding to the target driving scene of the intelligent vehicle based on the auxiliary prompt voice vector; establishing an emotion prosody mapping model of an emotion state to an acoustic parameter, and performing emotion prosody control on the auxiliary prompt voice data through the emotion prosody mapping model and the target driving scene to obtain emotion expression voice data with emotion prosody; determining vehicle driving noise data based on a current driving state of the intelligent vehicle in the target driving scene, and optimizing the emotion expression voice data according to the vehicle driving noise data to obtain noise-optimized voice data; performing acoustic characteristic compensation on the noise-optimized voice data based on vehicle attribute characteristics of the intelligent vehicle, and performing spatial acoustic optimization on the noise-optimized voice data according to acoustic layout characteristics of the intelligent vehicle to obtain acoustic-optimized voice data; and performing context emotion expression adjustment on the acoustic-optimized voice data based on text content of the driving assistance prompt text to generate auxiliary prompt synthesized speech of the intelligent vehicle corresponding to the target driving scene. That is, in the technical solution of the application, the synthesized speech can be adapted to multiple emotion data in different scenes, and the context emotion expression in the synthesized speech can also be adjusted. In the prior art, a large number of voice segments are pre-recorded and spliced, and the emotion expression of different voice segments cannot be presented. Therefore, compared with the prior art, the multi-voice emotion adaptive speech synthesis method, device, equipment and medium provided by the embodiment of the application can meet the needs of users for emotional expression in different scenes, thereby effectively improving user stickiness.
[0064] Figure 2 The flowchart of the multi-voice emotion adaptive speech synthesis method provided by another embodiment of the application is shown. Based on the above technical solution, further optimization and expansion can be performed, and the above-mentioned various optional embodiments can be combined. As shown in Figure 2 the multi-voice emotion adaptive speech synthesis method can comprise the following steps:
[0065] S201, input the auxiliary prompt voice vector into a preset network model, and determine initial voice data based on the output of the preset network model.
[0066] In this step, the preset network model is, for example, a small-parameter efficient model. The small-parameter efficient model architecture can solve the problem of limited computing resources in the intelligent cockpit environment in speech synthesis, while ensuring the quality of speech synthesis.
[0067] The hierarchical lightweight design of the small-parameter high-efficiency model specifically includes: a hybrid neural architecture design: integrating an improved non-autoregressive structure and an efficient encoder, introducing a depth separable convolution to replace a standard convolution, and significantly reducing the model parameter quantity; the non-autoregressive structure and the efficient encoder can improve the synthesis speed, reduce the parameter quantity, maintain the synthesis quality, enhance the parallel computing capability, and improve the stability of long sentences; a parameter-shared multi-task encoder: adopting a key feature layer sharing strategy to greatly reduce the feature extraction parameter quantity and improve the cross-task feature extraction capability; the key feature layer sharing strategy is a technical method adopted in the parameter-shared multi-task encoder, specifically including: feature extraction layer sharing, a selective sharing mechanism, hierarchical feature mapping, task-specific branch design, and cross-task knowledge transfer; a non-autoregressive decoding framework: abandoning autoregressive dependence, realizing high-speed parallel decoding, and designing an adaptive length control mechanism to solve the prosody prediction problem; the adaptive length control mechanism mainly solves the prosody prediction problem in the non-autoregressive speech synthesis; and an efficient acoustic parameter conversion: designing a parameter-efficient acoustic feature generator to realize high-fidelity audio reconstruction with a small parameter quantity; the acoustic feature generator plays a role in the speech synthesis system: audio reconstruction core, high-fidelity restoration, feature compression and expression, cross-feature mapping color and emotion integration, computing resource optimization, and cross-feature mapping.
[0068] A hierarchical distillation framework is designed to efficiently transfer large model knowledge to small parameter models in a multi-stage manner, a dynamic sparse training technology based on parameter importance is used to retain key connections while greatly reducing redundant calculations, a mixed precision quantization strategy is used to allocate different quantization bit widths to different levels and implement different quantization strategies, and the model volume is further compressed.
[0069] S202. Split the initial timbre data into multiple controllable timbre attributes, and configure attribute parameters for the multiple controllable timbre attributes to obtain target timbre data with different timbre attributes.
[0070] In this step, the timbre is split into multiple controllable timbre attributes (such as pitch, roughness, breathiness, etc.), and different timbres are generated by adjusting these parameters. For example: set pitch = high, breathiness = strong, roughness = low to generate a “young female” timbre; adjust pitch = low, breathiness = weak, roughness = high to generate a “calm male” timbre.
[0071] S203. Limit the timbre condition parameters of the target timbre data based on the scene category to which the target driving scene belongs, and mix the target timbre data after the timbre condition parameter limitation to obtain the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene.
[0072] In this step, the spectral feature adaptive normalization method designed by small sample audio reduces the noise interference caused by environmental and recording condition differences to limit the timbre condition parameters of the target timbre data.
[0073] For mixed timbre generation: generate a mixed timbre by weighted combination of multiple base timbre vectors; for example: mix the "professional male voice" and "mild female voice" timbre vectors in a 7:3 ratio to create a new timbre that is authoritative but not losing affinity.
[0074] The multi-timbre emotion adaptive speech synthesis method proposed in the embodiments of the present application inputs the auxiliary prompt timbre vector into the preset network model, determines the initial timbre data based on the output of the preset network model, splits the initial timbre data into multiple controllable timbre attributes, configures attribute parameters for the multiple controllable timbre attributes, obtains target timbre data with different timbre attributes, limits the timbre condition parameters of the target timbre data based on the scene category to which the target driving scene belongs, and mixes the target timbre data after the timbre condition parameter limitation to obtain the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene, thereby obtaining the auxiliary prompt timbre data adapted to the target driving scene.
[0075] Figure 3 The flowchart of the multi-timbre emotion adaptive speech synthesis method provided by another embodiment of the present application. Based on the above technical solutions, further optimization and expansion can be performed, and the method can be combined with the above various optional embodiments. As shown in Figure 3 The multi-timbre emotion adaptive speech synthesis method can include the following steps:
[0076] S301, detect the speech data capacity corresponding to the auxiliary prompt synthesis speech.
[0077] S302, if the speech data capacity corresponding to the auxiliary prompt synthesis speech exceeds the preset capacity threshold, perform speech splitting on the auxiliary prompt synthesis speech based on the driving assistance prompt text to obtain multiple auxiliary prompt speech segments.
[0078] In this step, if the auxiliary prompt synthesis speech whose speech data capacity exceeds the preset capacity threshold is "front complex roundabout, please take the third exit to turn into Qingshan Road", it is determined that the speech content complexity of the auxiliary prompt synthesis speech is high, and it is split into three auxiliary prompt speech segments: "the front is a complex roundabout" + "please take the third exit" + "turn into Qingshan Road".
[0079] S303, generate the speech response intensity corresponding to each auxiliary prompt speech segment based on the speech context logic between the multiple auxiliary prompt speech segments, and adjust the response intensity of the auxiliary prompt synthesis speech based on the speech response intensity corresponding to each auxiliary prompt speech segment.
[0080] In this step, in combination with the above, the number "third" is emphasized in a voice intensity enhancement manner, and a turning prompt tone is added for assistance, and key information is repeated as appropriate, such as "third", "green mountain".
[0081] The multi-voice emotion adaptive speech synthesis method proposed in the embodiment of the application detects the speech data capacity corresponding to the auxiliary prompt synthesis speech, and segments and splits complex speech data to emphasize key information, so that complex navigation speech is easier to understand and reduces the cognitive burden of drivers.
[0082] In addition, a text semantic importance evaluation mechanism can also be established, and an optimized expression strategy is adopted for safety-related information. For example, a multi-dimensional semantic analysis framework and a hierarchical semantic analysis: a multi-level analysis structure of vocabulary-phrase-sentence is constructed; a key information identification algorithm: the core elements of the content are extracted using a lightweight Transformer variant; information entropy evaluation: calculate the information amount and irreplaceability index of the text segment. Importance scoring model, multi-feature vectorization: map semantic units to importance feature space; cascaded classifier architecture: combine rule-based models and neural network models for two-stage evaluation; context-dependent weight dynamic adjustment: adjust the importance of local semantic units according to the complete context. Scene-based importance mapping, safety keyword identification: establish a special vocabulary table related to driving safety and priority; navigation key information extraction: design special evaluation rules for direction, distance, landmark, and other navigation key elements; functional expression identification: identify important functional expressions that represent system operations and control instructions. Semantic importance quantification mechanism, multi-level importance division: divide the text content into four levels: key, important, ordinary, and auxiliary; probability distribution representation: assign importance probability distribution to each semantic unit instead of a single score; differentiated decay model: different types of important information use different context decay models.
[0083] Optionally, it further comprises:
[0084] The auxiliary prompt synthesis speech corresponding to the target driving scene is played in the intelligent vehicle; if the warning prompt level of the auxiliary prompt synthesis speech is higher than the preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesis speech, the scene key speech in the auxiliary prompt synthesis speech is extracted; the voice intensity of the auxiliary prompt synthesis speech is enhanced based on the voice intensity of the scene key speech, and the voice intensity of the auxiliary prompt synthesis speech is adjusted based on the voice intensity after the voice intensity enhancement; the auxiliary prompt synthesis speech is continuously played in the intelligent vehicle.
[0085] In this step, the timbre feature enhancement mechanism includes data layer enhancement, feature layer enhancement, model layer enhancement and inference layer enhancement. In data layer enhancement, random mask enhancement: in the time-frequency domain, random mask in part of the area, forcing the network to learn more stable timbre representation, rather than relying on local features that may not be stable; synthetic sample augmentation: based on limited samples to generate additional synthetic samples to expand the training data, such as creating variants by adding controllable noise, variable speed, etc. In feature layer enhancement, spatial constraint regularization: introduce distance constraint in timbre embedding space to ensure that the representation of different samples of the same speaker is tightly clustered; contrast noise filtering: design a contrast learning framework to identify and filter out features unrelated to timbre identity, leaving stable identity information; timbre feature decomposition: decompose the timbre feature into two parts: content-independent and content-dependent, and focus on extracting and enhancing the content-independent part. In model layer enhancement, multi-view feature fusion: extract timbre features from different angles (time domain, frequency domain, cepstrum domain) and fuse multi-view information through attention mechanism; hierarchical feature aggregation: make full use of shallow and deep network features, shallow layer captures basic timbre features, and deep layer extracts abstract identity features; dynamic confidence weighting: assign dynamic confidence weights to different feature dimensions, and prefer to retain high-confidence features when aggregating. In inference layer enhancement, integrated inference strategy: extract features from multiple segments in small samples, and obtain more stable representation through weighted average or median filtering; uncertainty modeling: establish an uncertainty model for timbre representation, and dynamically adjust the feature influence according to the uncertainty level during synthesis; self-calibration mechanism: introduce a reference timbre library to automatically calibrate the extracted timbre features and reduce bias.
[0086] Therefore, by enhancing the voice intensity of the scene key voice in the auxiliary prompt synthesized voice, and adjusting the voice intensity of the auxiliary prompt synthesized voice based on the scene key voice after voice intensity enhancement; the way of continuously playing the auxiliary prompt synthesized voice in the intelligent vehicle can continuously prompt the driver, which can effectively avoid dangerous driving.
[0087] Optionally, it also includes:
[0088] The core voice segment of the auxiliary prompt synthesized voice corresponding to the target driving scene of the intelligent vehicle is extracted, and a plurality of core voice segments in the auxiliary prompt synthesized voice are obtained. The voice is reorganized according to the context relationship between the plurality of core voice segments, and the auxiliary prompt core voice is obtained. The auxiliary prompt core voice is backed up to play the voice when the intelligent vehicle moves to the next target driving scene.
[0089] In this step, the synthesized voice is played in the same scene, saving voice synthesis resources, or setting a backup scheme for the basic timbre of the core function, which can ensure basic voice services even if the main system fails.
[0090] Core functions include safety warning class functions, key navigation instructions, system status notifications, and emergency communication functions. Safety warning class functions such as collision warning prompts ("Vehicle approaching, please emergency brake"); system failure warning ("Brake system abnormal, please park safely"); vehicle state warning ("Tire pressure is too low, please check as soon as possible"); emergency avoidance prompt ("Obstacle ahead, please turn immediately"). Key navigation instructions such as upcoming turn prompts ("Turn right at the next intersection"); high-speed exit reminders ("500 meters to exit the highway"); lane change instructions ("Please change to the left lane"); route re-planning notifications ("Route is being re-planned"). System status notifications such as automatic driving mode switching ("Exiting automatic driving mode"); key system activation / deactivation ("Emergency braking system activated"); severe fuel / charge shortage reminder ("Battery level less than 10%, please charge as soon as possible"). Emergency communication functions such as emergency call status ("Calling emergency rescue phone"); safety help confirmation ("Do you need to contact rescue services?"); automatic accident reporting ("Accident location has been automatically reported").
[0091] In addition, an acoustic feature enhancement module is designed for the acoustic characteristics of the intelligent cockpit to selectively enhance key frequency band information. For example, a mapping of human ear sensitivity to different frequencies is established, higher gain is applied to the key frequency band (1 kHz-4 kHz) for voice recognition, and dynamic adjustment of different frequency band weight distribution is performed; the typical noise spectrum characteristics in the vehicle are predicted; the voice features orthogonal to the noise frequency band are specifically improved; the spectral valley filling and peak protection strategies are implemented. Real-time identification of voice content key areas; enhanced processing of information-intensive areas such as consonant boundaries; application of attention mechanism to highlight voice recognition key elements. Contrast enhancement. Improve the contrast of vowels-consonants, voiced-unvoiced, etc.; enhance the performance of prosodic changes to ensure clear emotional expression.
[0092] By introducing a specialized adversarial training framework, the intelligibility of synthesized speech in the vehicle noise environment is improved. Self-attention noise adaptation layer: Introduce a noise adaptation self-attention layer in the generator to dynamically adjust the attention weight according to the noise condition, selectively enhance the speech features related to the current noise characteristics; multi-scale spectral discriminator: the discriminator evaluates the speech quality at multiple time-frequency resolutions simultaneously, designs specific loss weights, emphasizes the areas overlapping with the in-vehicle noise frequency band, and applies perceptual weighted spectral loss, which aligns with the human ear masking characteristics; vehicle-specific noise modeling: establish a noise characteristic database for different vehicles, use vehicle-specific noise conditions in training, and realize targeted optimization of synthesized speech for specific acoustic environments.
[0093] The emotion transition and mixing mechanism is established to smoothly transition from calm to alertness, and the emotion parameters are gradually changed, including pitch, speech rate, and timbre adjustment sentence by sentence. The situation drives the fusion of emotions, including automatic adjustment of emotion intensity based on vehicle speed changes and adjustment of voice emotional color based on weather conditions. Multi-modal emotion coordination, including voice emotion and in-vehicle light coordination, volume and timbre adjustment linked to driving mode. In addition, user preferences can be adapted, such as learning user preferences for timbre and emotional expression through interaction history; automatically adjusting timbre and emotional parameters based on user preferences to provide personalized experience; and supporting cloning of driver voice as system timbre to create a unique personalized experience.
[0094] Figure 4 The structure diagram of the multi-timbre emotion adaptive speech synthesis device provided by the embodiment of the application is shown in the figure. Figure 4 As shown in the figure, the multi-timbre emotion adaptive speech synthesis device comprises:
[0095] The acquisition module 401 is configured to acquire a driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle.
[0096] The first processing module 402 is configured to encode the timbre of the driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle, obtain an auxiliary prompt timbre vector corresponding to the driving assistance prompt text, and generate auxiliary prompt timbre data corresponding to the target driving scene of the intelligent vehicle based on the auxiliary prompt timbre vector.
[0097] The second processing module 403 is configured to establish an emotion prosody mapping model of emotion state to acoustic parameters, and control the emotion prosody of the auxiliary prompt timbre data through the emotion prosody mapping model and the target driving scene, to obtain emotion expression timbre data with emotion prosody.
[0098] The third processing module 404 is configured to determine vehicle driving noise data based on the current driving state of the intelligent vehicle in the target driving scene, and optimize the emotion expression timbre data according to the vehicle driving noise data, to obtain noise-optimized timbre data.
[0099] The compensation and optimization module 405 is configured to compensate the acoustic characteristics of the noise-optimized timbre data based on the vehicle attribute characteristics of the intelligent vehicle, and optimize the spatial acoustics of the noise-optimized timbre data based on the acoustic layout characteristics of the intelligent vehicle, to obtain acoustically optimized timbre data.
[0100] The adjustment module 406 is configured to adjust the context emotion expression of the acoustically optimized timbre data based on the text content of the driving assistance prompt text, to generate auxiliary prompt synthesized speech corresponding to the target driving scene of the intelligent vehicle.
[0101] Optionally, the first processing module 402 is specifically configured to:
[0102] The auxiliary prompt timbre vector is input into a preset network model, initial timbre data is determined based on an output of the preset network model, the initial timbre data is split into a plurality of controllable timbre attributes, attribute parameter configuration is performed on the plurality of controllable timbre attributes, target timbre data with different timbre attributes is obtained, the target timbre data is subjected to timbre condition parameter limitation based on a scene category to which the target driving scene belongs, and the target timbre data subjected to the timbre condition parameter limitation is subjected to timbre mixing, so as to obtain the auxiliary prompt timbre data corresponding to the intelligent vehicle in the target driving scene.
[0103] Optionally, the compensation and optimization module 405 is specifically configured to:
[0104] An acoustic characteristic vector corresponding to a vehicle attribute feature of the intelligent vehicle is obtained, the acoustic characteristic vector is used to describe a plurality of acoustic characteristic dimensions, the plurality of acoustic characteristic dimensions include: a space volume, a reflectivity, and a distribution of an absorbing material; a mapping conversion model of the acoustic characteristic to the sound distortion is established, and a sound distortion degree corresponding to the acoustic characteristic vector is determined through the mapping conversion model; and the noise optimization timbre data is subjected to acoustic characteristic compensation based on the sound distortion degree corresponding to the acoustic characteristic vector through a characteristic compensation function.
[0105] Optionally, the compensation and optimization module 405 is specifically configured to:
[0106] The clarity of the timbre data at the driving position in the intelligent vehicle in the noise optimization timbre data is enhanced based on the pre-compensation filter, so as to perform directional optimization of the timbre at the driving position in the intelligent vehicle; and the clarity of the timbre data at the non-driving position in the intelligent vehicle in the noise optimization timbre data is reduced, so as to perform directional optimization of the timbre at the non-driving position in the intelligent vehicle.
[0107] Optionally, the method further comprises a fourth processing module.
[0108] The fourth processing module is configured to detect a voice data capacity corresponding to the auxiliary prompt synthesized voice; if the voice data capacity corresponding to the auxiliary prompt synthesized voice exceeds a preset capacity threshold, the auxiliary prompt synthesized voice is subjected to voice splitting based on the driving auxiliary prompt text, so as to obtain a plurality of auxiliary prompt voice segments; a voice response intensity corresponding to each auxiliary prompt voice segment is generated based on a voice context logic between the plurality of auxiliary prompt voice segments, and the auxiliary prompt synthesized voice is subjected to response intensity adjustment based on the voice response intensity corresponding to each auxiliary prompt voice segment.
[0109] Optionally, the method further comprises a fifth processing module.
[0110] The fifth processing module is configured to play the auxiliary prompt synthesized voice corresponding to the target driving scene in the intelligent vehicle; if the warning prompt level of the auxiliary prompt synthesized voice is higher than the preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesized voice, the scene key voice in the auxiliary prompt synthesized voice is extracted; the scene key voice in the auxiliary prompt synthesized voice is subjected to voice intensity enhancement, and the auxiliary prompt synthesized voice is subjected to voice intensity adjustment based on the scene key voice subjected to the voice intensity enhancement; and the auxiliary prompt synthesized voice is continuously played in the intelligent vehicle.
[0111] Optionally, the method further comprises: a sixth processing module.
[0112] The sixth processing module is configured to extract core voice segments from the auxiliary prompt synthesized voice corresponding to the target driving scene of the intelligent vehicle, to obtain a plurality of core voice segments in the auxiliary prompt synthesized voice, and to perform voice reorganization according to the context relationship between the plurality of core voice segments, to obtain an auxiliary prompt core voice; and the auxiliary prompt core voice is backed up for voice broadcast when the intelligent vehicle moves to the next target driving scene.
[0113] The multi-voice-color emotion adaptive voice synthesis device can execute the method provided by any embodiment of the present application, has the corresponding function modules and beneficial effects of executing the method. Technical details not described in detail in the present embodiment can be referred to the multi-voice-color emotion adaptive voice synthesis method provided by any embodiment of the present application.
[0114] Figure 5 The structure schematic diagram of the electronic device provided by the embodiments of the present application is shown. Figure 5 A block diagram of an exemplary electronic device suitable for implementing the present embodiments is shown. Figure 5 The electronic device 12 shown is merely an example and should not limit the function and use range of the embodiments of the present application.
[0115] As shown in Figure 5 The electronic device 12 is in the form of a general computing device. The components of the electronic device 12 can include but are not limited to one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0116] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0117] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0118] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 5 Not shown; usually referred to as a "hard drive"). Although Figure 5 As not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this application.
[0119] A program / utility 40 having a set (at least one) of program modules may be stored, for example, in memory 28. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules typically perform the functions and / or methods described in the embodiments of this application.
[0120] The electronic device 12 can also be in communication with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; can also be in communication with one or more devices that enable a user to interact with the electronic device 12; and / or can be in communication with any devices (such as a network card, a modem or the like) that enable the electronic device 12 to communicate with one or more other computing devices. Such communication can be facilitated by an Input / Output (I / O) interface 22. Still yet, the electronic device 12 can be in communication with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or the Internet) through a network adapter 20. As an example, the network adapter 20 can be capable of Figure 5 Other hardware and / or software modules that can be used in conjunction with the electronic device 12, such as microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc. are not shown in FIG. 1 but can be incorporated into the electronic device 12.
[0121] The processing unit 16 performs various function applications and data processing by running programs stored in the system memory 28, such as implementing the multi-voice emotion adaptive speech synthesis method provided by the embodiments of the present application.
[0122] The embodiments of the present application also provide a computer storage medium.
[0123] The computer readable storage medium of the embodiments of the present application can adopt any combination of one or more computer readable mediums. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium, for example, can be, but is not limited to, an electrical, a magnetic, an optical, an electromagnetic, an infrared, or a semiconductor system, device or apparatus, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus or device.
[0124] A computer readable signal medium can include a propagated data signal with computer executable progra m code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal can take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable signal medium can be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport program for use by or in connection with an instruction execution system, apparatus, or device.
[0125] Program code embodied on a computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0126] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In an embodiment, electronic program guide data can be received from a remote computer that is connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0127] Embodiments of the present application also provide a computer program product.
[0128] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, in integrated circuitry, in one or more processors, in computer hardware, in firmware, in software, and / or in combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable on one or more programmable processing devices, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, one or more computer-readable storage media.
[0129] It is to be noted that the above-mentioned embodiments illustrate rather than limit the application, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the application. The word "comprising" does not exclude the presence of elements or steps other than those listed in a claim. In a claim, the word "a" or "an" preceding the commencement of the recitation of a list of elements or steps does not exclude the presence of more than one of such element or step. It is appreciated that features of the application that are, at this time, considered to be the most preferred embodiments can evolve. Therefore, the claims represent the extent of the inventor's current contribution to the art and are drafted to maintain this contribution in the face of future developments in this technology.
Claims
1. A method of multi-voice timbre emotion adaptive speech synthesis, characterized in that, The method comprises: acquiring driving assistance prompt text corresponding to the intelligent vehicle in a target driving scene; performing timbre coding on the driving assistance prompt text corresponding to the intelligent vehicle in the target driving scene to obtain an assistance prompt timbre vector corresponding to the driving assistance prompt text, and generating assistance prompt timbre data corresponding to the intelligent vehicle in the target driving scene based on the assistance prompt timbre vector; establishing a sentiment prosody mapping model of sentiment state to acoustic parameters, and performing sentiment prosody control on the assistance prompt timbre data through the sentiment prosody mapping model and the target driving scene to obtain sentiment expression timbre data with sentiment prosody; determining vehicle driving noise data based on a current driving state of the intelligent vehicle in the target driving scene, and optimizing the sentiment expression timbre data according to the vehicle driving noise data to obtain noise-optimized timbre data; performing acoustic characteristic compensation on the noise-optimized timbre data based on vehicle attribute features of the intelligent vehicle, and performing spatial acoustic optimization on the noise-optimized timbre data according to acoustic layout features of the intelligent vehicle to obtain acoustic-optimized timbre data; performing context sentiment expression adjustment on the acoustic-optimized timbre data based on the text content of the driving assistance prompt text to generate assistance prompt synthesized speech corresponding to the target driving scene of the intelligent vehicle.
2. The method of claim 1, wherein, The generation of the assistance prompt timbre data corresponding to the intelligent vehicle in the target driving scene based on the assistance prompt timbre vector comprises: inputting the assistance prompt timbre vector into a preset network model, determining initial timbre data based on the output of the preset network model; splitting the initial timbre data into a plurality of controllable timbre attributes, and configuring attribute parameters for the plurality of controllable timbre attributes to obtain target timbre data with different timbre attributes; limiting the target timbre data based on the scene category to which the target driving scene belongs to timbre condition parameters, and mixing the target timbre data after the timbre condition parameter limitation to obtain the assistance prompt timbre data corresponding to the intelligent vehicle in the target driving scene.
3. The method of claim 1, wherein, The acoustic characteristic compensation on the noise-optimized timbre data based on the vehicle attribute features of the intelligent vehicle comprises: acquiring an acoustic characteristic vector corresponding to the vehicle attribute features of the intelligent vehicle, the acoustic characteristic vector being used to describe a plurality of acoustic characteristic dimensions, the plurality of acoustic characteristic dimensions comprising: spatial volume, reflectivity, and sound-absorbing material distribution; establishing a mapping conversion model of acoustic characteristics to sound distortion, and determining a sound distortion degree corresponding to the acoustic characteristic vector through the mapping conversion model; performing acoustic characteristic compensation on the noise-optimized timbre data based on the sound distortion degree corresponding to the acoustic characteristic vector through a characteristic compensation function.
4. The method of claim 3, wherein, The spatial acoustic optimization on the noise-optimized timbre data according to the acoustic layout features of the intelligent vehicle to obtain acoustic-optimized timbre data comprises: The sound color data at the driving position in the intelligent vehicle in the noise-optimized sound color data is enhanced in intelligibility based on a pre-compensation filter, so as to optimize the sound color orientation at the driving position in the intelligent vehicle. The sound color data at the non-driving position in the intelligent vehicle in the noise-optimized sound color data is reduced in intelligibility, so as to optimize the sound color orientation at the non-driving position in the intelligent vehicle.
5. The method of claim 1, wherein, The method further comprises: detecting the voice data volume corresponding to the auxiliary prompt synthesized voice; if the voice data volume corresponding to the auxiliary prompt synthesized voice exceeds a preset volume threshold, splitting the auxiliary prompt synthesized voice based on the driving assistance prompt text to obtain a plurality of auxiliary prompt voice segments; generating the voice response intensity corresponding to each auxiliary prompt voice segment based on the voice context logic between the plurality of auxiliary prompt voice segments, and adjusting the response intensity of the auxiliary prompt synthesized voice based on the voice response intensity corresponding to each auxiliary prompt voice segment.
6. The method of claim 1, wherein, The method further comprises: playing the auxiliary prompt synthesized voice corresponding to the target driving scene in the intelligent vehicle; if the warning prompt level of the auxiliary prompt synthesized voice is higher than a preset prompt level, and the current driving state of the intelligent vehicle does not match the expected driving state corresponding to the auxiliary prompt synthesized voice, extracting the scene key voice in the auxiliary prompt synthesized voice; enhancing the voice intensity of the scene key voice in the auxiliary prompt synthesized voice, and adjusting the voice intensity of the auxiliary prompt synthesized voice based on the scene key voice after voice intensity enhancement; continuously playing the auxiliary prompt synthesized voice in the intelligent vehicle.
7. The method of claim 1, wherein, The method further comprises: extracting the core voice segment of the auxiliary prompt synthesized voice corresponding to the target driving scene of the intelligent vehicle to obtain a plurality of core voice segments in the auxiliary prompt synthesized voice, and recombining the voice according to the context relationship between the plurality of core voice segments to obtain an auxiliary prompt core voice; backing up the auxiliary prompt core voice for voice broadcast when the intelligent vehicle moves to the next target driving scene.
8. A multi-phonetic emotion-adaptive speech synthesis apparatus, characterized by, The device comprises: an acquisition module configured to acquire the driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle; a first processing module configured to encode the driving assistance prompt text corresponding to the target driving scene of the intelligent vehicle into sound color to obtain an auxiliary prompt sound color vector corresponding to the driving assistance prompt text, and generate auxiliary prompt sound color data corresponding to the target driving scene of the intelligent vehicle based on the auxiliary prompt sound color vector; a second processing module configured to establish a sentiment prosody mapping model of sentiment state to acoustic parameters, and control the sentiment prosody of the auxiliary prompt sound color data through the sentiment prosody mapping model and the target driving scene to obtain sentiment expression sound color data with sentiment prosody; The third processing module is configured to determine vehicle driving noise data based on a current driving state of the intelligent vehicle in the target driving scene, and optimize the emotional expression timbre data according to the vehicle driving noise data to obtain noise-optimized timbre data. The compensation and optimization module is configured to perform acoustic characteristic compensation on the noise-optimized timbre data based on vehicle attribute features of the intelligent vehicle, and perform spatial acoustic optimization on the noise-optimized timbre data based on acoustic layout features of the intelligent vehicle, to obtain acoustic-optimized timbre data. The adjustment module is configured to perform context emotional expression adjustment on the acoustic-optimized timbre data based on text content of the driving assistance prompt text, to generate the auxiliary prompt synthesized voice of the intelligent vehicle corresponding to the target driving scene.
9. An electronic device, comprising: Comprise: One or more processors; Memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-timbre emotional self-adaptive speech synthesis method of any one of claims 1 to 7.
10. A storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the multi-timbre emotional self-adaptive speech synthesis method of any one of claims 1 to 7.
Citation Information
Patent Citations
Speech playing method and device, and electronic equipment
CN111627417A
Speech synthesis method and device and device for speech synthesis
CN113409765A