Emotional voice broadcasting method, device and computer-readable storage medium
By generating personalized voiceprint models and combining them with deep learning models to extract timbre and emotional features, the problem of inconsistent emotional expression with users' real habits in speech synthesis systems has been solved, enabling personalized and emotional broadcasting on TV, and improving the voice interaction experience and device compatibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO HAIER MULTI MEDIA CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-02
Smart Images

Figure CN122135723A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, such as a method, apparatus and computer-readable storage medium for emotional speech broadcasting. Background Technology
[0002] With the rapid development of artificial intelligence technology, speech synthesis technology has been widely applied in consumer electronics products such as smart TVs and smart speakers. Current mainstream speech synthesis systems can convert text into natural and fluent speech output, providing users with a convenient interactive experience. Especially in the field of smart TVs, voice broadcasting has become a standard feature, widely used in scenarios such as weather reports, schedule reminders, and news readings.
[0003] Currently, a method for voice replication and low-latency streaming speech synthesis based on ultra-short samples has been disclosed. This solution generates personalized timbre data from short audio samples uploaded by users and supports dynamic adjustment of speech rate, volume, and emotional parameters based on text content in real-time dialogues, achieving rapid voice replication and low-latency voice interaction. This solution can solve the problems of monotonous timbre and high customization costs in traditional speech synthesis, and improves the personalized experience of voice interaction to a certain extent.
[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art: While related technologies have achieved rapid voice replication and dynamic emotion adjustment, their emotional expression is primarily based on real-time dialogue context calculations, with emotional features derived from the machine's instantaneous analysis of the current interaction state. This dynamically generated emotional tone differs from the user's actual expression habits, resulting in synthesized speech that, while capable of presenting different emotional tendencies, struggles to accurately reproduce the unique intonation, rhythm, and expression of a specific speaker under different emotional states. Therefore, how to simultaneously replicate a user's timbre and emotional expression features based on a small number of speech samples, enabling synthesized speech to be broadcast in the voice of a specific user and their authentic emotional expression, has become a pressing technical problem that needs to be solved.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0007] This disclosure provides an emotional voice broadcasting method, apparatus, and computer-readable storage medium to simultaneously replicate a user's timbre and emotional expression characteristics based on a small number of voice samples, enabling synthesized voice to be broadcast in the voice of a specific user and their authentic emotional expression.
[0008] In some embodiments, the emotional voice broadcasting method includes: receiving multiple voice samples, which are recorded by a user for different preset emotional types, including greetings, reminders, and blessings; using a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters from the multiple voice samples to generate a personalized voiceprint model; wherein the personalized voiceprint model includes user timbre features and emotional features corresponding to the preset emotional types; storing the personalized voiceprint model in association with a user identifier; responding to a model synchronization request carrying a user identifier sent by the TV terminal, matching the corresponding personalized voiceprint model based on the user identifier, and synchronizing the matched personalized voiceprint model to the TV terminal so that the TV terminal can synthesize emotional broadcasting voice based on the text to be broadcast, the personalized voiceprint model, and preset emotional intensity adjustment parameters during broadcasting.
[0009] In some embodiments, the emotional voice broadcasting method includes: sending a model synchronization request carrying a user identifier to a cloud server; receiving and storing a personalized voiceprint model returned by the cloud server based on the user identifier and obtaining emotional intensity adjustment parameters; wherein the personalized voiceprint model is generated from multiple voice samples recorded by the user for different preset emotional types uploaded by the mobile terminal, including the user's timbre features and emotional features corresponding to the preset emotional types; in response to a broadcasting trigger event, obtaining the text to be broadcast; synthesizing emotional broadcasting voice through a speech synthesis engine based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters, and playing the emotional broadcasting voice.
[0010] In some embodiments, the emotional voice broadcasting method includes: responding to a user's voice timbre replication instruction, guiding the user on an interactive interface to record multiple voice samples for different preset emotion types, including greetings, reminders, and blessings; uploading the multiple voice samples to a cloud server, so that the cloud server can use a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters of the multiple voice samples, generating a personalized voiceprint model including the user's voice timbre features and the emotional features corresponding to the preset emotion types; receiving the emotional intensity adjustment parameters set by the user, and outputting the emotional intensity adjustment parameters, so that the TV terminal can synthesize emotional broadcasting voice based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters during broadcasting.
[0011] In some embodiments, the emotional voice broadcasting device includes a processor and a memory storing program instructions, the processor being configured to execute the aforementioned emotional voice broadcasting method when running the program instructions.
[0012] In some embodiments, the computer-readable storage medium stores program instructions that, when executed, cause the computer to perform the aforementioned emotional voice broadcasting method.
[0013] The emotional voice broadcasting method, apparatus, and computer-readable storage medium provided in this disclosure can achieve the following technical effects: This solution utilizes a deep learning model to extract both the user's timbre and personalized emotional features under different emotional states from multiple voice samples recorded by the user for various preset emotional types, such as greetings, reminders, and blessings. This generates a personalized voiceprint model encompassing timbre and multiple emotional dimensions. When broadcast on television, the synthesized speech is based on this voiceprint model and emotional intensity adjustment parameters. By pre-capturing the user's authentic intonation, rhythm, and expression under specific emotional types, this solution enables the synthesized speech to not only mimic the user's timbre but also their emotional expression habits. This resolves the technical issue of discrepancies between dynamically generated emotions and the user's authentic expression habits, significantly enhancing the realism and emotional resonance of the voice broadcast.
[0014] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description
[0015] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a schematic diagram of an emotional voice broadcasting system provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of an emotional voice broadcasting method provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of another emotional voice broadcasting method provided in this embodiment of the disclosure; Figure 4 This is a schematic diagram of a method for determining emotion intensity adjustment parameters provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of another emotional voice broadcasting method provided in this embodiment of the disclosure; Figure 6 This is a schematic diagram of an emotional voice broadcasting device provided in an embodiment of this disclosure. Detailed Implementation
[0016] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0017] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0018] Unless otherwise stated, the term "multiple" means two or more.
[0019] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0020] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0021] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.
[0022] Combination Figure 1 As shown, optionally, this disclosure provides an emotional voice broadcasting system, including a mobile terminal 10, a cloud server 20, and a television terminal 30. The mobile terminal 10 corresponds to the user side 401 and is used for a first user to replicate voice timbre and set emotional parameters; the television terminal 30 corresponds to the user side 402 and is used for a second user to receive emotional voice broadcasts. Here, the first user can be the same as or different from the second user.
[0023] Specifically, the mobile terminal 10 is configured to respond to the first user's voice replication command, guiding the first user to record multiple voice samples for different preset emotion types on the interactive interface. The preset emotion types include greetings, reminders, and blessings. The mobile terminal 10 is also used to receive emotion intensity adjustment parameters set by the first user through the interactive interface, upload multiple voice samples to the cloud server 20, and upload the association between the emotion intensity adjustment parameters and the subsequently generated personalized voiceprint model to the cloud server 20. The cloud server 20 is configured to receive the multiple voice samples uploaded by the mobile terminal 10, use a pre-trained deep learning model to extract voiceprint features and analyze emotion parameters from the multiple voice samples, and generate a personalized voiceprint model. The personalized voiceprint model includes the user's voice characteristics and the emotion features corresponding to the preset emotion types. The cloud server 20 also associates and stores the personalized voiceprint model with the first user's user identifier, and receives the association between the emotion intensity adjustment parameters and the personalized voiceprint model uploaded by the mobile terminal 10, as well as associates and stores the emotion intensity adjustment parameters with the personalized voiceprint model. The TV terminal 30 is configured to send a model synchronization request carrying a user identifier to the cloud server 20 in response to a second user's broadcast request or a system-triggered event. The cloud server 20 matches the corresponding personalized voiceprint model and associated emotional intensity adjustment parameters based on the user identifier, and synchronizes the matched personalized voiceprint model and emotional intensity adjustment parameters to the TV terminal 30. After receiving and storing the personalized voiceprint model and emotional intensity adjustment parameters, the TV terminal 30, in response to a broadcast trigger event, such as a second user's voice command or scene change, obtains the text to be broadcast, and synthesizes an emotional broadcast voice through a speech synthesis engine based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters. This emotional broadcast voice is then played through the audio output unit of the TV terminal 30 for the second user to listen to.
[0024] The emotional voice broadcasting system provided in this embodiment collects voice samples and emotional intensity adjustment parameters of a first user for different emotional types via a mobile terminal 10. A personalized voiceprint model containing timbre and emotional features is generated by a cloud server 20, and the emotional broadcasting voice is synthesized by a television terminal 30 based on this model and parameters. This achieves simultaneous replication and independent adjustment of timbre and emotional features, enabling the synthesized voice to both reproduce the user's authentic timbre and express emotions at the user's preset emotional intensity. This effectively solves the technical problem of inconsistencies between emotional expression and user's actual habits in related technologies, improving the voice interaction experience in home settings. Simultaneously, data synchronization between the mobile terminal 10 and the television terminal 30 is achieved through the cloud server 20, supporting multi-terminal adaptation, simplifying the user's personalization process, and lowering the technical threshold.
[0025] Combination Figure 2As shown, optionally, this disclosure provides an emotional voice broadcasting method applied to a cloud server, including: S11: The cloud server receives multiple voice samples, which are recorded by the user for different preset emotion types, including greetings, reminders, and blessings.
[0026] In S12, the cloud server uses a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters from multiple speech samples, generating a personalized voiceprint model. This personalized voiceprint model includes the user's timbre characteristics and emotional features corresponding to preset emotional types.
[0027] S13, the cloud server stores personalized voiceprint models associated with user identifiers.
[0028] S14, the cloud server responds to the model synchronization request carrying the user identifier sent by the TV terminal, matches the corresponding personalized voiceprint model based on the user identifier, and synchronizes the matched personalized voiceprint model to the TV terminal so that the TV terminal can synthesize emotional broadcast voice based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters when broadcasting.
[0029] In this solution, the cloud server receives multiple voice samples, which are recorded and uploaded by the user via a mobile terminal. Specifically, the user opens a mobile app that works with the TV. On the app interface, the system displays recording prompts corresponding to greeting, reminder, and blessing emotions, respectively, each with a sample text. For example, a greeting prompt might say, "Good morning," a reminder prompt might say, "Remember to bring an umbrella," and a blessing prompt might say, "Happy birthday." The user records three voice samples according to the prompts, each sample lasting 10-15 seconds to ensure complete capture of the user's vocal characteristics and emotional expression without overburdening the user. After recording, the mobile terminal uploads these three voice samples, each containing a different emotional type, to the cloud server via an encrypted API. The cloud server receives these voice samples and preliminarily associates them with the user's identification information, preparing for subsequent voiceprint feature extraction and emotional parameter analysis. Using the above method, the cloud server receives voice samples recorded by users for three preset emotional types: greetings, reminders, and blessings. This provides a data foundation for generating personalized voiceprint models that include the user's timbre and emotional features, thus solving the technical problem in related technologies where the lack of emotional features leads to a lack of intonation variation in synthesized speech.
[0030] In this scheme, the cloud server can preprocess each speech sample. Preprocessing includes three stages: signal digitization, pre-emphasis processing, and sound framing. Signal digitization converts analog sound waves into digital signals through sampling and quantization for subsequent computer processing. Pre-emphasis processing uses techniques such as spectral subtraction or adaptive filtering to improve speech quality, compensate for high-frequency attenuation during transmission, and make the speech spectrum flatter. Sound framing divides the continuous speech signal into short segments of 10-40 milliseconds for analysis, allowing for processing of signals that can be considered stable within a short timeframe. Next, the cloud server can input the preprocessed speech samples into a pre-trained self-supervised learning feature extractor to extract deep acoustic features. This self-supervised learning feature extractor uses an improved self-supervised learning pre-trained model, learning rich acoustic representations from the original audio waveform through a contrastive learning framework. Based on the wav2vec2.0 architecture, the feature extractor can effectively capture key features in speech. The extracted deep acoustic features include timbre features, which reflect the speaker's unique vocal characteristics. In this way, for each emotion type of speech sample, the corresponding emotional features are extracted from the deep acoustic features. Since different emotion types, such as greetings, reminders, and blessings, have different acoustic expressions in terms of tone, intonation, and rhythm, emotion parameter analysis can separate the feature components related to emotion expression from the deep acoustic features. These emotional features can recreate the user's true expression in a specific emotional state. Finally, the timbre features in the deep acoustic features are identified as the user's timbre features, and these user timbre features are associated with the emotional features of each emotion type to generate a personalized voiceprint model. This personalized voiceprint model includes both the user's inherent timbre information and the user's emotional features in different emotional states, achieving an organic fusion of timbre and emotional features. Using the above method, the cloud server generates a personalized voiceprint model containing user timbre features and multiple emotional features by preprocessing, self-supervised feature extraction, emotional feature extraction, and feature association of multiple emotion speech samples. This solves the technical problem of the lack of emotional features leading to a lack of intonation variation in synthesized speech in related technologies, providing an accurate speech foundation for subsequent emotionalized broadcasting on television.
[0031] In practical applications, such as in a family setting, the cloud server receives three voice samples uploaded by the mother via a mobile app. These samples correspond to greetings ("Good morning, baby"), reminders ("Remember to bring your water bottle today"), and blessings ("Happy Birthday"). The cloud server first digitizes the three voice samples, performs pre-emphasis processing, and performs 10-40 millisecond sound framing. Then, the pre-processed voice is input into a self-supervised learning feature extractor based on the wav2vec2.0 architecture. A contrastive learning framework is used to extract deep acoustic features containing the mother's vocal timbre. Next, for each of the three emotional types—greeting, reminder, and blessing—corresponding emotional features are extracted from the deep acoustic features, such as the rising intonation of greetings, the concerned tone of reminders, and the pleasant rhythm of blessings. Finally, the mother's vocal timbre is correlated and fused with these three emotional features to generate a complete personalized voiceprint model. This model will then be synchronized to the television to synthesize broadcast voices with the mother's vocal timbre and specific emotional nuances.
[0032] In this solution, after generating a personalized voiceprint model, the cloud server associates and stores the model with a user identifier. The user identifier is uniquely identifying a user; it can be a registered mobile phone number, account ID, or other identifier that distinguishes different users. By associating the personalized voiceprint model with the user identifier, the cloud server establishes a mapping relationship between the user and their voice model, laying the foundation for subsequent multi-terminal access to the same voice model. Furthermore, this associated storage method supports multi-terminal adaptation, allowing the same user to access the same voice model on different TV devices using the same user identifier, achieving a personalized broadcast experience across devices.
[0033] Furthermore, the cloud server also receives emotion intensity adjustment parameters sent by the mobile terminal. These parameters are set by the user on the mobile terminal through an interactive interface, such as adjusting the emotion intensity within a range of 0-100 using a slider control. The system can preview the broadcast effect in real time. The cloud server associates and stores the emotion intensity adjustment parameters with the personalized voiceprint model, making the emotion intensity parameters part of the personalized voiceprint model, ensuring that they can be sent to the TV during subsequent synchronization. Subsequently, the cloud server responds to the model synchronization request carrying the user identifier sent by the TV. In practical applications, the TV can automatically send a model synchronization request carrying the user identifier to the cloud server upon startup to detect and update the locally stored personalized voiceprint model. After receiving the request, the cloud server matches the corresponding personalized voiceprint model in the stored models based on the user identifier carried in the request, and synchronizes the matched personalized voiceprint model and its associated emotion intensity adjustment parameters to the TV via an encrypted API. After the cloud server synchronizes the personalized voiceprint model to the TV, the TV can synthesize emotional broadcast speech based on the text to be broadcast, the personalized voiceprint model, and the emotion intensity adjustment parameters during broadcast. The TV integrates a speech synthesis engine, supporting real-time access to synchronized voiceprint models and automatically matching emotional modes based on the content being played. Users can also manually switch emotional modes. The TV also supports voice command activation, allowing users to trigger personalized broadcasts using natural language commands such as a mother's voice to report the weather or a gentle voice to tell a story. Using this approach, the cloud server associates and stores personalized voiceprint models with user identifiers and responds to TV requests by matching and synchronizing models based on user identifiers. This achieves centralized management and multi-terminal distribution of user voice models, resolving the technical issue of poor data interoperability between TV systems and mobile devices, which leads to cumbersome personalization processes and improves the convenience of cross-device collaboration.
[0034] In practical applications, such as in a family setting, the cloud server generates a personalized voiceprint model that includes the mother's vocal characteristics and three emotional features: greetings, reminders, and blessings, and stores this model along with the mother's user identifier. The mother sets the emotional intensity to 80 points via a mobile app, and this parameter is also stored along with the model. When the child turns on the TV in the morning, the TV automatically sends a model synchronization request carrying the mother's user identifier to the cloud server. Upon receiving the request, the cloud server matches the mother's personalized voiceprint model and its associated emotional intensity parameter based on the user identifier, and synchronizes the model and parameters to the TV via an encrypted API. Subsequently, when the child tells the TV, "Mom reminded me to bring this today," the TV synthesizes a reminder message with the mother's voice and a caring tone based on the text to be played, the mother's personalized voiceprint model, and the preset emotional intensity of 80 points, and plays it, allowing the child to feel the mother's presence.
[0035] Optionally, the method further includes: The cloud server establishes a mapping relationship between the personalized voiceprint model and multiple TV device identifiers associated with the same user identifier.
[0036] In response to a model synchronization request carrying a user identifier sent by any TV terminal, the matching personalized voiceprint model is synchronized to the corresponding TV terminal.
[0037] In this solution, when a user owns multiple TV devices under the same account, such as a living room TV and a bedroom TV, the cloud server generates a personalized voiceprint model and associates and stores this model with the user's unique identifier and the device identifiers of the two TVs, establishing a one-to-many mapping relationship. Subsequently, when any TV sends a model synchronization request carrying the user's identifier, the cloud server matches the corresponding personalized voiceprint model based on the user identifier in the request, and synchronizes the matched personalized voiceprint model to the corresponding TV based on the device identifier carried in the request or according to the preset association relationship between the TV and the user identifier. Using this method, the cloud server, by establishing a mapping relationship between the user identifier and multiple device identifiers, enables the same user to call the same set of voice models on different TV devices, breaking down the data interoperability barriers between the TV system and mobile terminals, as well as between different TV terminals. This allows users to enjoy personalized and emotional broadcast services on all home TV devices with only one recording, avoiding the cumbersome process of repeated setup and significantly improving the user experience in multi-terminal scenarios.
[0038] Optionally, the method further includes: In response to the emotional broadcast event on the first TV station, the cloud server obtains the emotional intensity adjustment parameters currently used by the first TV station.
[0039] The cloud server synchronizes the emotional intensity adjustment parameters to the second TV terminal associated with the same user ID, so that the second TV terminal can adjust the emotional expression of subsequent broadcasts according to the synchronized emotional intensity adjustment parameters to maintain the consistency of emotional expression across devices.
[0040] In this solution, the cloud server also supports cross-device emotional state synchronization to maintain consistency in emotional expression across multiple terminals. Specifically, when an emotional broadcast event occurs on the first TV terminal, such as when a user triggers broadcasting by telling a story in a gentle mode via voice command, or when the TV terminal detects that the currently playing content has switched to a children's program scene and automatically matches a lively emotional mode, the first TV terminal will use the currently determined emotional intensity adjustment parameters when synthesizing the emotional broadcast voice. The cloud server responds to the emotional broadcast event on the first TV terminal by obtaining the emotional intensity adjustment parameters currently used by the first TV terminal. These parameters reflect the emotional expression intensity selected by the user on the current device or adapted by the system. Subsequently, the cloud server synchronizes these emotional intensity adjustment parameters to the second TV terminal associated with the same user identifier; for example, the living room TV synchronizes the emotional parameters to the bedroom TV. After receiving the synchronized emotional intensity adjustment parameters, the second TV terminal adjusts its emotional expression based on these parameters in subsequent broadcasts, ensuring that the emotional intensity of the broadcast on the second TV terminal is consistent with that of the first TV terminal. By using the above method, the cloud server synchronizes the emotional state across devices, enabling the same user to have a consistent emotional broadcasting experience on different TV devices. This avoids the tedious operation of repeatedly setting emotional parameters on multiple devices, while ensuring that TV devices in different locations in the home can broadcast with the same emotional intensity, thus enhancing the integrity and coherence of voice interaction.
[0041] Optionally, the cloud server stores the personalized voiceprint model in association with the user identifier, including: The cloud server encrypts the personalized voiceprint model, generating encrypted voiceprint model data.
[0042] The cloud server stores the encrypted voiceprint model data in association with the user identifier.
[0043] In response to the model synchronization request from the TV, the cloud server synchronizes the encrypted voiceprint model data to the TV, which then decrypts and uses it locally.
[0044] In this solution, the cloud server also encrypts the personalized voiceprint model to ensure user data security. Specifically, after generating the personalized voiceprint model, the cloud server uses an encryption algorithm to encrypt the model data, generating encrypted voiceprint model data, and then stores the encrypted voiceprint model data in association with the user identifier. When the TV sends a model synchronization request carrying the user identifier, the cloud server matches the corresponding encrypted voiceprint model data based on the user identifier and synchronizes the encrypted data to the TV. After receiving the encrypted voiceprint model data, the TV decrypts it locally using a decryption algorithm to restore the usable personalized voiceprint model, and stores it in a local secure area on the TV for subsequent broadcasting. By using the above method, the cloud server ensures the security of user voice and emotional data throughout the entire data link by encrypting the storage and transmission of personalized voiceprint models. Even if the data is intercepted during transmission or illegally accessed while stored in the cloud, the usable voiceprint model content cannot be obtained, effectively protecting the user's biometric privacy and meeting the security needs of sensitive personal data in a home setting.
[0045] Optionally, the method further includes: The cloud server constructs the user's emotional profile based on personalized voiceprint models and the user's emotional interaction records at different times and in different scenarios.
[0046] The cloud server stores emotional profiles in association with user identifiers.
[0047] In response to the request from the TV station, the cloud server synchronizes the emotional profile and personalized voiceprint model to the TV station, so that the TV station can adjust the emotional expression of the broadcast according to the emotional profile.
[0048] In this solution, the cloud server also constructs an emotional profile of the user based on a personalized voiceprint model and records of the user's emotional interactions at different times and in different scenarios. Specifically, during long-term service, the cloud server records emotional data when the user interacts with the TV, including information such as when and in what scenarios the user tends to use which emotional mode, the user's feedback to the emotional broadcast, and the user's preference for manually switching emotional modes. The cloud server comprehensively analyzes these emotional interaction records with the timbre and emotional features in the personalized voiceprint model to uncover the user's emotional expression patterns and preference trends, constructing an emotional profile that comprehensively reflects the user's emotional characteristics and habits. This emotional profile can include multi-dimensional information such as the user's emotional tendencies at different times, emotional preferences in different scenarios, and the range of emotional intensity they can accept. The cloud server stores the constructed emotional profile in association with the user's identifier, forming the user's emotional archive. When the TV sends a model synchronization request, the cloud server synchronizes the emotional profile and the personalized voiceprint model to the TV. During subsequent broadcasts, the television not only synthesizes voices with user timbre and emotional characteristics based on personalized voiceprint models, but also personalizes the emotional expression of the broadcast according to emotional profiles. For example, it uses a warmer tone during the evening when users are usually in a low mood, or automatically enhances emotional intensity in entertainment scenarios where users prefer a lively mode. Using this method, the cloud server constructs user emotional profiles and synchronizes them to the television, enabling the television to proactively and personally adjust emotional expression based on the user's historical emotional habits. This upgrades voice interaction from a one-time response to a continuous service with emotional memory, solving the technical problems of lacking emotional continuity and personalized adaptability in related technologies, and achieving truly emotional intelligent interaction.
[0049] Optionally, in S12, the cloud server uses a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters from multiple speech samples to generate a personalized voiceprint model, including: The cloud server preprocesses each speech sample, including signal digitization, pre-emphasis processing, and audio framing.
[0050] The cloud server inputs the pre-processed speech samples into a pre-trained self-supervised learning feature extractor to extract deep acoustic features, including timbre features.
[0051] For each emotion type of speech sample, the cloud server extracts the corresponding emotion features from deep acoustic features.
[0052] The cloud server identifies the timbre features from the deep acoustic features as the user's timbre features, and associates the user's timbre features with the emotional features of various emotion types to generate a personalized voiceprint model.
[0053] In this solution, the cloud server preprocesses each speech sample, including three key steps: signal digitization, pre-emphasis processing, and audio framing. Signal digitization converts continuous analog sound waves into discrete digital signals through sampling and quantization. The sampling process extracts analog signal values at fixed time intervals; the quantization process approximates the continuously changing amplitude values using a finite number of amplitude values; and the encoding process represents the quantized value of each sample using binary numbers, thus converting the analog speech signal into a computer-processable digital form. Pre-emphasis processing uses techniques such as spectral subtraction or adaptive filtering to improve speech quality. By compensating for high-frequency attenuation during transmission, it flattens the speech spectrum and effectively suppresses environmental noise interference, providing a cleaner speech signal for subsequent feature extraction. Understandably, speech signals have short-term stationary characteristics and can be considered stable signals for processing within a time range of 10-40 milliseconds. Therefore, audio framing divides the continuous speech signal into short segments of 10-40 milliseconds for analysis. By segmenting the speech into short frames and setting overlaps between adjacent frames, smooth transitions between speech frames are ensured, fully preserving the dynamic changes in the speech. By employing the aforementioned preprocessing method, the cloud server can effectively improve the quality of speech samples, providing a clean, stable, and analyzable speech data foundation for subsequent voiceprint feature extraction and emotional parameter analysis. This solves the problems of noise interference and signal distortion that may exist in the original recording. In practical applications, for example, when a mother records a "Good morning" greeting for her child via a mobile app, the cloud server receives the speech sample and first samples the analog signal at a sampling rate of 16kHz and converts it into a digital signal using 16-bit quantization encoding. Then, it uses spectral subtraction to filter out environmental background noise that may exist during the recording process and compensates for high-frequency signal attenuation. Finally, it segments the preprocessed speech into short frames of 25 milliseconds and sets a frame shift of 10 milliseconds, generating a series of continuous speech frames for subsequent analysis and processing by the self-supervised learning feature extractor.
[0054] Furthermore, the cloud server inputs the preprocessed speech samples into a pre-trained self-supervised learning feature extractor to extract deep acoustic features, including timbre features. Specifically, this self-supervised learning feature extractor is based on the wav2vec2.0 architecture and learns rich acoustic representations from the original audio waveform through a contrastive learning framework. In one example, the feature extractor processes the preprocessed speech samples using a feature encoder composed of multiple convolutional neural networks. This feature encoder contains seven convolutional layers, each employing group normalization and the GELU activation function. Through layer-by-layer convolutional operations, it captures local features at different time scales from the original waveform, such as short-term energy fluctuations, fundamental frequency variations, and formant structures—acoustic details. The latent speech representation output by the feature encoder is input into a context network composed of multiple Transformer modules. This network models the global dependencies of the entire speech segment through a self-attention mechanism, learning long-range contextual information in the speech signal. In the self-supervised pre-training phase, the model is trained through a contrastive learning task: partial time steps of the feature encoder output are masked, and then the context network predicts the true latent representation of the masked position based on the unmasked context information. Simultaneously, the latent representation is discretized using a quantization module as the objective of contrastive learning, forcing the model to learn discriminative features that can distinguish different speech segments. After self-supervised pre-training on large-scale unlabeled speech data, this feature extractor can extract high-quality deep acoustic features from the input speech samples. These features include timbre information representing the speaker's unique vocal characteristics, as well as prosodic features such as pitch and speech rate in the speech signal, providing a rich feature foundation for subsequent sentiment parameter analysis and personalized voiceprint model construction. Using this approach, the cloud server, through a self-supervised learning feature extractor based on the wav2vec2.0 architecture, can efficiently extract deep acoustic representations containing timbre features from ultra-short user-recorded speech samples. This solves the technical problem of traditional methods relying on large amounts of labeled data and manually designed features, improving feature extraction capabilities in low-sample scenarios. In practical applications, for example, a cloud server preprocesses a recording of a mother greeting her child in the morning and inputs it into a pre-trained feature extractor based on the wav2vec2.0 architecture. This feature extractor's seven-layer convolutional network first extracts local features from the waveform at 10-millisecond intervals per frame. Then, it uses the self-attention mechanism of a Transformer context network to capture the intonation and emotional nuances of the entire speech, ultimately outputting a 512-dimensional deep acoustic feature vector. This vector contains the mother's unique timbre and intonation features reflecting the emotional tone of the greeting, providing high-quality feature input for subsequent generation of personalized voiceprint models.
[0055] In an alternative approach, the cloud server can also employ a hybrid architecture based on convolutional neural networks and recurrent neural networks. This involves extracting local spectral features through stacked one-dimensional convolutional layers, capturing long-term dependencies in speech through a bidirectional long short-term memory network, and finally outputting a fixed-dimensional deep acoustic feature vector through a global pooling layer. Alternatively, a HuBERT model can be used, generating pseudo-labels through offline clustering and then performing mask prediction pre-training to learn rich speech representations. Both alternative approaches can extract deep acoustic features containing timbre, pitch, and speech rate from preprocessed speech samples, providing high-quality feature input for subsequent speaker recognition and sentiment parameter analysis, thus addressing the feature extraction needs under different technical approaches.
[0056] Furthermore, after obtaining deep acoustic features through a self-supervised learning feature extractor, the cloud server performs emotional parameter analysis on voice samples of three different emotional types recorded by the user: greetings, reminders, and blessings. The emotional feature extraction process is implemented through the emotional classification branch in the deep neural network. This branch analyzes intonation variation features based on the fundamental frequency curve in the deep acoustic features; the rising or falling trend of the fundamental frequency curve reflects intonation patterns in different emotional states such as questioning, affirmation, or excitement. It also analyzes volume variation features based on energy distribution; the dynamic range of energy distribution reflects the intensity changes in emotional expression, for example, blessings often have higher energy peaks. Finally, it analyzes speech rate and pause features based on rhythm patterns; the speed of speech and the duration of pauses differ significantly in different emotional states, for example, reminders typically have a moderate speech rate and clear pauses to emphasize key information. The cloud server then fuses and maps these acoustic parameters related to emotional expression to form the emotional feature vector corresponding to that emotional type. Taking greetings as an example, their emotional features might include a rising fundamental frequency curve, a moderately high energy distribution, and a relatively coherent rhythmic pattern. For reminders, their emotional features might include a stable fundamental frequency curve, a moderate energy distribution, and a rhythmic pattern with emphatic pauses. And for blessings, their emotional features might include a fluctuating fundamental frequency curve, a high energy distribution, and a brisk rhythmic pattern. In this way, the cloud server can extract emotional features from speech samples of different emotional types that accurately represent how users express themselves in specific emotional states. By extracting corresponding emotional features from speech samples of different emotional types, the cloud server achieves digital modeling of users' true emotional expressions, solving the problems of singular and mechanical emotional expression in related technologies, and providing a precise emotional parameter basis for the subsequent synthesis of broadcast speech with rich emotional color. In practical applications, for example, a cloud server extracts emotional features from a mother's greeting "Good morning, baby," analyzing it to obtain a rising fundamental frequency curve, a moderately high energy distribution, and a coherent rhythmic pattern, forming a greeting emotional feature vector. Analyzing a reminder message like "Remember to bring your water bottle today," it obtains a stable fundamental frequency curve, a moderate energy distribution, and a rhythmic pattern with emphatic pauses, forming a reminder emotional feature vector. Analyzing a birthday greeting message like "Happy Birthday," it obtains a fluctuating fundamental frequency curve, a high energy distribution, and a brisk rhythmic pattern, forming a birthday greeting emotional feature vector. These emotional feature vectors, together with the mother's vocal timbre characteristics, constitute a personalized voiceprint model, enabling the subsequently synthesized broadcast speech to realistically reproduce the mother's expressions under different emotional states.
[0057] In one optimized approach, for each speech sample of a particular emotion type, the corresponding emotion features for that emotion type are extracted from deep acoustic features, including: The cloud server uses an adversarial training mechanism to separate pure emotional features unrelated to timbre features from deep acoustic features, so that the extracted emotional features do not contain user timbre information.
[0058] In this scheme, the cloud server employs an adversarial training mechanism to separate pure emotional features unrelated to timbre from deep acoustic features, ensuring that the extracted emotional features do not contain user timbre information. Specifically, the cloud server constructs an adversarial training network consisting of a feature extractor, an emotion classifier, and a timbre discriminator. The feature extractor is responsible for extracting emotional features from deep acoustic features, the emotion classifier ensures that the extracted features can accurately identify the emotion type, and the timbre discriminator attempts to identify the speaker's timbre identity from the extracted emotional features. During training, the feature extractor aims to maximize the accuracy of emotion classification while minimizing the accuracy of timbre discrimination, i.e., attempting to extract features that accurately express emotion but cannot be identified by the timbre discriminator; while the timbre discriminator aims to identify the timbre identity from the emotional features as accurately as possible. Through this adversarial game, the feature extractor is forced to discard timbre-related information from the emotional features, retaining only pure acoustic features related to emotional expression, such as intonation patterns, dynamic energy distribution, and rhythmic features, which are unrelated to the speaker's identity. Ultimately, the pure emotional features obtained by the cloud server through adversarial training are independent of the user's timbre features, achieving effective decoupling between emotional and timbre features. Using this method, the cloud server separates pure emotional features unrelated to timbre through adversarial training, enabling emotional features to be expressed and controlled independently of the speaker's identity. This solves the technical problem in related technologies where the coupling of emotional and timbre features affects the accuracy of timbre reproduction during emotional adjustment. It achieves independent adjustment of timbre and emotion, allowing users to freely adjust the intensity of emotional expression while maintaining their original timbre, greatly improving the flexibility and naturalness of speech synthesis.
[0059] In this scheme, the cloud server identifies the timbre feature from the deep acoustic features as the user's timbre feature and associates it with the emotional features of various emotion types to generate a personalized voiceprint model. Specifically, the cloud server identifies the core acoustic component that uniquely represents the speaker's identity—the timbre feature—from the deep acoustic features obtained through a self-supervised learning feature extractor. This feature reflects the inherent attributes of the user's vocal organs, such as physiological structure, resonance characteristics, and pronunciation habits, and is a key identifier for distinguishing different speakers. After identifying this timbre feature as the user's timbre feature, the cloud server associates and fuses it with the emotional features previously extracted from speech samples of greetings, reminders, and blessings. This association is not a simple feature concatenation, but rather establishes a mapping relationship between the timbre feature and various emotional features through a feature fusion layer in a deep neural network, forming a unified feature representation space. During the fusion process, the cloud server maintains timbre features as the base information and associates emotional features of different emotion types as adjustable dimensions with the timbre features. This ensures that the final personalized voiceprint model includes both the user's inherent timbre information and the user's expressive features under different emotional states, achieving effective integration of timbre and emotional features. This personalized voiceprint model is stored in a structured form, including a fixed timbre feature base and multiple emotional feature components, providing a foundation for the subsequent synthesis of speech on the television screen to call upon the corresponding emotional features as needed. Using this method, the cloud server generates a personalized voiceprint model with both identity identification and emotional expression capabilities by associating and fusing the user's timbre features with the emotional features of various emotion types. This solves the problem that current voiceprint models can only represent timbre information and lack an emotional dimension, providing a complete speech feature foundation for truly emotional voice broadcasting on television.
[0060] In practical applications, the cloud server identifies the timbre features from the mother's deep acoustic characteristics as the user's timbre features. These timbre features encompass the mother's unique vocal qualities, resonance characteristics, and pronunciation habits. The cloud server then integrates these timbre features with the greeting emotion features extracted from greeting voices, the reminder emotion features extracted from reminder voices, and the blessing emotion features extracted from blessing voices to generate a complete personalized voiceprint model. Specifically, the greeting emotion features extracted from greeting voices include the mother's characteristic rising fundamental frequency curve, moderate to high energy distribution, and smooth rhythmic pattern when saying "Good morning, baby"; the reminder emotion features extracted from reminder voices include the mother's characteristic stable fundamental frequency curve, moderate energy distribution, and rhythmic pattern with emphatic pauses when saying "Remember to bring your water bottle today"; and the blessing emotion features extracted from blessing voices include the mother's characteristic fluctuating fundamental frequency curve, high energy distribution, and light and lively rhythmic pattern when saying "Happy Birthday." In this model, the mother's vocal timbre is stored as fixed base information, while three emotional features—greetings, reminders, and blessings—are mapped to the vocal timbre in adjustable component form. When the television station needs to synthesize a greeting voice with the mother's vocal timbre, it can directly combine the greeting emotional features and vocal timbre features from this model to generate a broadcast voice that has both the mother's unique vocal timbre and is full of greeting emotion. When it needs to synthesize a reminder voice, it combines the reminder emotional features and vocal timbre features to generate a broadcast voice with the mother's vocal timbre and a caring tone.
[0061] Optionally, the cloud server receives the emotion intensity adjustment parameters sent by the mobile terminal, which are set by the user on the mobile terminal through an interactive interface.
[0062] The cloud server stores the emotion intensity adjustment parameters in association with the personalized voiceprint model.
[0063] In response to a model synchronization request from the TV, the cloud server synchronizes the emotional intensity adjustment parameters and the personalized voiceprint model to the TV.
[0064] In this solution, the cloud server receives emotion intensity adjustment parameters sent by the mobile terminal, which are set by the user on the mobile terminal through an interactive interface. Specifically, after the user completes the recording of multiple emotion type voice samples in the mobile terminal's mini-program interface, the system provides an emotion intensity adjustment function. The user adjusts the emotion intensity within a range of 0 to 100 points using a slider control. For example, setting the emotion level for a greeting to 80 points to express warmth, and setting the emotion level for a reminder to 60 points to express concern, the system can preview the playback effect in real time for the user's reference. After the user completes the emotion intensity setting, the mobile terminal uploads the emotion intensity adjustment parameters to the cloud server. After receiving the emotion intensity adjustment parameters, the cloud server associates and stores them with the previously generated personalized voiceprint model, making the emotion intensity parameters part of the personalized voiceprint model, ensuring that they can be sent to the TV during subsequent synchronization. When the TV sends a model synchronization request carrying a user identifier, the cloud server, in response to the request, synchronizes the emotion intensity adjustment parameters and the personalized voiceprint model to the TV so that the TV can synthesize emotional broadcast speech based on the text to be broadcast, the personalized voiceprint model, and the emotion intensity adjustment parameters during broadcast.
[0065] In another alternative implementation, the emotion intensity adjustment parameters can be sent directly from the mobile terminal to the TV instead of being sent to the cloud server for association processing. Specifically, after the user sets the emotion intensity adjustment parameters on the mobile terminal, the mobile terminal sends the parameters directly to the TV via a local network direct connection, such as Wi-Fi Direct or Bluetooth communication within the same local area network. Upon receiving the parameters, the TV associates and stores them locally with the personalized voiceprint model previously synchronized from the cloud server. When the TV triggers a broadcast event, it synthesizes an emotional broadcast voice based on the locally stored personalized voiceprint model and the emotion intensity adjustment parameters. This method is suitable for scenarios where the user only wants to apply the emotion settings on the currently used TV device, or where unstable network conditions cause cloud synchronization delays. Both methods can coexist; users can choose the transmission path for the emotion intensity adjustment parameters according to their actual needs. The cloud server provides parameter storage and synchronization services to support consistency across multiple terminals, while the direct connection method offers a faster and more private local setting experience.
[0066] By adopting the above approach, the cloud server receives and associates the stored emotion intensity adjustment parameters, and synchronizes the parameters with the model when responding to requests from the TV terminal, thereby realizing cloud-based management and multi-terminal distribution of user emotion preferences. At the same time, by providing an alternative solution for direct connection to the TV terminal, it meets users' needs for quick local settings and data privacy in specific scenarios, solves the technical problems of cumbersome emotion setting process, inconvenient multi-terminal synchronization, and user concerns about the privacy of emotion data in related technologies, and improves the flexibility and user-friendliness of the emotion-based voice broadcasting system.
[0067] In one optimized approach, the method further includes: The cloud server receives broadcast templates uploaded by mobile terminals, and the broadcast templates are associated with personalized voiceprint models.
[0068] The cloud server associates and stores the broadcast template with the personalized voiceprint model.
[0069] When responding to a model synchronization request from the TV, the cloud server synchronizes the broadcast template and the personalized voiceprint model to the TV so that the TV can call the broadcast template for broadcasting in the corresponding application scenario.
[0070] In this solution, the mobile terminal's mini-program interface provides a multi-scenario broadcast template selection function. Users can choose or customize broadcast templates according to their needs, such as weather broadcast templates, schedule reminder templates, and educational content explanation templates. Each template includes the broadcast text structure, emotional mode preference, and broadcast style settings for a specific scenario. The mobile terminal uploads the broadcast template selected by the user to the cloud server and establishes an association between the template and the personalized voiceprint model. After receiving the broadcast template, the cloud server associates and stores it with the personalized voiceprint model, making the broadcast template part of the user's personalized configuration. When the TV sends a model synchronization request carrying the user's identifier, the cloud server, in response to the request, synchronizes the broadcast template, personalized voiceprint model, and emotional intensity adjustment parameters to the TV. After receiving and storing this data, the TV automatically calls the corresponding broadcast template for broadcasting in the corresponding application scenario. For example, when a weather broadcast event is triggered, the weather broadcast template is called, and the personalized voiceprint model and emotional intensity adjustment parameters are combined to synthesize a broadcast voice that meets the scenario requirements. Users can also manually switch to other templates. By adopting the above method, the cloud server receives, stores, and synchronizes broadcast templates, realizing complete cloud management of user broadcast preferences. This enables the TV to automatically adapt the broadcast templates according to the application scenario, further simplifying the user operation process and improving the consistency and personalized experience of broadcasts in multiple scenarios.
[0071] Combination Figure 3 As shown, optionally, this disclosure provides another method for emotional voice broadcasting, applied to a television terminal, the method including: S21, the TV sends a model synchronization request carrying the user's identifier to the cloud server.
[0072] S22, the TV receives and stores the personalized voiceprint model returned by the cloud server based on user identifier matching, as well as the emotional intensity adjustment parameters. The personalized voiceprint model is generated from multiple voice samples recorded by the user for different preset emotional types, uploaded by the user from the mobile terminal, including the user's timbre characteristics and the emotional characteristics corresponding to the preset emotional types.
[0073] S23, the TV responds to the broadcast trigger event and obtains the text to be broadcast.
[0074] S24: The TV terminal synthesizes emotional voice messages based on the text to be read, a personalized voiceprint model, and emotional intensity adjustment parameters through a speech synthesis engine, and then plays the emotional voice messages.
[0075] In this solution, the TV automatically sends a model synchronization request to the cloud server upon startup to detect and update the locally stored personalized voiceprint model. The cloud server matches the corresponding personalized voiceprint model based on the user identifier carried in the request and synchronizes the model to the TV. The TV receives and stores the personalized voiceprint model returned by the cloud server, and simultaneously obtains the emotion intensity adjustment parameters associated with the model. The personalized voiceprint model is generated from multiple voice samples recorded by the user for different preset emotion types such as greetings, reminders, and blessings, uploaded by the user from the mobile terminal. This model includes the user's timbre characteristics and the emotional characteristics corresponding to each preset emotion type.
[0076] Furthermore, the TV responds to the broadcast trigger event and obtains the text to be broadcast. The broadcast trigger event can be that the TV receives a voice command from the user, such as a user triggering a personalized broadcast by using their mother's voice to announce the weather; or it can be that the TV detects that the currently playing content has switched to a preset application scenario, such as entering a weather broadcast scenario, a schedule reminder scenario, or an educational content explanation scenario.
[0077] Furthermore, the TV terminal synthesizes and plays emotionally-influenced speech based on the text to be read, a personalized voiceprint model, and emotion intensity adjustment parameters using a speech synthesis engine. During the synthesis process, the TV terminal inputs the text to be read into the speech synthesis engine. The engine determines the user's basic timbre, intonation, and speech rate characteristics based on the personalized voiceprint model. Then, based on these basic features, it dynamically adjusts the fundamental frequency, energy, and duration parameters of the synthesized speech according to the emotion intensity adjustment parameters. This superimposes the emotional color corresponding to the emotion intensity adjustment parameters onto the user's basic timbre, generating the final emotionally-influenced speech. Using this method, the TV terminal requests and obtains the personalized voiceprint model from the cloud server, retrieves the text to be read and emotion intensity adjustment parameters based on the broadcast trigger event, synthesizes and plays the emotionally-influenced speech, achieving personalized broadcasting based on the user's real timbre and preset emotion intensity on the TV terminal, thus improving the voice interaction experience in home scenarios.
[0078] Optionally, the method further includes: The TV receives manual switching commands input by the user, which are used to specify the target emotional mode.
[0079] The TV switches the current emotional mode to the target emotional mode based on the manual switching command, and determines the emotional intensity adjustment parameters based on the target emotional mode.
[0080] In this solution, the TV provides an emotion mode selection menu during playback or on the standby screen. Users can manually switch emotions via remote control, voice commands, or virtual controls on the TV interface, specifying the desired emotion mode, such as a warm mode, a lively mode, a formal mode, or selecting from a variety of preset emotion templates. Upon receiving the manual switching command, the TV switches the current broadcast emotion mode to the target emotion mode specified in the command and determines the corresponding emotion intensity adjustment parameters based on that target emotion mode. Different emotion modes correspond to different ranges and adjustment benchmarks for emotion intensity adjustment parameters; for example, a warm mode might correspond to a medium-to-high emotion intensity parameter, a lively mode to a high emotion intensity parameter, and a formal mode to a moderate emotion intensity parameter. After determining the emotion intensity adjustment parameters, the TV synthesizes emotional broadcast voice based on these parameters in subsequent broadcasts. Using this method, the TV, by receiving manual switching commands from users, allows users to autonomously select emotion modes according to personal preferences or specific scenario needs, further enhancing user control over emotional expression and the flexibility of the interactive experience, making emotional broadcasts more tailored to users' real-time needs.
[0081] In one optimized approach, the method further includes: The TV terminal obtains the user's historical interaction records, which include historical emotional states and historical broadcast responses.
[0082] The TV app uses historical interaction records to construct a trajectory of changes in the user's emotional state.
[0083] The TV terminal verifies whether the currently determined emotional intensity adjustment parameters conform to the user's emotional evolution pattern based on the trajectory of emotional state changes.
[0084] If the conditions are not met, the television terminal will smooth out the emotional intensity adjustment parameters.
[0085] In this solution, the TV client acquires users' historical interaction records during long-term service. These records include users' historical emotional states at different times and in different scenarios, such as a preference for mild emotions in the morning and a preference for lively emotions in the evening, as well as users' responses to historical broadcasts, such as feedback ratings for certain emotional broadcasts and whether they manually switched emotional modes. Based on these historical interaction records, the TV client constructs a trajectory of users' emotional state changes. This trajectory reflects the patterns and trends of users' emotional preferences changing over time and in different scenarios, such as emotional fluctuation curves at different times of the day and changes in emotional tendencies on different holidays or special dates. When the TV client determines the emotional intensity adjustment parameters based on the current broadcast trigger event, it compares these parameters with the user's emotional state change trajectory to verify whether the current parameters conform to the user's emotional evolution patterns. For example, if the user typically prefers mild emotions at the current time, and the emotional intensity corresponding to the current parameters is too high, it may not conform to the user's patterns. If the verification reveals that the current emotional intensity adjustment parameter does not match the user's emotional evolution pattern, the TV will smooth the parameter. For example, it can use weighted averaging or linear interpolation to adjust the current parameter towards one that aligns with the user's historical patterns. This ensures that the final parameter meets the needs of the current scenario while maintaining consistency with the user's historical emotional habits. By constructing a trajectory of the user's emotional state changes and verifying and smoothing the emotional intensity adjustment parameter, the TV ensures that emotional expression follows the user's emotional evolution pattern. This avoids abrupt changes in emotional expression caused by single-scene adaptation or manual settings, achieving a more natural and consistent emotional interaction experience. Users feel that the TV understands and respects their emotional habits. In practical applications, for example, if the TV discovers from historical records that a mother typically prefers a gentle and relaxed emotional mode after returning home from get off work in the evening, and one evening a child requests to have a story read in the mother's voice, the system automatically adapts to a lively emotional mode based on the children's program scenario. The TV program compared the currently determined lively emotional parameters with the mother's emotional state change trajectory and found that the parameters did not match the mother's emotional patterns in the evening. Therefore, the emotional intensity parameters were smoothed and the lively parameters were appropriately adjusted towards a gentler direction. Finally, a story broadcast voice that has the mother's voice tone and moderate liveliness but does not lose its gentleness was synthesized. This satisfies the child's entertainment needs and conforms to the mother's current emotional habits, so that all family members can have a comfortable experience.
[0086] In one optimized approach, the method further includes: After playing the emotionally charged audio message, the TV collects user feedback information, which includes at least one of the user's voice response, facial expressions, or physiological signals.
[0087] The TV terminal performs sentiment analysis on the feedback information to determine the user's emotional acceptance of the broadcast voice.
[0088] The TV terminal stores the parameters related to emotional acceptance and emotional intensity adjustment, which serve as the basis for optimizing secondary emotional adjustment.
[0089] In this solution, the TV terminal collects user voice responses via built-in or external microphones, such as evaluative voices like "I like this tone" or "It's too exaggerated" after listening to the broadcast; it also collects user facial expressions via a camera, such as smiles, frowns, or surprise; and it collects physiological signals from the user via wearable devices or sensors, such as heart rate changes and skin conductance, reflecting physiological fluctuations. The TV terminal performs sentiment analysis on this feedback information, using voice emotion recognition technology to analyze the emotional tendency in the user's voice responses, facial expression recognition technology to analyze the type and intensity of emotions in the user's facial expressions, and physiological signal analysis technology to interpret the user's emotional state. By integrating multiple factors, the TV terminal determines the user's emotional acceptance of the broadcast voice, which can be quantified as a numerical score from 0 to 100 or rated as excellent, good, average, or poor. The TV terminal associates and stores the emotional acceptance with the emotional intensity adjustment parameters corresponding to the currently used broadcast voice, forming a historical record of the emotional adjustment effect. When subsequent emotionally-oriented broadcasts are needed for this user, the TV can optimize and adjust the emotional intensity adjustment parameters based on historically stored emotional receptivity data. For example, if a certain emotional intensity parameter repeatedly receives low receptivity scores, its weight will be automatically reduced in subsequent broadcasts, or it will be adjusted towards a similar but more receptive parameter. Using this method, the TV collects user feedback and performs emotional analysis, using user emotional receptivity as the basis for optimization. This gives the emotional broadcast system closed-loop learning and continuous optimization capabilities, allowing the system to constantly adapt to changes in user emotional preferences, providing an increasingly personalized broadcast experience, and enhancing the depth of emotional interaction between the user and the TV.
[0090] In one optimized approach, the method further includes: Based on historical emotional acceptance data, a mapping model is constructed between emotional intensity adjustment parameters and user emotional acceptance on the TV platform.
[0091] The television platform uses a mapping model to pre-optimize the emotional intensity modulation parameters to be used, in order to predict and select the emotional intensity modulation parameters that are most acceptable to users.
[0092] In this solution, the TV terminal has accumulated a large amount of historical data over long-term service, including the emotional intensity adjustment parameters used for each broadcast and user emotional acceptance ratings for that broadcast. The TV terminal uses this data as training samples and employs machine learning algorithms to build a mapping model. This model can learn the non-linear relationship between different emotional intensity parameters and user acceptance; for example, some users show the highest acceptance for moderately high emotional intensity, while exhibiting low acceptance for excessively high or low emotional intensity. When the TV terminal needs to determine the emotional intensity adjustment parameters for an upcoming broadcast, it inputs candidate parameters into the mapping model for prediction. The model outputs the expected user acceptance for each candidate parameter, and the TV terminal selects the emotional intensity adjustment parameter with the highest predicted acceptance for the current broadcast. As historical data accumulates, the mapping model is continuously updated and optimized, thus improving prediction accuracy. Using the above method, the TV terminal pre-optimizes the emotional intensity adjustment parameters by constructing a mapping model, realizing an upgrade from passive response to active prediction. This enables the system to predict and select the emotional intensity that best matches the user's preferences based on historical experience, giving the emotional broadcasting system the ability to learn and continuously evolve. As the usage time increases, the emotional expression of the broadcast becomes more and more in line with the user's true preferences, providing the user with a personalized experience that becomes more and more considerate the longer it is used.
[0093] In one optimized approach, the television terminal synthesizes emotionally-influenced speech based on the text to be broadcast, a personalized voiceprint model, and preset emotional intensity adjustment parameters, including: The TV will store the personalized voiceprint model and emotion intensity adjustment parameters in a local secure area on the TV.
[0094] The emotional broadcast voice is synthesized locally on the TV, and the user's voice data and emotional data are not uploaded to the cloud server.
[0095] In this solution, the personalized voiceprint model and emotional intensity adjustment parameters, synchronously obtained from the cloud server, are stored in a local secure area on the TV. This local secure area can be a built-in security encryption chip, a trusted execution environment, or an encrypted storage partition, ensuring that the stored voiceprint model and emotional parameters cannot be accessed by unauthorized applications or users. When a broadcast event is triggered, the TV directly completes the entire process of synthesizing the emotional broadcast speech locally. All processing steps, including text input, voiceprint model invocation, emotional parameter application, fundamental frequency energy duration adjustment, and speech waveform generation, are executed locally on the TV. Throughout the process, sensitive data such as the user's timbre characteristics, emotional characteristics, and emotional intensity adjustment parameters are not uploaded to the cloud server. By storing the personalized voiceprint model and emotional intensity adjustment parameters in a local secure area and completing speech synthesis locally, the TV achieves end-to-end processing of user biometric data and emotional preference data. This avoids the risk of leakage of sensitive data during network transmission and cloud storage, allowing privacy-sensitive users to confidently use the emotional broadcast function. It also meets the security needs of sensitive personal data in a home setting, improving system compliance and user trust.
[0096] Optionally, the broadcast trigger event includes the TV receiving a voice command from the user, and the method further includes: The TV terminal performs semantic recognition on voice commands and extracts identification information used to specify the broadcast tone. The TV terminal retrieves the corresponding personalized voiceprint model from local storage based on the identity information.
[0097] In this solution, users issue natural language commands via the microphone or voice input function of the TV's remote control. Upon receiving the voice command, the TV first converts the speech into text using automatic speech recognition technology. Then, it performs semantic understanding and keyword extraction on the text to identify the identity information used to specify the broadcast voice tone. Based on the extracted identity information, the TV searches its locally stored family member voice database for a matching personalized voiceprint model. This database pre-stores multiple family member voiceprint models recorded via mobile terminals and synchronized to the TV, each model associated with its corresponding identity information. When a matching personalized voiceprint model is found, the TV uses that model for subsequent speech synthesis. If no matching model is found, the TV can prompt the user that the member's voice has not yet been recorded or guide the user to record it. Using this method, the TV extracts identity information from the user's voice command through semantic recognition and calls the corresponding voiceprint model, enabling personalized broadcasting with a specified voice tone triggered by natural language commands. This improves the convenience and naturalness of voice interaction, allowing family members to enjoy personalized broadcasting services in the most intuitive way.
[0098] Optionally, the voice command may also include emotion identification information for specifying an emotion pattern; the method may further include: The TV terminal performs semantic recognition on voice commands and extracts emotional marker information.
[0099] The television terminal determines the corresponding emotional pattern based on the emotional identification information, and then determines the emotional intensity adjustment parameters based on the emotional pattern.
[0100] In this solution, users specify both the voice tone and emotional mode via natural language commands. Upon receiving the voice command, the TV extracts the identity information for specifying the voice tone and the emotional mode using semantic recognition technology. The TV then matches the corresponding emotional mode from a pre-defined emotional mode library based on the emotional mode information. After determining the emotional mode, the TV further determines the corresponding emotional intensity adjustment parameters. Different emotional modes have different preset emotional intensity ranges and adjustment benchmarks; for example, a warm mode corresponds to a medium-to-high emotional intensity parameter, while a lively mode corresponds to a high emotional intensity parameter. The TV then invokes the corresponding personalized voiceprint model based on the determined identity information and combines it with the determined emotional intensity adjustment parameters and the text to be broadcast to synthesize an emotionally-driven voice message. Using this method, the TV enables users to specify both the voice tone and emotional mode simultaneously through a single natural language command by recognizing the emotional indicator information in the voice command, achieving more accurate and convenient emotionally-driven broadcast control.
[0101] Combination Figure 4 As shown, the broadcast trigger event includes the TV detecting that the currently playing content has switched to a preset application scenario, and the method also includes: S31, the TV terminal identifies the current application scenario type, which includes weather broadcast scenario, schedule reminder scenario, or educational content explanation scenario.
[0102] S32, the TV automatically matches the preset emotional mode according to the type of application scenario.
[0103] S33: The TV terminal determines the emotional intensity adjustment parameters based on the emotional pattern.
[0104] In this solution, the TV monitors the currently playing content in real time. When it detects a switch to a preset application scenario, such as a weather report, a schedule reminder, or an educational content explanation, the TV first identifies the current application scenario type. Then, it automatically matches a preset emotional mode based on the identified scenario type. For example, a weather report scenario might match a formal or warm mode, a schedule reminder scenario might match a concerned mode, and an educational content explanation scenario might match a lively mode. During the emotional mode matching process, the TV also performs text sentiment analysis on the text to be broadcast, extracting the emotional tendency of the text content. For example, a weather forecast text like "Today is sunny and suitable for outdoor activities" presents a positive emotional tendency, while a rainstorm warning, "Please be careful," presents a negative emotional tendency. The TV integrates the emotional tendency of the text content with the application scenario type to determine a target emotional mode that is dually compatible with both the text content and the scenario. For example, in a weather report scenario, if the text's emotional tendency is positive, a cheerful mode is selected; if the text's emotional tendency is negative, a concerned or serious mode is selected. Furthermore, the TV also records the user's manual switching history of emotional modes in different application scenarios. For example, a user might manually switch to immersive mode multiple times while watching a movie, or switch to formal mode during news broadcasts. Based on this manual switching history, the TV uses machine learning algorithms to learn the user's emotional preferences and generate personalized scene-emotion mapping relationships. When entering the same or similar application scenario again, the TV automatically matches the emotional mode that matches the user's preferences based on this personalized scene-emotion mapping relationship, rather than simply using the system's preset general mode. After determining the emotional mode, the TV determines the corresponding emotional intensity adjustment parameters based on that emotional mode. Different emotional modes have different preset emotional intensity value ranges and adjustment benchmarks; for example, the joy mode corresponds to a higher emotional intensity parameter, while the concern mode corresponds to a medium-to-high emotional intensity parameter. By adopting the above method, the TV achieves more accurate and personalized emotional mode matching by integrating text sentiment analysis and user historical preferences for scene emotional adaptation. This allows emotional broadcasts to truly understand the emotional connotation of the content and respect the user's personalized habits, improving the emotional authenticity of the broadcasts and user satisfaction.
[0105] In practical applications, for example, when the TV detects that the current playback content has switched to a weather report scene, and the text to be broadcast is "The first snow of the year will arrive tomorrow; please take precautions against the cold," the TV first performs sentiment analysis on the text, extracting the emotional tendency of concern and reminder. Simultaneously, the TV checks the user's historical manual switching records and finds that the user has manually switched to a "warm and caring" mode multiple times during such cold weather broadcasts. The TV integrates the text's sentiment tendency, the weather scene type, and the user's preferences, ultimately determining the target emotional mode as a "warm and caring" mode, and accordingly setting an emotional intensity adjustment parameter of 75 points. Subsequently, based on the mother's personalized voiceprint model and the 75-point emotional intensity parameter, the TV synthesizes a weather broadcast voice that is both characteristic of a mother's voice and full of warm and caring tones, allowing family members to feel warmth in the cold weather.
[0106] Optionally, in S24, the television terminal synthesizes emotionally-inspired speech using a speech synthesis engine based on the text to be read, a personalized voiceprint model, and emotional intensity adjustment parameters, including: The TV inputs the text to be broadcast into the speech synthesis engine.
[0107] The speech synthesis engine on the TV is based on a personalized voiceprint model to determine the user's basic timbre, intonation, and speech rate characteristics.
[0108] On the TV, based on the user's basic timbre, intonation, and speech rate characteristics, the fundamental frequency, energy, and duration parameters of the speech to be synthesized are dynamically adjusted according to the emotional intensity adjustment parameters, so as to superimpose the emotional color corresponding to the emotional intensity adjustment parameters on the basic timbre and generate emotional broadcast speech.
[0109] In this solution, the television terminal first inputs the text to be broadcast into a speech synthesis engine, which integrates a deep neural network acoustic model and a vocoder module. The speech synthesis engine determines the user's basic timbre, intonation, and speech rate characteristics based on a personalized voiceprint model. The basic timbre characteristics reflect the physiological structure of the user's vocal organs and pronunciation habits, serving as a core identifier for user identity. The intonation characteristics reflect the user's habitual pitch variation patterns, and the speech rate characteristics reflect the user's habitual speaking speed and pause rhythm. After determining the user's basic characteristics, the speech synthesis engine dynamically adjusts the fundamental frequency, energy, and duration parameters of the synthesized speech according to the emotion intensity adjustment parameters. This superimposes the emotional color corresponding to the emotion intensity adjustment parameters onto the basic timbre, generating emotionally-driven broadcast speech. Specifically, while maintaining the user's basic timbre, intonation, and speech rate characteristics, the speech synthesis engine independently adjusts the fundamental frequency, energy, and duration parameters related to emotional expression to achieve the superposition of emotional color without altering the user's identity characteristics. Fundamental frequency parameters determine the pitch and intonation variations of speech. By adjusting the amplitude and dynamic range of the fundamental frequency curve, different emotional states can be expressed. For example, pleasant emotions typically have a large fundamental frequency dynamic range and rising intonation, while calm emotions have a relatively flat fundamental frequency curve. Energy parameters determine the volume and loudness variations of speech. By adjusting the energy distribution, emotional intensity can be expressed. For example, excited emotions have a high energy peak and a large energy dynamic range, while mild emotions have a relatively uniform energy distribution. Duration parameters determine the speech rate and rhythm variations of speech. By adjusting the phoneme duration and pause positions, emotional prosody can be expressed. For example, lively emotions typically have a faster speech rate and shorter pauses, while solemn emotions have a slower speech rate and longer pauses. During the dynamic adjustment process, the speech synthesis engine first performs prosodic structure analysis on the text to be broadcast, determining the stress and pause positions of the text. Prosodic structure analysis uses natural language processing techniques to identify grammatical boundaries, keywords, and semantic units in text. For example, in the sentence "The weather is so nice today," "so nice" might be an accented word, expressing strong emotion. In long sentences, commas, periods, and other punctuation marks are usually natural pauses. The speech synthesis engine jointly optimizes the adjustment range of fundamental frequency, energy, and duration parameters based on the emotion intensity adjustment parameters and the identified prosodic structure to ensure that the emotional expression is coordinated with the prosodic structure of the text. For example, for accented words, the fundamental frequency variation and energy enhancement are appropriately increased according to the emotion intensity parameters; for pauses, the pause duration is adjusted according to the emotion intensity parameters—a lively emotion can shorten the pause, while a solemn emotion can lengthen the pause; for ordinary text areas, a smooth transition of fundamental frequency, energy, and duration is maintained. Through this joint optimization, the expression of emotional coloring is ensured to blend with the natural rhythm of the text, avoiding unnatural speech effects caused by emotional superposition. The specific adjustment range is determined by the emotion intensity adjustment parameters and the prosodic structure of the text, which will not be elaborated further here.Furthermore, the speech synthesis engine inputs the adjusted fundamental frequency, energy, and duration parameters into the vocoder. The vocoder generates the final speech waveform based on these acoustic parameters, outputting an emotionally resonant voice that combines the user's natural timbre with the target emotional tone and harmonizes with the text's rhythm. Using this method, the television terminal achieves a deep integration of emotional expression and text content by independently adjusting emotion-related parameters while maintaining the user's identity characteristics, and by jointly optimizing these parameters with the text's prosodic structure. This allows the synthesized speech to accurately reproduce the user's unique timbre, flexibly express emotions according to their intensity, and perfectly harmonize with the text's prosodic structure, enhancing the naturalness and impact of the emotionally resonant broadcast.
[0110] In practical applications, such as on a television, a warm reminder voice message might be synthesized based on the mother's personalized voiceprint model: "Remember to bring an umbrella tomorrow, the weather forecast says it will rain." The voice synthesis engine first extracts the mother's unique basic timbre, gentle tone, and moderate speaking speed from her voiceprint model. The engine performs prosodic structure analysis on the text, determining that there should be a slight pause after the "baby" part of the vocative, with the umbrella being the emphasis in "Remember to bring an umbrella," and the "weather forecast says it will rain" providing supplementary information. Based on preset 80 points of warm emotional intensity adjustment parameters, while maintaining the mother's basic timbre, gentle tone, and moderate speaking speed, the engine independently adjusts the fundamental frequency parameter to slightly raise the "umbrella" part to express concern, adjusts the energy parameter to slightly enhance the "umbrella" part to highlight the reminder, and adjusts the duration parameter to add a short pause after the "baby" part to enhance intimacy, while maintaining a smooth transition in other areas. The final generated audio message retains the mother's unique voice while being full of warm and caring tone. It also blends naturally with the rhythmic structure of the text, allowing the child to clearly perceive that it is the mother's voice and feel the importance of the mother's love and reminder.
[0111] In one optimized approach, the method further includes: The television terminal performs deep semantic understanding of the text to be broadcast, identifying key emotional points and emotional turning points in the text.
[0112] The television broadcaster dynamically adjusts the emotional intensity adjustment parameters during the broadcast process based on key emotional points and emotional turning points, so as to ensure that the emotional expression matches the semantic depth of the text.
[0113] In this solution, after acquiring the text to be broadcast, the television terminal performs deep semantic analysis using natural language processing technology to identify keywords, phrases, or sentences containing emotional connotations as emotional key points. These include words directly expressing emotion such as "really happy" and "what a pity," as well as conjunctions or contextual changes indicating emotional shifts such as "but," "however," and "suddenly." Based on these identified emotional key points and emotional shift points, the television terminal constructs an emotional change trajectory for the text and dynamically adjusts the emotional intensity parameter during broadcasting. For example, when broadcasting a story that shifts from sadness to encouragement, the television terminal uses a lower emotional intensity parameter in the sad sections, gradually increasing the emotional intensity at the turning point, reaching a higher emotional intensity parameter in the encouraging sections, forming a smooth emotional transition curve. This dynamic adjustment does not simply unify the entire text to a single emotional intensity, but rather adjusts the emotional expression intensity of each sentence and even each word in real time according to the fluctuations in the text's semantics, ensuring a deep match between the emotional expression of the synthesized speech and the emotional connotation of the text content. By employing the above method, the TV terminal performs deep semantic understanding of the text and dynamically adjusts the emotional intensity change curve, enabling the emotional broadcast to accurately follow the emotional context of the text. This makes the emotional expression of the synthesized speech more delicate, natural, and infectious, allowing users to truly feel the emotional fluctuations contained in the text and obtain an immersive auditory experience.
[0114] Combination Figure 5 As shown, optionally, this disclosure provides another method for emotional voice broadcasting, applied to a mobile terminal, the method including: S41, in response to the user's voice reproduction command, the mobile terminal guides the user to record multiple voice samples for different preset emotion types on the interactive interface. The preset emotion types include greetings, reminders and blessings.
[0115] S42, the mobile terminal uploads multiple voice samples to the cloud server, so that the cloud server can use a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters of the multiple voice samples, and generate a personalized voiceprint model that includes the user's timbre features and the emotional features corresponding to the preset emotional type.
[0116] S43, the mobile terminal receives the emotional intensity adjustment parameters set by the user and outputs the emotional intensity adjustment parameters so that the TV terminal can synthesize emotional broadcast voice based on the text to be broadcast, the personalized voiceprint model and the emotional intensity adjustment parameters when broadcasting.
[0117] In this solution, the mobile terminal responds to the user's voice reproduction command, guiding the user to record multiple voice samples for different preset emotion types on the interactive interface. These preset emotion types include greetings, reminders, and blessings. The mobile terminal uploads these voice samples to a cloud server, where a pre-trained deep learning model extracts voiceprint features and analyzes emotional parameters, generating a personalized voiceprint model that includes the user's voice characteristics and the corresponding emotional features for each preset emotion type. The mobile terminal receives and outputs the emotion intensity adjustment parameters set by the user, enabling the TV to synthesize emotionally-influenced broadcast voice based on the text to be broadcast, the personalized voiceprint model, and the emotion intensity adjustment parameters. During the process of guiding the user to record voice samples, the mobile terminal displays recording prompts corresponding to greeting, reminder, and blessing emotions in a sequential manner on the mini-program interface. Each prompt includes a corresponding example text. For example, the prompt for greeting might display "Please say in a greeting tone: Good morning"; the prompt for reminder might display "Please say in a reminder tone: Remember to bring an umbrella"; and the prompt for blessing might display "Please say in a blessing tone: Happy birthday". The mobile terminal collects the user's voice recordings of each example text, serving as voice samples for the corresponding emotion type. Each sample is 10 to 15 seconds long to ensure both complete capture of the user's vocal characteristics and emotional expression without placing an excessive recording burden on the user. After recording, the mobile terminal can play back the recorded voice sample for user confirmation; users can re-record if dissatisfied. Furthermore, after uploading the voice sample, the mobile terminal provides an emotion intensity adjustment function. Users can adjust the emotion intensity within a range of 0 to 100 using a slider control. For example, setting a greeting emotion to 80 points to express warmth, a reminder emotion to 60 points to express concern, and a blessing emotion to 90 points to express joy. The mobile terminal can preview the playback effect in real time for user reference and adjustment. After the user completes the emotion intensity setting, the mobile terminal outputs the emotion intensity adjustment parameter, which can be synchronized to the TV via a cloud server or directly sent to the TV within the same local area network for use in subsequent broadcasts. Using the above method, the mobile terminal guides users to record voice samples of multiple emotional types and set emotional intensity parameters through the mini-program interface. This provides high-quality raw data for the cloud server to generate personalized voiceprint models. At the same time, it provides users with a convenient and intuitive entry point for emotional settings, allowing ordinary family users to customize personalized emotional voices without professional equipment and technical knowledge. This lowers the technical threshold and improves the user experience.
[0118] An optimized solution also includes: After receiving the emotional intensity adjustment parameters set by the user, the mobile terminal synthesizes a preview voice based on the personalized voiceprint model and the currently set emotional intensity adjustment parameters, and plays it for the user to preview the broadcast effect.
[0119] In this solution, after the user adjusts the emotion intensity parameter using the slider control, the mobile terminal can obtain a personalized voiceprint model generated by the cloud server in real time, or use a locally cached voiceprint model. Combined with the currently set emotion intensity adjustment parameters and preset sample text, a lightweight speech synthesis engine is invoked to quickly synthesize a preview speech. The synthesis process of this preview speech is consistent with the actual broadcast process on the television, both using the same voiceprint model and emotion intensity parameter adjustment mechanism to ensure that the preview effect truly reflects the emotional expression of the actual broadcast. The mobile terminal plays the preview speech through its built-in speaker or headphones, allowing the user to intuitively hear the actual effect of the broadcast speech under the current emotion intensity setting, including timbre fidelity, emotional intensity, and naturalness of tone changes. If the user is not satisfied with the preview effect, they can continue to adjust the emotion intensity slider, and the system updates the preview speech in real time until the user achieves a satisfactory emotional expression effect. Using this method, the mobile terminal allows users to instantly perceive the broadcast effect when setting the emotion intensity, enabling them to accurately adjust the emotion intensity to the expected level. This improves the accuracy of emotion settings and user satisfaction, avoiding the repeated adjustments required for subsequent unsatisfactory broadcast effects on the television due to improper settings.
[0120] Combination Figure 6 As shown, this embodiment of the disclosure provides an emotional voice broadcasting device 300, including a processor 301 and a memory 302. Optionally, the device 300 may further include a communication interface 303 and a bus 304. The processor 301, communication interface 303, and memory 302 can communicate with each other via the bus 304. The communication interface 303 can be used for information transmission. The processor 301 can call logical instructions in the memory 302 to execute the emotional voice broadcasting method of the above embodiment.
[0121] Furthermore, the logic instructions in the aforementioned memory 302 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0122] The memory 302, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 301 executes functional applications and data processing by running the program instructions / modules stored in the memory 302, thereby implementing the emotional voice broadcasting method in the above embodiments.
[0123] The memory 302 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 302 may include high-speed random access memory and may also include non-volatile memory.
[0124] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described emotional voice broadcasting method.
[0125] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0126] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0128] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. An emotional voice broadcasting method, characterized in that, Applied to cloud servers, the methods include: Receive multiple voice samples, which are recorded by the user for different preset emotion types, including greetings, reminders, and blessings; A pre-trained deep learning model is used to extract voiceprint features and analyze emotional parameters from the multiple speech samples to generate a personalized voiceprint model; wherein, the personalized voiceprint model includes user timbre features and emotional features corresponding to the preset emotional type; The personalized voiceprint model is associated with and stored with the user identifier; In response to a model synchronization request carrying a user identifier sent by the TV terminal, a corresponding personalized voiceprint model is matched based on the user identifier, and the matched personalized voiceprint model is synchronized to the TV terminal so that the TV terminal can synthesize emotional broadcast speech based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters during broadcast.
2. The method according to claim 1, characterized in that, Using a pre-trained deep learning model, voiceprint features are extracted and emotion parameters are analyzed from the multiple speech samples to generate a personalized voiceprint model, including: Each of the speech samples is preprocessed, including signal digitization, pre-emphasis processing, and sound framing. The preprocessed speech samples are input into a pre-trained self-supervised learning feature extractor to extract deep acoustic features, including timbre features. For each emotion type of speech sample, the emotion feature corresponding to that emotion type is extracted from the deep acoustic features; The timbre features in the deep acoustic features are identified as the user timbre features, and the user timbre features are associated with the emotional features of each of the emotional types to generate the personalized voiceprint model.
3. The method according to claim 1, characterized in that, Also includes: Receives emotion intensity adjustment parameters sent by a mobile terminal, wherein the emotion intensity adjustment parameters are set by the user on the mobile terminal side through an interactive interface; The emotional intensity adjustment parameters are associated with and stored in the personalized voiceprint model; In response to a model synchronization request from the television, the emotional intensity adjustment parameters and the personalized voiceprint model are synchronized together to the television.
4. An emotional voice broadcasting method, characterized in that, For television applications, the methods include: Send a model synchronization request carrying the user's identifier to the cloud server; The system receives and stores the personalized voiceprint model returned by the cloud server based on the user identifier and obtains the emotion intensity adjustment parameters; wherein, the personalized voiceprint model is generated by multiple voice samples recorded by the user for different preset emotion types and uploaded by the mobile terminal, including the user's timbre features and the emotion features corresponding to the preset emotion types. In response to the broadcast trigger event, obtain the text to be broadcast; Based on the text to be broadcast, the personalized voiceprint model, and the emotion intensity adjustment parameters, an emotional broadcast voice is synthesized through a speech synthesis engine and then played.
5. The method according to claim 4, characterized in that, The broadcast trigger event includes the TV receiving a voice command from the user, and the method further includes: The voice commands are semantically recognized to extract the identification information used to specify the broadcast timbre; Based on the identity information, the corresponding personalized voiceprint model is retrieved from local storage.
6. The method according to claim 4, characterized in that, The broadcast trigger event includes the TV detecting that the currently playing content has switched to a preset application scenario, and the method further includes: Identify the current application scenario type, which includes weather forecast scenario, schedule reminder scenario, or educational content explanation scenario; Based on the application scenario type, a preset emotional pattern is automatically matched; The emotion intensity adjustment parameter is determined based on the emotion pattern.
7. The method according to claim 4, characterized in that, Based on the text to be broadcast, the personalized voiceprint model, and the preset emotion intensity adjustment parameters, an emotional broadcast voice is synthesized through a speech synthesis engine, including: The text to be broadcast is input into the speech synthesis engine; The speech synthesis engine determines the user's basic timbre, intonation, and speech rate characteristics based on the personalized voiceprint model. Based on the user's basic timbre, intonation, and speech rate characteristics, the fundamental frequency, energy, and duration parameters of the speech to be synthesized are dynamically adjusted according to the emotional intensity adjustment parameters, so as to superimpose the emotional color corresponding to the emotional intensity adjustment parameters on the basic timbre to generate emotional broadcast speech.
8. An emotional voice broadcasting method, characterized in that, Applied to mobile terminals, the methods include: In response to the user's voice reproduction command, the user is guided on the interactive interface to record multiple voice samples for different preset emotional types, including greetings, reminders and blessings; The multiple speech samples are uploaded to a cloud server so that the cloud server can use a pre-trained deep learning model to extract voiceprint features and analyze emotional parameters of the multiple speech samples, and generate a personalized voiceprint model that includes user timbre features and emotional features corresponding to the preset emotional type. The system receives and outputs the emotional intensity adjustment parameters set by the user, so that the TV terminal can synthesize emotional broadcast voice based on the text to be broadcast, the personalized voiceprint model, and the emotional intensity adjustment parameters during broadcast.
9. An emotional voice broadcasting device, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the emotional voice broadcasting method as described in any one of claims 1 to 8 when running the program instructions.
10. A computer-readable storage medium storing program instructions, characterized in that, When the program instructions are executed, they cause the computer to perform the emotional voice broadcasting method as described in any one of claims 1 to 8.