Voice recognition and voice broadcast system

By adopting GAN and deep neural network technology in the speech synthesis system, flexible adjustment of speech parameters and automatic format conversion are achieved, which solves the shortcomings in the speech nature, emotional expression and device compatibility in the existing technology, and improves the speech nature and user experience.

CN120048246APending Publication Date: 2025-05-27GUANGZHOU NINE CHIP ELECTRON SCI & TECH CO LTD

Patent Information

Application Number
CN202510190098.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing speech synthesis system has shortcomings in nature, emotional expression and device compatibility, which leads to the generated speech appearing dull and unable to meet the needs of different application scenarios. At the same time, there are also problems with file management and device compatibility.

Method used

Through the speech recognition and voice broadcasting system, the technology combined with generative adversarial network (GAN) and deep neural network is adopted to optimize the speech synthesis process, realize the flexible adjustment of voice parameters and real-time pre-listening functions, and provide automatic format conversion and cross-device synchronization.

Benefits of technology

It improves the naturalness and emotional expression of voice, enhances device compatibility and file management convenience, ensures that voice files are played compatible on different platforms, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048246A_ABST
    Figure CN120048246A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent equipment and man-machine interaction, and discloses a voice recognition and voice broadcast system, which comprises a text input module used for receiving text content input by a user; the voice synthesis module is used for generating a corresponding voice signal according to the text input by the user; the parameter configuration module is used for allowing a user to customize voice parameters including timbre, speed, volume, tone and brightness; the file generation module is used for converting the synthesized voice into a downloadable audio format file; and the downloading management module is used for providing storage and downloading of the voice file. A user can adjust voice parameters such as timbre, speech speed, volume, tone, brightness and the like, audio files in various formats are generated, and management of local storage and cloud storage is supported. Compared with the prior art, more flexible voice customization, higher voice quality and higher equipment compatibility are provided, and the user experience is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent devices and human-computer interaction, and specifically to a system for speech recognition and speech broadcast. Background Art

[0002] With the rapid development of intelligent voice technology, existing speech synthesis systems have been widely used in many fields. Most traditional speech synthesis methods generate speech based on rules or models. Although they can produce standard speech outputs, there are often significant gaps in the naturalness and emotional expression of the speech. Speech synthesis systems in the prior art usually rely on preset timbre, speech rate, and pitch parameters, which makes the generated speech appear rigid and lack variation, and cannot fully meet the requirements of speech expression in different application scenarios.

[0003] Specifically, the parameter settings in traditional speech synthesis technology are mostly static, lacking the ability to dynamically adjust emotional colors and tone changes. When users use speech synthesis systems, they often can only select preset timbres and speech rates, and cannot flexibly adjust the emotional tone or the strength of the tone of the speech according to actual needs. For example, in scenarios such as customer service and intelligent assistants, the emotional expression and tone changes of the speech are crucial for enhancing the user experience. However, the prior art is difficult to achieve natural emotional expression and tone changes, resulting in the speech sounding mechanical and lacking emotional interaction with users.

[0004] In addition, existing speech file generation and storage technologies also have obvious deficiencies in device compatibility and file management. Traditional audio file generation modules usually only support one or a few formats, and the storage and conversion processes of audio files are complex. The user experience of file sharing and playing between different devices is poor. Most of the prior art does not have the function of automatically optimizing file formats and compression ratios, resulting in the generated speech files may not be compatible for playback on different platforms, or waste too much resources in terms of storage space. Therefore, the deficiencies of the prior art in file format conversion, device compatibility, and cross-device file synchronization greatly reduce the management and use experience of speech files; Therefore, a system for speech recognition and speech broadcast is proposed to solve the above problems. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a system for speech recognition and speech broadcast, which solves the deficiencies of existing speech synthesis systems in terms of naturalness, emotional expression, and device compatibility. Through flexible speech parameter adjustment, real-time preview function, and multi-device support, users can generate personalized speech files according to their needs. The system optimizes the naturalness of speech output and emotional transmission, and at the same time provides functions of automatic format conversion and cross-device synchronization to ensure the compatibility and convenient management of speech files.

[0006] To achieve the above object, the present invention is implemented through the following technical solutions: A system for speech recognition and speech broadcast, including: A text input module for receiving the text content input by the user; A speech synthesis module for generating corresponding speech signals according to the text input by the user; A parameter configuration module for allowing the user to customize speech parameters, including timbre, speech rate, volume, pitch, and brightness; A file generation module for converting the synthesized speech into a downloadable audio format file; A download management module for providing storage and download of speech files, enabling the user to download the generated speech files to the local or burn them into the target device.

[0007] Preferably, the speech synthesis module includes: A text encoding unit for encoding the input text information to generate an intermediate representation; A speech generation unit for generating corresponding speech signals according to the encoded text information; An optimization unit for adjusting parameters such as pitch, speech rate, volume, pitch, and brightness according to the characteristics of the generated speech to optimize the naturalness and emotional expression of the speech.

[0008] Preferably, the speech synthesis module is optimized through a generative adversarial network model, and the generative adversarial network model includes: A generator for generating synthetic speech as close as possible to real speech; A discriminator for distinguishing between generated speech and real speech and continuously improving the quality of the synthetic speech through adversarial training.

[0009] Preferably, the speech synthesis module optimizes the generated speech signal by minimizing the Kullback-Leibler divergence to minimize the difference between it and the target speech signal. The Kullback-Leibler divergence is defined as: where D KL (P||Q) is the Kullback-Leibler divergence; P(x) is the probability distribution of the target speech signal, representing the feature distribution of the target speech signal; Q(x) is the probability distribution of the generated speech signal, representing the feature distribution of the synthetic speech signal; x is the sample point or feature of the audio signal, representing the state or feature of the speech waveform at a specific moment.

[0010] Preferably, the download management module allows users to save the generated voice files in multiple formats and supports selecting appropriate encoding methods and bitrates according to device requirements to ensure the compatibility and playback quality of voice files on different hardware devices.

[0011] Preferably, the parameter configuration module provides a real-time preview function. After adjusting parameters such as tone, speech rate, volume, pitch, and brightness, users can listen to the generated effect in real time to optimize the final voice output.

[0012] Preferably, the file generation module includes a voice file format conversion function, which can automatically select the optimal file format according to the compatibility of the target device and compress or optimize the file to improve storage efficiency and transmission quality.

[0013] Preferably, the download management module supports local storage and cloud storage options and provides functions for managing, backing up, and restoring voice files to ensure that users can access and use the stored voice files at any time.

[0014] Preferably, the speech synthesis module further combines multiple speech models, including but not limited to natural human speech, robot speech, and dialect speech, and allows users to select or customize speech models to adapt to different scenario requirements.

[0015] The present invention also provides a method for speech recognition and speech broadcast, including the following steps: Receiving text information input by the user; Using the speech synthesis module to generate a target voice signal according to the input text; Optimizing parameters such as pitch, speech rate, volume, pitch, and brightness of the voice according to the generated voice signal to enhance the naturalness and emotional expression of the voice; Converting the synthesized voice signal into a downloadable audio file; Providing the user with the generated audio file for downloading, saving, or burning to the target device.

[0016] The present invention provides a system for speech recognition and speech broadcast. It has the following beneficial effects: 1. The present invention adopts a technical solution that combines a generative adversarial network (GAN) with a deep neural network, optimizes the speech synthesis process, and realizes high-quality speech generation, especially the improvement in the naturalness and emotional expression of speech. Compared with the speech synthesis methods that rely on traditional parameter models in the prior art, the present invention effectively solves the deficiencies of mechanization and lack of emotion in speech synthesis through the adversarial training method of the generator and discriminator, making the generated speech closer to the natural fluency of real human pronunciation.

[0017] 2. The parameter configuration module of the present invention not only supports the adjustment of traditional parameters such as timbre, speech rate, volume, pitch, and brightness, but also innovatively introduces a real-time emotion and tone adjustment function, allowing users to adjust the emotional color of speech based on actual needs. Compared with the static settings in the prior art, the present invention provides users with more personalized speech adjustment options, enabling better expression of the emotion and tone of speech in different application scenarios, and significantly enhancing the affinity and expressiveness of voice interaction.

[0018] 3. The file generation module of the present invention realizes automated audio format conversion and compression optimization, supports multiple storage formats for generating voice files, and can be optimized and adjusted according to the compatibility of different devices. Compared with the systems in the prior art that only support single-format output, the present invention, through adaptive format selection technology, automatically judges the requirements of the target device, provides higher device compatibility, solves the problem of incompatibility between audio file formats and playback devices, and ensures that users can obtain the best voice playback effect on any device.

[0019] 4. The present invention innovatively combines the dual functions of cloud storage and local storage, and realizes cross-device synchronization, backup, and recovery functions for files. Compared with the solutions in the prior art that only rely on local storage or simple cloud storage functions, the present invention greatly improves the flexibility and security of file management through intelligent synchronization and convenient recovery. Users can access voice files on different devices at any time, greatly enhancing the user experience and solving the problems of data sharing and recovery between devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of the system framework of the present invention; Figure 2 is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Please refer to the attached Figure 1 , the embodiments of the present invention provide a system for speech recognition and speech broadcast, including: A text input module for receiving the text content input by the user; The text input module is the first step of the speech recognition and speech broadcasting system. Its main function is to receive the text content input by the user and process it to ensure that the text can be accurately passed to the speech synthesis module. This module supports multiple input methods, including keyboard input, speech-to-text, handwriting input, etc., so that the system can adapt to various devices and application scenarios. The accuracy and flexibility of the text input module are crucial to the overall performance of the system; In this embodiment, the text input module is responsible for receiving the text data input by the user and processing the text. Generally, the user can input the text content through a keyboard, a touch screen or a voice input device. The text input module standardizes, cleans and formats the received text content to ensure the efficiency and accuracy of subsequent processing.

[0023] Specifically, the text input module will process the input text in the following steps: Text cleaning: Remove irrelevant symbols or formatting issues in the text, such as spaces, special characters, etc., for subsequent processing.

[0024] Text standardization: For multilingual text input, the system will encode the text in a unified manner. Common standardization methods include UTF-8 encoding to ensure the consistency of multilingual text.

[0025] Word segmentation and grammatical analysis: In languages ​​without spaces, such as Chinese, the text input module uses word segmentation technology to divide continuous text into words. For English text, the system separates words in the text by spaces. In addition, the system also performs grammatical analysis and dependency analysis to identify subject-verb-object structures and other language relationships in the text.

[0026] Emotion and tone analysis: In order to make subsequent speech synthesis more natural, the system can analyze the emotional tendencies in the input text, such as joy, sadness, anger, etc., and provide the necessary information for the speech synthesis module to achieve emotional expression.

[0027] Mapping and output: The processed text content is converted into a standardized data structure and then passed to the speech synthesis module. The text input module not only passes the original text, but also provides other additional information, such as emotional tags, voice style requirements, etc.

[0028] In one possible implementation, the text input module combines natural language processing (NLP) technology to perform complex text analysis and understanding. These technologies include but are not limited to: text semantic analysis, named entity recognition (NER), sentiment analysis, keyword extraction, etc.

[0029] Input text analysis: Assume that the input text is T = {t 1 ,t 2 ,...,tn}, where T is the text sequence input by the user, and t i represents the i-th word or character of the input. For Chinese input, the text input module first segments the text through a word segmentation algorithm to generate a word sequence T′: T′ = {t 1 ′, t 2 ′,..., t m ′} where, t′ 1 , t′ 2 ,..., t′ m are the results after word segmentation, and m represents the number of words after word segmentation. For English text, the system splits it by spaces to obtain a word sequence.

[0030] Sentiment analysis and tone judgment: During the sentiment analysis process, the system analyzes the sentiment tendency of the input text according to its content. The formula for sentiment analysis is expressed as: S = f(T′) where, S is the sentiment score, T′ is the text after word segmentation processing, f represents the sentiment analysis function, and this function maps the text to a sentiment space to obtain the sentiment score. This score will be provided to the speech synthesis module to adjust the emotional color of the speech during synthesis.

[0031] Text normalization and encoding: After the input text is normalized, the system will convert it into a unified encoding format (such as UTF-8). For each word t i , the system will perform unified encoding according to the requirements of the character set: t i = encode(t i ) Syntax analysis and mapping: Among them, encode is the normalization encoding function to ensure that the text can be processed internally by the system. Syntax analysis performs structured processing on the text to generate a dependency tree or a syntax tree. For example, the input sentence "I like listening to music" can be mapped to the dependency relationship: Sentence = {Subject: I, Predicate: like, Object: music} Through this mapping, the system can understand the syntax structure of the text and provide more accurate information for subsequent speech synthesis.

[0032] To further improve the accuracy and reliability of text input, the system uses a probability model to evaluate the occurrence probability of each word during the text parsing process, especially during text word segmentation. Specifically, the occurrence probability of a word is calculated by the following formula: where: P(ti |T′) is the word t i The conditional probability in the given text context T′; C(t i ,T′) is the word t i The number of times it appears in the text T′; C(T′) is the total number of occurrences of all words in the text T′.

[0033] This probability-based word segmentation method can effectively solve the problems of ambiguity and incorrect word segmentation in text input.

[0034] In this embodiment, the text input module can efficiently and accurately process different types of input text. Through steps such as cleaning, standardizing, word segmentation, and syntactic analysis of the input text, the system can ensure the integrity and accuracy of the text information. In addition, through sentiment analysis and tone recognition, the system can provide clues about sentiment and intonation for subsequent speech synthesis, making the synthesized speech more in line with user needs and improving the naturalness and expressiveness of the speech.

[0035] This module also supports multiple input methods, such as manual input, speech-to-text, and handwriting input. Each input method can be parsed through different processing algorithms to ensure that the system can handle complex input situations and meet various practical application requirements.

[0036] The speech synthesis module is used to generate corresponding speech signals according to the text input by the user; The speech synthesis module is the core part of the entire system. The text input module provides processed text information, and the speech synthesis module is responsible for converting the text into corresponding speech signals for subsequent storage and playback. This module not only needs to ensure the basic intelligibility of the speech but also meet the user's personalized requirements for timbre, speech rate, pitch, etc. To achieve high-quality speech output, this module uses deep neural networks, optimization algorithms, and speech parameter adjustment mechanisms to improve the naturalness, clarity, and emotional expression ability of the speech; In this embodiment, the speech synthesis module receives the text data from the text input module and performs a series of processing on it to generate the expected speech signals. Generally, this module includes a text encoding unit, a speech generation unit, and an optimization unit, which are used for the structured processing of the text, the synthesis of speech signals, and the optimization of the final speech, respectively.

[0037] The main task of the text encoding unit is to convert the input text into a standardized intermediate representation. The text first undergoes syntactic analysis to identify the syntactic structure and perform prosody prediction to ensure that the speech has natural pauses, accents, and rhythm changes during pronunciation. For texts in different languages, the system will use different annotation rules to ensure the accuracy of the synthesized speech.

[0038] Specifically, the text encoding unit uses a phoneme sequence mapping technique to convert the input text into corresponding phonemes to form a phoneme sequence P = {p 1 , p 2 ,..., p n}, where: p i represents the i-th phoneme; n is the total number of phonemes after text conversion.

[0039] In a possible implementation, the text encoding unit uses a model based on a deep neural network (DNN) to predict the correspondence between the phoneme sequence and prosodic parameters. The speech generation unit is responsible for converting the phoneme sequence into a speech signal. In this process, the system uses a waveform generation model, such as an architecture based on Tacotron, WaveNet, or Transformer-TTS, to generate a high-fidelity speech signal. The core calculation of speech generation can be expressed as follows: S = G(P, M) where: S is the generated speech signal; G(·) is the speech generation function; P is the phoneme sequence; M is the prosodic parameter matrix, including information such as speech rate, pitch, and stress.

[0040] In some embodiments, the speech generation unit can further incorporate a conditional variational autoencoder (CVAE) to improve the stylization ability of the speech and make it more diverse. For example, different users can select different speech styles, such as formal broadcast, natural conversation, children's speech, etc.

[0041] The optimization unit is used to adjust the final speech signal to ensure that the synthesized speech is more natural and meets the parameter requirements set by the user. This unit mainly involves the adjustment of key parameters such as timbre, speech rate, volume, pitch, and brightness. The key calculation process of optimization can be expressed as: S′ = O(S, A) where: S′ is the optimized speech signal; O(·) is the optimization function; A is the speech parameters set by the user, including timbre, speech rate, etc.

[0042] During the optimization process, the system adopts an optimization strategy based on minimizing perceptual distortion to ensure that the optimized speech is as close as possible to real speech in terms of auditory perception.

[0043] As an option, the optimization unit can also apply an autoregressive filter to reduce the noise in the speech signal and improve the clarity and stability of the speech. This technology is particularly suitable for speech broadcast scenarios that need to be played in a noisy environment.

[0044] In some possible implementations, the optimization unit also incorporates adaptive timbre mapping to match the audio playback characteristics of different devices, so that the speech has a consistent listening experience when played on different devices.

[0045] In this embodiment, all processing steps of the speech synthesis module are carried out under a real-time computing framework to ensure that the speed of speech synthesis meets the user interaction requirements. For longer texts, the system adopts incremental speech synthesis technology to process the text in segments, avoiding delays or lags during the speech playback process.

[0046] To ensure the smoothness and naturalness of the speech signal, the optimization unit also uses the Short-Time Fourier Transform (STFT) for spectral analysis to optimize the frequency components of the speech. This process can be expressed as: where: X(t,f) is the speech signal represented in the time-frequency domain; s(n) is the speech signal in the time domain; w(n) is the window function; t is the time frame index; f is the frequency index.

[0047] Through STFT analysis, the system can adaptively adjust different frequency components during the optimization stage to improve the naturalness and clarity of the speech.

[0048] A parameter configuration module, used to allow users to customize speech parameters, including timbre, speech rate, volume, pitch, and brightness; The parameter configuration module plays a crucial role in the speech synthesis system. It allows users to customize multiple parameters of the speech, such as timbre, speech rate, volume, pitch, and brightness, etc., in order to generate personalized speech output that meets the user's needs. This module can adjust these speech parameters in real time and provides a preview function, enabling users to listen to the adjustment effects during the adjustment process, thereby further optimizing the speech output. Through this module, users can flexibly customize the expression mode of the speech to ensure that the finally generated speech meets the requirements of specific scenarios.

[0049] In this embodiment, the parameter configuration module is a key component in the speech synthesis system, and its main task is to provide users with the function of customizing speech parameters. After the user inputs text, the system converts it into a speech signal through the speech synthesis module. However, in order to better adapt to different usage scenarios, users usually need to adjust various features of the speech. The parameter configuration module provides multiple adjustable parameters such as timbre, speech rate, volume, pitch, and brightness for this purpose.

[0050] Specifically, timbre refers to the quality and style of speech. Through this module, users can select different types of timbres, such as male, female, child, or robotic voices. The speaking speed determines how fast the speech is spoken, and users can adjust it to a fast or slow speed according to their needs. The volume parameter allows users to adjust the loudness of the speech. The pitch controls the high and low of the speech, and users can adjust the pitch to achieve different tonal effects. Finally, the clarity refers to the clarity and brightness of the speech, which is suitable for clearly conveying information or creating a certain emotional atmosphere.

[0051] In some embodiments, the parameter configuration module is also capable of supporting speech stylization, allowing users to adjust the style and emotional expression of the speech according to specific application requirements, such as customer service voice, broadcast voice, news broadcast voice, etc.

[0052] In this embodiment, the parameter configuration module has a real-time preview function, that is, when the user adjusts the parameters, they can immediately hear the adjusted speech effect. Through this interaction method, users can more intuitively understand the impact of different parameters on the speech and further optimize the settings according to their needs.

[0053] Specifically, by adjusting each speech parameter, the parameter configuration module can match the generated speech with the user's expected speech output. For different application scenarios, the system can adjust the quality of the speech according to the usage background of the speech. For example, in a smart home application, the affinity of the speech can be adjusted; in a navigation system, the clarity and speaking speed of the speech can be enhanced to ensure the accurate conveyance of instructions.

[0054] In a possible implementation, the parameter configuration module adopts an adaptive algorithm to automatically optimize multiple features of the speech according to the user's settings. For example, when adjusting the speaking speed, the system will automatically adjust the pitch and timbre to maintain the naturalness and fluency of the speech. In this way, users can more easily achieve the ideal speech output effect.

[0055] Technical implementation: The technical implementation of the parameter configuration module can be achieved through a deep learning model to perform multi-dimensional adjustment of the speech. Each speech parameter corresponds to a specific network structure, and by adjusting the parameters of these networks, the generation of the final speech can be controlled. For example, in the adjustment of timbre, the system processes the training data of different timbres to learn how to adjust the vibration mode of the vocal tract when generating speech, thereby affecting the quality of the speech.

[0056] In a possible implementation, the algorithm for timbre adjustment can be expressed as: S new =f timbre (S origin ,θ) Where: S new is the adjusted speech signal; ftimbre is the timbre adjustment function; S origin is the original voice signal; θ is the parameter set for timbre adjustment, including the timbre type selected by the user. Similarly, parameters such as speech rate, volume, and pitch can also be adjusted by a similar method. For the adjustment of speech rate, its function can be expressed as: S new = f speed (S origin , v) where: f speed is the speech rate adjustment function; v is the speech rate value set by the user, usually a positive real number, indicating the speed of speech.

[0057] Through these functions, the parameter configuration module can finely adjust each parameter to ensure that the voice output meets the user's needs.

[0058] To enable the user to listen to the generated voice effect in real time during the adjustment process, the parameter configuration module realizes two-way feedback of voice generation and real-time playback through linkage with the speech synthesis module. Specifically, when the user changes a certain voice parameter, the system will immediately generate the adjusted voice and play it to the user through the speaker or headphones so that the user can perceive the adjustment effect in the actual environment.

[0059] Through the parameter configuration module, the system can greatly improve the personalization and flexibility of speech synthesis. Generally, the user can adjust multiple voice parameters according to different application scenarios, such as voice assistants, automatic customer service, navigation systems, etc., to ensure the applicability and effectiveness of speech synthesis in different scenarios.

[0060] As an option, the parameter configuration module can also combine emotion recognition technology to automatically adjust voice parameters according to the user's emotional state or environmental changes. For example, when the user is anxious, the system can automatically slow down the speech rate and lower the volume to achieve a soothing effect; while in a happy or exciting scenario, the system can enhance the ups and downs of the tone to make the voice more contagious.

[0061] The file generation module is used to convert the synthesized voice into a downloadable audio format file; The file generation module plays the function of converting the synthesized voice signal into an audio file that can be stored, downloaded, and played in the speech recognition and speech broadcast system. The main goal of this module is to ensure that the generated voice file has a format suitable for the playback device and appropriate encoding to meet the requirements of different devices for audio file formats. Through this module, the user can save or transmit the generated voice file in various common formats to achieve device compatibility and high-quality voice playback. The file generation module not only supports conventional audio format conversion but also has the ability to optimize storage and compression; In this embodiment, the file generation module is responsible for converting the original voice signal generated by the voice synthesis module into an audio file format. Generally, the generated voice signal may exist in the form of original waveform data, and before storage or playback, it must be converted into a standard audio format, such as MP3, WAV, AAC, etc.

[0062] In the file generation module, audio encoding and file format conversion are one of the core functions. Specifically, after the voice synthesis module generates a voice signal, the file generation module will select an appropriate audio encoding method and format according to the requirements of the user device. Common audio formats include WAV (lossless audio format), MP3 (lossy compression format), and AAC, etc. The conversion between different formats can be achieved through specialized audio conversion algorithms. The specific conversion process is as follows: F output = C(F input , P) where: F output is the output audio file; F input is the input original voice signal; C(·) represents the audio format conversion function; P is the encoding parameter set by the user (such as bit rate, sampling rate, etc.).

[0063] In some embodiments, a compression algorithm will be applied during the file generation process. Especially when storage space needs to be saved, the system will automatically select an appropriate compression ratio according to the requirements. The compressed file not only maintains a high voice quality but also reduces the occupancy of storage space.

[0064] In some embodiments, the file generation module will also adjust the sampling rate, bit rate, volume, number of channels, etc. of the audio. The optimization of these parameters can ensure that the voice file provides the best playback effect on different playback devices. For example, for low-bitrate audio encoding, the system will use audio enhancement technology to retain as many audio details as possible to ensure the sound quality.

[0065] Specifically, the sampling rate determines the frequency range of the audio. A higher sampling rate can retain more sound details but will increase the file size. Generally, the standard sampling rate for CD-quality audio is 44.1 kHz, and the voice synthesis system may select a lower sampling rate according to the needs. The bit rate affects the quality of the audio after compression. A higher bit rate can provide clearer audio quality. The file generation module can select appropriate bit rates and sampling rates according to different device requirements and storage limitations.

[0066] For example, for high-quality voice files, the file generation module may use the MP3 format with a sampling rate of 44.1 kHz and a bit rate of 192 kbps. For low-bandwidth environments, the system may select lower bit rates and sampling rates to ensure smooth playback of the voice file on the device.

[0067] After the file generation module completes the audio conversion, the system stores the generated audio file in a local device or cloud storage. Users can choose different storage options, such as saving the audio file as a high-quality WAV format file or compressing it into an MP3 format to save storage space. The specific storage process is as follows: F store = save(F output , storage_location) Where: F store represents the stored audio file; F output is the audio file after conversion; storage_location is the location where the file is saved, which may be local storage or cloud storage.

[0068] In some embodiments, the file generation module provides dual options of cloud storage and local storage. Cloud storage has better cross-device accessibility, and users can access their voice files anytime and anywhere, while local storage is suitable for scenarios that require higher access speeds. Users can choose the appropriate storage method according to actual needs.

[0069] The file generation module supports burning the generated voice file into a target device. The system can automatically identify the storage format of the target device and convert the audio file into a format compatible with the device. For example, if the target device is a smart home speaker or a car system, the system will convert the voice file into a format supported by the device to ensure normal playback of the voice. The types of devices supported by the system include embedded devices, smart home devices, smart speakers, mobile devices, etc.

[0070] During the actual storage process, the file generation module also performs audio sampling and quantization. Specifically, the audio signal can be represented by the sampling function f(t), and the quantization process is implemented by the quantization function Q(x). This process can be expressed as: Where: f(t) is the sampled value of the audio signal; x n is the original value of the audio signal; sample rate(n) is the sampling rate of the nth sampling point.

[0071] Next, the audio signal is quantized through the quantization function to meet the requirements of the target audio format, usually by digitally processing the amplitude of each sample: Q(x) = round(x, quantization levels) Where: Q(x) is the quantized audio signal; quantization levels represents the discrete level during quantization.

[0072] The download management module is used to provide storage and download of voice files, enabling users to download the generated voice files to the local or burn them into the target device; The download management module is an important part of the speech recognition and speech broadcast system, responsible for the storage, management, and download functions of the generated voice files. This module ensures that users can easily download the synthesized voice files and save them in different formats to meet the requirements of different devices. Through this module, users can select a suitable audio format (such as WAV, MP3, etc.) and automatically adjust the encoding method and bit rate according to the requirements of the device. In addition, the download management module also supports backup, recovery, and cross-device synchronization of voice files, providing users with convenient management and access functions.

[0073] In this embodiment, the main function of the download management module is to store, convert the format, download, and manage the voice files generated by the speech synthesis module. The module allows users to select different audio formats and perform encoding conversion to ensure the compatibility of the audio files with the target device. Generally, this module supports common audio file formats such as WAV, MP3, AAC, etc., and users can select a suitable format for download or storage according to their needs.

[0074] In some embodiments, the download management module provides a format selection function, allowing users to select the storage format of the audio file according to the device requirements. For example, mobile devices may prefer to select a smaller compression format (such as MP3), while high-fidelity devices may prefer to use a lossless audio format (such as WAV). To ensure the quality of the audio file, the system also provides a customization function for encoding parameters such as bit rate, sampling rate, number of channels, etc., which users can adjust according to actual needs.

[0075] In this embodiment, the download management module can automatically convert the generated voice files into the target format through the built-in audio conversion tool. For example, the system may convert the synthesized original WAV file into a smaller MP3 format or convert it into AAC format to improve the compression efficiency. The specific conversion process is as follows: F output = C(F input , P) Where: F output is the target audio file; F inputis the originally generated voice signal; C(·) is the format conversion function; P is the encoding parameter defined by the user, such as bit rate, sampling rate, etc.

[0076] In addition, the download management module also supports compressing audio files to reduce the storage space of the files. Generally, the compression methods of audio files include lossless compression (such as FLAC) and lossy compression (such as MP3). In some embodiments, the system will automatically select the appropriate compression method according to the usage scenario of the file.

[0077] In some embodiments, the download management module also includes cloud storage support and multi-device synchronization functions. The voice files generated by the user can be uploaded to cloud storage to ensure sharing and synchronization among multiple devices. Through cloud storage, users can access the downloaded audio files on any device, whether it is a mobile phone, a PC, a smart speaker or a vehicle-mounted device.

[0078] Specifically, the download management module provides the following functions: File backup: Upload the generated voice files to the cloud to ensure the security and persistence of the files; File recovery: Support recovering voice files from the cloud to ensure that users can still recover voice files in case of device replacement or loss; Cross-device synchronization: Ensure that users can access consistent voice files on different devices. For example, users can download voice files on their mobile phones and play them on a smart speaker, and all devices can synchronously update the files.

[0079] In some embodiments, the download management module automatically synchronizes the audio files between the cloud and the local device through integration with the cloud service platform. Users do not need to intervene manually, and the file synchronization will be automatically completed after the user generates the voice file.

[0080] For the audio files stored locally, users can save them to a computer or other devices through the download management module. When the user selects to download, the system will automatically adjust the file format according to the storage format of the file and the requirements of the target device to ensure that the voice file can be played smoothly. For example, the system will automatically select the most suitable file format and encoding method during the download process to improve the compatibility of the playback device.

[0081] In the specific implementation, the download management module will check the storage type and file format support of the user device and adjust the file conversion process according to the detection results to ensure file compatibility and the best playback effect. The basic steps of this process are as follows: F download = D(F store ,S) where: F downloadis the downloaded audio file; D(·) is the download function; F store is the stored audio file; S is the storage and format requirements of the target device.

[0082] In the download management module, the core of file format conversion, audio compression, and synchronization functions lies in optimizing the storage and transmission efficiency of audio files. For each audio file, the system needs to perform bitrate control and file size estimation to ensure that the generated audio file meets both the device compatibility requirements and is reasonably optimized in terms of storage. For bitrate control, the system selects the most appropriate bitrate B according to the target file size and audio quality requirements set by the user and dynamically adjusts it according to the file format. The calculation formula for the bitrate is as follows: where: B is the bitrate, in kbps; S target is the target file size, in bytes (Byte); T file is the duration of the audio file, in seconds.

[0083] According to the target bitrate, the system will select the corresponding encoding method for compression, thereby reducing the size of the audio file on the basis of ensuring audio quality and optimizing the storage and transmission process.

[0084] Please refer to the appendix Figure 2 , the present invention also provides a method for speech recognition and speech broadcast, including the following steps: S1. Receive the text information input by the user; In this embodiment, the system receives the text information input by the user through the text input module. This text information can be obtained in various ways, such as manual input, speech-to-text conversion, or import from an external data source. The system ensures that regardless of the input method, all text data can be accurately received and prepared for subsequent processing. The text input module is responsible for cleaning, standardizing, and formatting the input text for subsequent speech synthesis and optimization steps; S2. Generate a target speech signal using the speech synthesis module according to the input text; After the text is input, the system passes it to the speech synthesis module. In this module, the text will go through the processes of encoding and speech generation and finally be converted into a speech signal. The speech synthesis module uses advanced technologies such as deep neural networks and generative adversarial networks (GANs) to generate high-quality speech signals that meet the text content and speech requirements. The generated speech signal will extract language features, adjust emotions, and generate speech styles according to the input text content; S3. Optimize the pitch, speech rate, volume, tone, and brightness parameters of the speech according to the generated speech signal to improve the naturalness and emotional expression of the speech; After the target voice signal is generated, the system enters the voice optimization stage. The voice synthesis module makes detailed adjustments to the generated voice signal through an optimization algorithm. Specifically, the system adjusts multiple parameters such as pitch, speech rate, volume, tone, and brightness to ensure that the generated voice is more natural, fluent, and has appropriate emotional expression. For example, pitch can be used to represent the emotional changes in speech, speech rate can control the clarity and rhythm of speech, and volume and brightness have important impacts on the expression effect and audibility of speech. Through these optimizations, the voice signal will better meet the user's expectations; S4. Convert the synthesized voice signal into a downloadable audio file; Once the voice signal is optimized, the system converts it into a standard audio format (such as WAV, MP3, AAC, etc.) for the user to download and use. The file generation module is responsible for selecting the appropriate encoding method, bit rate, and audio format according to the requirements of the target device. This process involves format conversion, compression, and parameter setting of the audio file to ensure that the final file has both good sound quality and can be played compatibly on different devices; S5. Provide the user with the generated audio file for download, for the user to save or burn to the target device; Finally, the generated audio file will be provided to the user for download. The user can save the file to local storage or cloud storage, or burn it to the target device (such as a smart speaker, mobile phone, in-vehicle system, etc.) as needed for playback. The download management module provides storage options, format selection, and download links for this purpose to ensure that the user can conveniently obtain and use the generated voice file according to the device requirements.

[0085] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for speech recognition and speech broadcasting, characterized in that: include: A text input module, used to receive text content input by a user; The speech synthesis module is used to generate corresponding speech signals according to the text input by the user; Parameter configuration module, used to allow users to customize voice parameters, including timbre, speaking speed, volume, pitch and brightness; A file generation module, used for converting the synthesized speech into a downloadable audio format file; The download management module is used to provide storage and downloading of voice files, so that users can download the generated voice files to the local or burn them to the target device.

2. The system for speech recognition and speech broadcasting according to claim 1, characterized in that: The speech synthesis module comprises: A text encoding unit, used to encode input text information and generate an intermediate representation; A speech generation unit, used to generate a corresponding speech signal according to the encoded text information; The optimization unit is used to adjust the pitch, speaking speed, volume, tone and brightness parameters according to the characteristics of the generated speech to optimize the naturalness and emotional expression of the speech.

3. The system for speech recognition and speech broadcasting according to claim 1, characterized in that: The speech synthesis module is optimized by a generative adversarial network model, and the generative adversarial network model includes: A generator for generating synthetic speech that is as close to real speech as possible; The discriminator is used to distinguish between generated speech and real speech, and continuously improve the quality of synthesized speech through adversarial training.

4. The system for speech recognition and speech broadcasting according to claim 3, characterized in that: The speech synthesis module optimizes the generated speech signal to minimize the difference between it and the target speech signal by minimizing the Kullback-Leibler divergence, which is defined as: Among them, D KL (P||Q) is the Kullback-Leibler divergence; P(x) is the probability distribution of the target speech signal, which represents the characteristic distribution of the target speech signal; Q(x) is the probability distribution of the generated speech signal, which represents the characteristic distribution of the synthesized speech signal; x is the sample point or feature of the audio signal, which represents the state or feature of the speech waveform at a specific moment.

5. The system for speech recognition and speech broadcasting according to claim 1, characterized in that: The download management module allows the user to save the generated voice file in multiple formats, and supports selecting a suitable encoding method and bit rate according to device requirements to ensure the compatibility and playback quality of the voice file on different hardware devices.

6. The system of speech recognition and speech broadcasting according to claim 1, characterized in that: The parameter configuration module provides a real-time preview function, and the user can listen to the generated effect in real time after adjusting the timbre, speaking speed, volume, pitch and brightness parameters to optimize the final voice output.

7. The system of speech recognition and speech broadcasting according to claim 1, characterized in that: The file generation module includes a voice file format conversion function, which can automatically select the optimal file format according to the compatibility of the target device and compress or optimize the file to improve storage efficiency and transmission quality.

8. The system of speech recognition and speech broadcasting according to claim 1, characterized in that: The download management module supports local storage and cloud storage options, and provides voice file management, backup and recovery functions to ensure that users can access and use stored voice files at any time.

9. The system for speech recognition and speech broadcasting according to claim 1, characterized in that: The speech synthesis module further combines multiple speech models, including but not limited to natural human speech, robot speech, and dialect speech, and allows users to select or customize speech models to adapt to different scenario requirements.

10. A method for speech recognition and speech broadcasting, applied to the speech recognition and speech broadcasting system as claimed in any one of claims 1 to 9, characterized in that: The following steps are involved: Receive text information input by the user; Based on the input text, a speech synthesis module is used to generate a target speech signal; Based on the generated speech signal, the pitch, speaking speed, volume, tone and brightness parameters of the speech are optimized to improve the naturalness and emotional expression of the speech; Convert the synthesized speech signal into a downloadable audio file; Provide users with the ability to download generated audio files for saving or burning to a target device.

Citation Information

Patent Citations

  • Speech synthesis method, device, system and equipment

    CN109147760A

  • Voice broadcast method, device, equipment and medium

    CN113066474A

  • Voice tone conversion method and system

    CN116741144A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN118711560A

  • Intelligent robot speech synthesis method based on deep learning

    CN119446117A

Cited By

  • Intelligent voice interaction and motion control method and system

    CN122637780A