Voice processing method and device, electronic equipment and storage medium

By extracting speech features and content from audio data and combining them with a speech synthesis model, new media stream data is generated, which solves the problem of audio data distortion after enhancement processing and improves auditory experience and adaptability.

CN120998214APending Publication Date: 2025-11-21GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410626390.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The speech content in the existing enhanced audio data is distorted, resulting in a poor auditory experience for the audience.

Method used

By extracting speech features and content from audio data, new media stream data is generated, preserving the speech features and content of the language expression while filtering out noise and echo. A pre-trained speech synthesis model is then used to synthesize the speech features and repaired content.

Benefits of technology

It enhances the auditory experience of audio audiences with media stream data, ensures the intelligibility, naturalness, and comfort of language expression, reduces data transmission volume, and adapts to complex and ever-changing real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998214A_ABST
    Figure CN120998214A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice processing method and device, electronic equipment and a storage medium. The method comprises the following steps: performing voice feature extraction and voice content extraction on original to-be-processed first media stream data to obtain voice feature information and voice content information related to language expression in the first media stream data, and then generating new second media stream data on the basis of the extracted voice feature information and voice content information, when the second media stream data is played, voice features and voice content expressed by languages in the first media stream data can be completely reserved, meanwhile, ambient sounds such as noise and echoes possibly appearing in the first media stream data are filtered out, and for sound audiences of the media stream data, the voice audiences of the media stream data are effectively protected while interference of external sounds on the language expressions is eliminated. The intelligibility, the naturalness and the comfort level of the language expression part are ensured, and the auditory feeling of sound audiences on the media stream data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to voice processing methods, apparatus, electronic devices, and storage media. Background Technology

[0002] Using electronic devices to collect data while a user speaks to generate media stream data, and then recording the spoken content from the audio data in the media stream data, is a common data processing method based on information technology to achieve information recording, interpersonal communication, and information transmission.

[0003] In real-life scenarios, users' speaking environment is likely to be affected by various noises, echoes, and other ambient sounds. In order to improve the auditory presentation of the spoken content to the audience when playing media stream data, the audio data is usually enhanced when generating media stream data with the spoken content as the recording object (or one of the recording objects). That is, the ambient sounds are suppressed or eliminated so that the audio data retains only the recording of the spoken content as much as possible.

[0004] When the inventors tested the playback of media stream data enhanced using existing technology, they found that suppressing or eliminating ambient noise caused data loss in the part of the audio data that recorded the speech content. This resulted in the actual playback of the speech content being distorted compared to the original natural speech, and the audience had a poor auditory experience of the enhanced speech content. Summary of the Invention

[0005] This invention provides a speech processing method, apparatus, electronic device, and storage medium to solve the technical problem that the actual spoken content played in the enhanced audio data is distorted compared to the original natural speech, resulting in a poor auditory experience for the audience.

[0006] In a first aspect, embodiments of this application provide a speech processing method, which includes:

[0007] First information extraction and second information extraction are performed on the first media stream data containing audio data to obtain the speech feature information and speech content information of the audio data, respectively.

[0008] The speech feature information and speech content information are input into a preset speech synthesis model to generate second media stream data. When the second media stream data is played, the speech content corresponding to the speech content information is presented based on the speech features corresponding to the speech feature information.

[0009] The above describes the process of extracting speech features and speech content from the original first media stream data to be processed, obtaining speech feature information and speech content information related to language expression in the first media stream data. Then, based on the extracted speech feature information and speech content information, a new second media stream data is generated. When the second media stream data is played, it can completely preserve the speech features and speech content of the language expression in the first media stream data, while filtering out environmental noise, echoes and other sounds that may appear in the first media stream data. For the audio audience of the media stream data, while eliminating the interference of external sounds on the language expression, it ensures the intelligibility, naturalness and comfort of the language expression, and improves the audio audience's auditory experience of the media stream data.

[0010] The second information extraction is performed using a pre-trained speech recognition model, and the speech content information includes speech recognition text.

[0011] As described above, the first media stream data is processed by a speech recognition model to obtain speech-recognized text. Subsequently, the second media stream data can be generated based on the speech-recognized text, enabling rapid processing of the content expressed in language.

[0012] The second information extraction is performed using a pre-trained speech coding model, and the speech content information includes the speech coding vector.

[0013] The first media stream data is encoded using a pre-trained speech coding model to obtain a speech coding vector. Subsequently, the second media stream data can be generated based on the speech coding vector, enabling rapid processing of the content of the language expression.

[0014] The first media stream data contains only audio data;

[0015] Accordingly, the speech feature information and speech content information are input into a preset speech synthesis model to generate second media stream data, including:

[0016] Based on a pre-trained speech restoration model, the speech content information is restored to obtain the restored content information.

[0017] The speech feature information and the repaired content information are input into a preset speech synthesis model to generate second media stream data.

[0018] As mentioned above, in response to potential issues such as stuttering, reduplicated words, and misspellings in human speech, the first media stream data, which consists only of audio data, can be repaired using a speech restoration model to obtain restored content information that enables fluent and accurate speech. Based on the speech feature information and the restored content information, a second media stream data is generated. For the audio audience of the media stream data, this eliminates potential speech errors that may occur when the speaker is expressing themselves, further enhancing the auditory experience of the audio audience.

[0019] The repair includes correcting repeated characters and correcting misspellings.

[0020] The above-mentioned repairs to reduplicated characters and misspelled characters are sufficient to basically guarantee the correction of possible expression errors.

[0021] The process includes extracting first and second information from the first media stream data containing audio data to obtain the speech feature information and speech content information of the audio data, respectively.

[0022] The speech feature information and speech content information are sent to a remote device, which then inputs the speech feature information and speech content information into a preset speech synthesis model to generate a second media stream data.

[0023] As mentioned above, when processing media stream data in a remote call scenario, voice feature information and voice content information can be sent directly to the remote device making the call. The remote device can then directly generate the second media stream data. Compared to directly transmitting audio data, transmitting voice feature information and voice content information can reduce the amount of data transmitted and lower the network bandwidth requirements in a remote call scenario.

[0024] Specifically, the first media stream data containing audio data undergoes first information extraction and second information extraction to obtain speech feature information and speech content information of the audio data, respectively, including:

[0025] First information extraction is performed on the starting segment of the first media stream data containing audio data to obtain the speech feature information of the audio data;

[0026] The second information extraction is continuously performed on the first media stream data containing audio data to obtain the voice content information of the audio data.

[0027] The above-mentioned method of extracting first information from the starting segment of the first media stream data containing audio data and using the speech feature information of the starting segment as the feature information of the corresponding media data can effectively reduce the amount of data processing in the speech feature information extraction part.

[0028] Among them, speech feature information includes one or more of voiceprint feature information, rhythm feature information, and emotion feature information.

[0029] As mentioned above, speech feature information includes multi-dimensional feature information, which can meet the different application needs of speakers.

[0030] Secondly, embodiments of this application provide a voice processing apparatus, which includes:

[0031] The information extraction unit is used to perform first information extraction and second information extraction on the first media stream data containing audio data, and respectively obtain the speech feature information and speech content information of the audio data.

[0032] The data generation unit is used to input speech feature information and speech content information into a preset speech synthesis model to generate second media stream data. When the second media stream data is played, the speech content corresponding to the speech content information is presented based on the speech features corresponding to the speech feature information.

[0033] The second information extraction is performed using a pre-trained speech recognition model, and the speech content information includes speech recognition text.

[0034] The second information extraction is performed using a pre-trained speech coding model, and the speech content information includes the speech coding vector.

[0035] The first media stream data contains only audio data;

[0036] Correspondingly, the data generation unit includes:

[0037] The information repair module is used to repair speech content information based on a pre-trained speech repair model to obtain repaired content information.

[0038] The data generation module is used to input speech feature information and repair content information into a preset speech synthesis model to generate second media stream data.

[0039] The repair includes correcting repeated characters and correcting misspellings.

[0040] The voice processing device also includes:

[0041] The information sending unit is used to send voice feature information and voice content information to a remote device, so that the remote device can input the voice feature information and voice content information into a preset voice synthesis model to generate a second media stream data.

[0042] The information extraction unit includes:

[0043] The first extraction module is used to extract first information from the starting segment of the first media stream data containing audio data to obtain the speech feature information of the audio data.

[0044] The second extraction module is used to continuously extract second information from the first media stream data containing audio data to obtain the voice content information of the audio data.

[0045] Among them, speech feature information includes one or more of voiceprint feature information, rhythm feature information, and emotion feature information.

[0046] Thirdly, embodiments of this application also provide an electronic device, which includes:

[0047] One or more processors;

[0048] Memory, used to store one or more computer programs;

[0049] When one or more computer programs are executed by one or more processors, an electronic device enables the speech processing method as described in the first aspect.

[0050] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech processing method as described in the first aspect. Attached Figure Description

[0051] Figure 1 This is a flowchart of a speech processing method provided in an embodiment of this application.

[0052] Figure 2 This is a schematic diagram of the data processing process of a speech processing method provided in an embodiment of this application.

[0053] Figure 3 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application.

[0054] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0055] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not the entire structure.

[0056] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art should be able to conceive after reading this application specification that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.

[0057] The embodiments of the present invention will be described in detail below.

[0058] Currently, methods for enhancing media stream data that record spoken content can be broadly categorized into time-domain methods, frequency-domain methods, and deep learning methods. Time-domain methods can be further divided into parametric filtering methods and signal subspace methods, while frequency-domain methods include spectral subtraction, Wiener filtering, and speech amplitude spectrum estimation methods based on minimum mean square error.

[0059] Existing audio enhancement schemes each have their own limitations. For example, time-domain and frequency-domain methods are poor at suppressing non-stationary burst noise, such as sudden car horns on the road. Furthermore, traditional algorithms often leave residual noise after enhancement, leading to poor subjective listening experience and even affecting the intelligibility of the speech information. Moreover, time-domain and frequency-domain methods rely heavily on assumptions during processing, limiting their effectiveness and making them unsuitable for complex and varied real-world scenarios. When media streams processed by these methods are played back, the "subtraction" process, which targets noise reduction in the overall audio signal, inevitably results in data loss of the spoken content. This leads to distortion of the played speech compared to the original natural speech, resulting in a poor auditory experience for the audience.

[0060] To address the above technical issues, this application proposes a speech processing method that extracts speech features and speech content from the original first media stream data to be processed, obtaining speech feature information and speech content information related to language expression in the first media stream data. Then, based on the extracted speech feature information and speech content information, a new second media stream data is generated. When the second media stream data is played, it can completely retain the speech features and speech content of the language expression in the first media stream data, while filtering out ambient noise, echoes, and other sounds that may appear in the first media stream data. For the audio audience of the media stream data, while eliminating the interference of external sounds on the language expression, it ensures the intelligibility, naturalness, and comfort of the language expression, thereby improving the audio audience's auditory experience of the media stream data.

[0061] Figure 1This application provides a flowchart of a voice processing method, applicable to various electronic devices capable of processing media stream data containing audio data. The method is implemented by electronic devices, specifically, various electronic devices that simultaneously record (and / or transmit) audio and video data, and electronic devices that record (and / or transmit) audio data. The specific form of the electronic device can be a mobile terminal, personal computer, interactive whiteboard, server, etc. Figure 1 As shown, the speech processing method includes, but is not limited to, steps S110-S120:

[0062] Step S110: Perform first information extraction and second information extraction on the first media stream data containing audio data to obtain the speech feature information and speech content information of the audio data, respectively.

[0063] Media streaming generally refers to the technology that uses streaming transmission to receive and play streaming media over a network. Specific data formats mainly include audio data, video data, and multimedia data containing both. Audio, video, or multimedia files transmitted over a network are not downloaded in their entirety before playback; the data stream is transmitted and played in real time, with only a slight initial delay. As an important form of information recording and transmission, media streaming primarily allows the audience to receive information through hearing and / or sight during playback.

[0064] In this application embodiment, the processed media stream data is not limited to specific application scenarios and application stages. Any data that can be transmitted and played in the form of a media stream can be implemented using this application embodiment. For example, it can be implemented during data acquisition, or when a server or other electronic device stores media stream data that can be used to implement this solution, or it can be implemented before playback after receiving media stream data by an electronic device. There are no specific limitations. In this application embodiment, the optimization focuses on the part of the media stream data playback process where the audience receives information through hearing. Therefore, it does not process all types of media stream data, but only media stream data containing audio data. Here, media stream data containing audio data is defined as first media stream data. That is, the data object processed in this application embodiment includes not only media stream data containing audio data, but also multimedia data containing both audio and video data. First media stream data can specifically be data acquired in real time, data acquired and sent by a data acquisition device and then centrally processed, or data received from other electronic devices and requiring local processing for playback. All of the above or any other data that needs to be processed through this application embodiment belongs to the original first media stream data to be processed.

[0065] The specific content of audio data recording may be the collection of speech, or it may be the collection of various sounds such as pure music, natural sounds, or other noise. Different recorded content leads to different ways of perceiving information for the audience. For example, the audio data obtained from the collection of speech actually conveys information through the recorded spoken content, and the core of information transmission is text. In the specific process of collecting audio data, the collection scenario for speech is also the most complex and variable collection scenario. For example, pure music usually has a dedicated recording environment to eliminate noise as much as possible at the sound source; the collection of natural sounds usually takes a rich variety of sounds as the object of collection; while the sound collection of speech has a clear sound collection target, but in reality, there may be various external sound interferences. Compared with the sound of speech, other sounds can be regarded as noise.

[0066] The processing procedure for the first media stream data in this embodiment is designed based on the conclusions drawn from the analysis of the basic components for achieving good information transmission when the audience receives the speech content in the audio data at the auditory level. First, in real, natural interpersonal communication, the core of the transmitted information is the speech content; therefore, the design involves extracting information from the first media stream data (i.e., second information extraction) to obtain speech content information. Second, accents, timbre, emotion, and rhythm in the speech expression process can help to express information more richly and accurately; therefore, the design involves extracting information from the first media stream data (i.e., first information extraction) to obtain speech feature information. In other words, speech content information is the information extracted from the speech content in the first media stream data, and speech feature information is the information extracted from the speech pronunciation features in the first media stream data.

[0067] In the specific extraction of the first information, speech features in the first media stream data can be extracted using a pre-trained feature extraction model or a zero-shot system. The first feature extraction model can be trained using a training set consisting of a large amount of speech data. The training principle of the first feature extraction model is largely the same as that of other feature extraction models in the field of network models; training can be performed accordingly and will not be elaborated here. Regardless of the method used to extract speech features, the extraction dimensions of the speech features can be constrained in the corresponding feature extraction model or zero-shot system. Corresponding to the speech expression-related elements such as timbre, rhythm, and emotion mentioned earlier, speech features include one or more of the following: voiceprint features, rhythm features, and emotional features. Specific methods for recording speech features include, for example, Mel spectrograms.

[0068] Voice feature information includes multi-dimensional features that can meet the different application needs of speakers. For example, if a user's Mandarin is not standard, and the electronic device can accurately recognize the corresponding voice content, the user can set the processing mode when communicating with others, choosing to retain only timbre, emotion, and rhythm. Subsequently, when generating second media stream data, the media stream data generated according to the user's timbre, emotion, and speaking rhythm, i.e., the second media stream data mentioned later, is equivalent to the user speaking directly with an accent that is relatively difficult for other users to understand, presenting a more standard expression of the user's accent to other users. Thus, even if the speaker's own heavy accent hinders communication, it can improve the auditory experience of the audience, especially improving intelligibility. Moreover, using the actual speaker's own speaking characteristics has better naturalness and comfort, making the other party feel that they are speaking to a real person, rather than an emotionless machine.

[0069] In the specific process of extracting the second information, a pre-trained speech recognition model can be used, with the speech content information including speech-recognized text. The first media stream data is processed by the speech recognition model to obtain speech-recognized text. At this point, the speech content information is presented to the user in text form, which is the text corresponding to the spoken words. Subsequently, the second media stream data can be generated based on the speech-recognized text, achieving rapid processing of the content of the spoken language. Alternatively, the second information extraction can be performed using a pre-trained speech coding model, with the speech content information being a speech coding vector. The first media stream data is encoded using the pre-trained speech coding model to obtain a speech coding vector. At this point, the speech content information is presented to the user in text form as garbled text or numbers that the user cannot understand. Subsequently, the second media stream data can be generated based on the speech coding vector, achieving rapid processing of the content of the spoken language. Extracting content information from speech using a speech recognition model or a speech coding model can employ relevant technologies in the field of speech-to-text conversion, which will not be elaborated upon here.

[0070] One possible implementation for information extraction is to continuously extract the first and second information based on the ongoing state of the first media stream data, thereby continuously obtaining the speech feature information and speech content information recorded on the timeline of the first media stream. This ensures a high degree of fidelity to the spoken content when generating the second media stream data, especially for multimedia data that combines audio and video data, and ensures that the final second media stream data is synchronized with the audio and video during playback.

[0071] Another optional implementation for information extraction is to perform first information extraction on the starting segment of the first media stream data containing audio data to obtain the speech feature information of the audio data; and then continuously perform second information extraction on the first media stream data containing audio data to obtain the speech content information of the audio data. Performing first information extraction on the starting segment of the first media stream data containing audio data, and using the speech feature information of the starting segment as the feature information of the corresponding media data, can effectively reduce the data processing volume of the speech feature information extraction part. The starting segment refers to a section of the first media stream data at the beginning, such as the segment within 3 seconds after the start; the specific length is not limited, as long as it is sufficient to extract speech feature information from it.

[0072] Whether it is continuous extraction or only the initial segment is extracted, either of the speech feature information extraction methods and speech content information extraction methods described above can be used.

[0073] It should be noted that in the specific implementation process, considering that the first media stream data may last for a long time, during which the speaker's voice may change, in order to ensure accurate processing of the corresponding speech content of each speaker in the first media stream data, if only the starting segment is processed, then each time the speech is interrupted and restarted, it is treated as a new starting segment for the first information extraction. This processing method ensures that when there is a speaker switch, the processing is based on the speech characteristics of the switched speaker. For example, in a remote conference scenario, if multiple people participate in a meeting in a real conference room and freely speak and participate in the discussion, by reconfirming the starting segment, the speech characteristics of the actual speaker can be extracted each time the speaker switches, and the second media stream data with the actual speaker characteristics can be generated accordingly. Even if the speaker remains the same, the speaker may change emotions during the communication process, and reconfirming the starting segment can also make the final generated second media stream data present a more realistic expression when played. Moreover, the first media stream data containing audio data does not necessarily contain a record of the speech content. If the extraction of speech feature information and / or speech content information fails, no further processing is required.

[0074] Step S120: Input the speech feature information and speech content information into a preset speech synthesis model to generate second media stream data. When the second media stream data is played, the speech content corresponding to the speech content information is presented based on the speech features corresponding to the speech feature information.

[0075] In one optional implementation, the speech synthesis model can be based on TTS (Text-to-Speech). The TTS-based speech synthesis model can be broadly divided into two stages: language analysis and sound synthesis. The language analysis stage mainly analyzes the input text information to generate a corresponding linguistic specification, determining how to pronounce it. The sound synthesis stage mainly generates the corresponding audio based on the linguistic specification provided by the speech analysis stage, realizing the function of speech production. The language analysis stage mainly includes processing information such as text structure, language identification, text standardization, text-to-phoneme conversion, and prosody corresponding to the speech content information. Specific processing details can be found in related technical implementations. The sound synthesis stage mainly involves retrieving the waveform sequences corresponding to the phonemes obtained in the language analysis stage from a pre-prepared database covering all syllable phonemes. Decoding and playing these waveform sequences at this stage is sufficient to present the speech content information. In the acoustic synthesis stage, the waveform sequence is further adjusted based on the speech feature information. This allows the adjusted waveform sequence to present the speech content corresponding to the speech content information during decoding and playback, using the speech features corresponding to the speech feature information. For specific adjustments to the speech features in the waveform file, please refer to the implementation of relevant audio adjustment technologies.

[0076] In another alternative implementation, a neural network model can be constructed and multiple samples can be collected. Each sample includes speech feature information and speech content information, as well as standard speech content. Using the speech feature information and speech content information from the samples as input and the corresponding standard speech content as the desired output, the neural network model is trained, ultimately yielding an end-to-end speech synthesis model. The internal structure of the speech synthesis model can be considered a black box. The speech synthesis model obtained by training the neural network model based on samples does not require extensive linguistic knowledge for text structure, grammar, and other processing during implementation, thus improving the adaptability of the speech processing method in this embodiment and reducing development costs.

[0077] In this embodiment, based on the data processing foundation described above, it is equivalent to generating text content and various elements of speech synthesis in real time. On this basis, speech synthesis is performed to obtain new audio data. Depending on the type of the first media stream data, the new audio data is directly used as the second media stream data, or it is combined with the video data in the first media stream data to obtain the second media stream data. Based on the speech feature information and speech content information obtained from the first media stream data, when the second media stream data is played, the speech content corresponding to the speech content information is presented with the speech features corresponding to the speech feature information.

[0078] It is important to note that, in the specific processing, if the first media stream data contains video data, in order to ensure audio-visual synchronization as much as possible, when extracting speech feature information and / or speech content information, the time information of the extracted speech feature information and / or speech content information in the first media stream data can be recorded. Subsequently, when generating the second media stream data, the speech content in the regenerated audio data can be in the corresponding time period as much as possible with the content recorded in the first media stream data, thereby ensuring audio-visual synchronization as much as possible.

[0079] In the specific process of generating the second media stream data, if the first media stream data only contains audio data, the speech content information can be repaired based on a pre-trained speech restoration model to obtain restored content information. The second media stream data is then generated based on the speech feature information and the restored content information. Addressing potential stuttering, reduplicated words, and misspellings in human speech, the first media stream data, composed only of audio data, can be repaired using a speech restoration model to obtain restored content information that enables fluent and accurate pronunciation. The second media stream data is then generated based on the speech feature information and the restored content information. For the audio audience, this eliminates potential errors in speech expression, further enhancing their auditory experience. In fact, during the second information extraction process, because the extracted information is directly text-related, stuttering and pauses are not recorded in the speech content information; these errors are essentially eliminated naturally. Repair at this point can include reduplicated word repair and misspelling repair. That is, repairing reduplicated words and misspellings that cannot be naturally eliminated during information extraction can essentially guarantee the repair of potential expression errors.

[0080] For users looking to improve their speaking skills, this method allows them to compare their actual speaking ability with their ideal ability, thus identifying areas for improvement and setting goals. For example, for users who stutter, the system can output two results from the audio data of a spoken passage: the first is the raw, unedited media stream data, and the second is the improved version. This provides a visual representation of a higher level of speaking ability, even if their actual ability falls short. See the detailed explanation below. Figure 2, assume that the original voice content to be recorded is the reading of "The sun along the mountain bows", but during the reading, it is expressed as "White ~ day ~ day ~ leans on ~ mountain ~ mountain ~ ends". When processing this first media stream data consisting only of audio data, the first information extraction can obtain the voice feature information (i.e., feature vector A), and the second information extraction can obtain the voice recognition text of "White day day leans on mountain mountain ends" as the voice content information. Repair the voice recognition text to obtain the repaired content information of "The sun along the mountain bows". At this time, re-synthesize to obtain a more standard expression of "White ~ day ~ leans on ~ mountain ~ ends". If the first media stream data and the second media stream data are output in two times, the user can clearly recognize the height that their expression ability can reach.

[0081] As described above, perform voice feature extraction and voice content extraction on the original first media stream data to be processed, obtain the voice feature information and voice content information related to language expression in the first media stream data. Then, based on the extracted voice feature information and voice content information, generate a new second media stream data. When the second media stream data is played, it can completely retain the voice features and voice content of the language expression in the first media stream data, and at the same time filter out the environmental sounds such as noise and echo that may appear in the first media stream data. For the sound audience of the media stream data, while eliminating the interference of external sounds on language expression, it ensures the intelligibility, naturalness and comfort of the language expression part, and improves the auditory experience of the sound audience for the media stream data. Compared with the existing enhancement processing method with the overall processing idea of "subtraction", this solution designs another technical idea, only extracts the information related to voice for audio synthesis, while achieving a good noise reduction effect, it also retains the proper speaking tone as much as possible and highly restores the pure voice.

[0082] In the specific implementation process, in a remote call scenario, the first media stream data containing audio data can undergo first information extraction and second information extraction to obtain the voice feature information and voice content information of the audio data, respectively. Then, the voice feature information and voice content information are sent to the remote device, which then inputs them into a preset speech synthesis model to generate the second media stream data. When processing media stream data in a remote call scenario, the voice feature information and voice content information can be directly sent to the remote device conducting the remote call, and the remote device can then directly generate the second media stream data. Compared to directly transmitting audio data, transmitting voice feature information and voice content information reduces data transmission volume and lowers the network bandwidth requirements in remote call scenarios. The electronic device used to implement the speech processing method in the embodiments of this application can extract speech feature information and speech content information from the first media stream data and then generate the second media stream data locally; or it can extract speech feature information and speech content information from the first media stream data and then send them to other electronic devices, which will then generate the second media stream data; or it can receive speech feature information and speech content information sent by other electronic devices and then synthesize and play them locally.

[0083] Figure 3 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Figure 3 As shown, the voice processing device includes an information extraction unit 210 and a data generation unit 220.

[0084] The information extraction unit 210 is used to perform first information extraction and second information extraction on the first media stream data containing audio data, and respectively obtain the speech feature information and speech content information of the audio data; the data generation unit 220 is used to input the speech feature information and speech content information into a preset speech synthesis model to generate second media stream data, and when the second media stream data is played, the speech content corresponding to the speech content information is presented with the speech features corresponding to the speech feature information.

[0085] Based on the above embodiments, the second information extraction is performed through a pre-trained speech recognition model, and the speech content information includes speech recognition text.

[0086] Based on the above embodiments, the second information extraction is performed through a pre-trained speech coding model, and the speech content information includes speech coding vectors.

[0087] Based on the above embodiments, the first media stream data only contains audio data;

[0088] Correspondingly, the data generation unit 220 includes:

[0089] The information repair module is used to repair speech content information based on a pre-trained speech repair model to obtain repaired content information.

[0090] The data generation module is used to input speech feature information and repair content information into a preset speech synthesis model to generate second media stream data.

[0091] Based on the above embodiments, the repair includes the repair of repeated characters and the repair of misspelled characters.

[0092] Based on the above embodiments, the voice processing device further includes:

[0093] The information sending unit is used to send voice feature information and voice content information to a remote device, so that the remote device can input the voice feature information and voice content information into a preset voice synthesis model to generate a second media stream data.

[0094] Based on the above embodiments, the information extraction unit 210 includes:

[0095] The first extraction module is used to extract first information from the starting segment of the first media stream data containing audio data to obtain the speech feature information of the audio data.

[0096] The second extraction module is used to continuously extract second information from the first media stream data containing audio data to obtain the voice content information of the audio data.

[0097] Based on the above embodiments, the speech feature information includes one or more of the following: voiceprint feature information, rhythm feature information, and emotion feature information.

[0098] The voice processing device provided in this application embodiment is included in an electronic device and can be used to execute the corresponding voice processing method provided in the above embodiment, and has corresponding functions and beneficial effects.

[0099] It is worth noting that in the embodiments of the above-mentioned voice processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0100] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device includes a processor 310 and a memory 320, and may also include an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device may be one or more. Figure 4Taking a processor 310 as an example; the processor 310, memory 320, input device 330, output device 340, and communication device 350 in the electronic device can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0101] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice processing method in the embodiments of this application. The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby implementing the aforementioned voice processing method.

[0102] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0103] Input device 330 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 340 may include display devices such as a display screen.

[0104] The aforementioned electronic device includes a voice processing unit, which can be used to execute any voice processing method and has corresponding functions and beneficial effects.

[0105] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program performs relevant operations in the speech processing method provided in any embodiment of this application and has corresponding functions and beneficial effects.

[0106] Those skilled in the art will understand that embodiments of this application may be provided as methods, systems, or computer program products.

[0107] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce implementations of the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0108] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0109] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0110] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0111] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A speech processing method, characterized in that, include: First information extraction and second information extraction are performed on the first media stream data containing audio data to obtain the speech feature information and speech content information of the audio data, respectively. The speech feature information and speech content information are input into a preset speech synthesis model to generate second media stream data. When the second media stream data is played, the speech content corresponding to the speech content information is presented based on the speech features corresponding to the speech feature information.

2. The speech processing method according to claim 1, characterized in that, The second information extraction is performed using a pre-trained speech recognition model, and the speech content information includes speech recognition text.

3. The speech processing method according to claim 1, characterized in that, The second information extraction is performed using a pre-trained speech coding model, and the speech content information includes a speech coding vector.

4. The speech processing method according to any one of claims 1-3, characterized in that, The first media stream data contains only audio data; Accordingly, the step of inputting the speech feature information and speech content information into a preset speech synthesis model to generate second media stream data includes: Based on a pre-trained speech restoration model, the speech content information is restored to obtain restored content information; The speech feature information and the repaired content information are input into a preset speech synthesis model to generate second media stream data.

5. The speech processing method according to claim 4, characterized in that, The repair includes repairing repeated characters and repairing misspelled characters.

6. The speech processing method according to any one of claims 1-3, characterized in that, After performing first information extraction and second information extraction on the first media stream data containing audio data to obtain the speech feature information and speech content information of the audio data respectively, the method further includes: The voice feature information and voice content information are sent to a remote device so that the remote device can generate a second media stream data based on the voice feature information and voice content information.

7. The speech processing method according to any one of claims 1-3, characterized in that, The first information extraction and the second information extraction of the first media stream data containing audio data, respectively obtaining the speech feature information and speech content information of the audio data, include: First information extraction is performed on the starting segment of the first media stream data containing audio data to obtain the speech feature information of the audio data; The second information extraction is continuously performed on the first media stream data containing audio data to obtain the voice content information of the audio data.

8. The speech processing method according to any one of claims 1-3, characterized in that, The speech feature information includes one or more of the following: voiceprint feature information, rhythm feature information, and emotion feature information.

9. A voice processing device, characterized in that, include: An information extraction unit is used to perform first information extraction and second information extraction on a first media stream data containing audio data, and respectively obtain the speech feature information and speech content information of the audio data; The data generation unit is used to input the speech feature information and speech content information into a preset speech synthesis model to generate second media stream data. When the second media stream data is played, the speech content corresponding to the speech content information is presented based on the speech features corresponding to the speech feature information.

10. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more computer programs; When the one or more computer programs are executed by the one or more processors, the electronic device implements the voice processing method as described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech processing method as described in any one of claims 1-8.