Voice audio processing method, device and electronic equipment
By generating target audio resource packages and synthesizing response voices, the problem of low voice audio processing efficiency in the prior art is solved, and the consistency of customized tone broadcasts for electronic devices in different network environments is realized, thereby improving user experience.
Patent Information
- Application Number
- CN202111565295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-12-20
AI Technical Summary
In the prior art, voice and audio processing efficiency is low, making it difficult to realize that electronic devices can broadcast voice with customized tones, and it is difficult to maintain timbre consistency when the network environment is poor.
By obtaining the target corpus of the target pronunciator, generating the target audio resource package, and finding the target audio resource package when the network state meets the conditions or locally, synthesize the response voice corresponding to the target pronunciator to ensure that the electronic device performs voice broadcasting with the target tone.
The voice and audio processing process is simplified, time cost investment is reduced, processing efficiency is improved, and the target tone is consistently broadcast in various network environments, improving user experience.
Smart Images

Figure CN114495893B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method, device and electronic equipment for voice and audio processing. Background Art
[0002] With the advancement of science and technology, voice interaction is being used in an increasing number of electronic devices, such as mobile communications, automobiles, and smart home products. Compared to manual user interaction, voice interaction is simpler, has a lower barrier to entry, requires less sensory input, is more efficient, and can convey more acoustic information, improving the user experience.
[0003] Normally, electronic devices with voice interaction functions can only make voice broadcasts in a few universal timbres. In the case where users have customized requirements for the timbre of the electronic device’s voice broadcasts (for example, users want the electronic device to make voice broadcasts in the timbre of a certain person), the prior art can record the voice audio of the speaker, and based on the above-mentioned voice audio and deep learning technology, enable the electronic device to make voice broadcasts in the timbre of the above-mentioned speaker. However, in order to avoid timbre distortion when the electronic device makes voice broadcasts in the timbre of the above-mentioned speaker, it is usually necessary to collect a large amount of the voice audio of the above-mentioned speaker to generate a response voice with the broadcast timbre of the above-mentioned speaker, and the audio quality requirements for the above-mentioned voice audio are high, which requires a lot of time and cost, and the efficiency of voice audio processing is low. Summary of the Invention
[0004] The present invention provides a speech and audio processing method, a device and an electronic device, which are used to solve the defect of low efficiency of speech and audio processing in the prior art and achieve more efficient speech and audio processing.
[0005] The present invention provides a speech and audio processing method, comprising: obtaining a target corpus of a target speaker; generating a target audio resource package corresponding to the target speaker based on the target corpus; and upon receiving a download request sent by a first electronic device indicating a request to download the target audio resource package, sending the target audio resource package to the first electronic device, wherein the target audio resource package is used to generate a response speech corresponding to the target speaker; wherein the target audio resource package includes at least one speech audio, the broadcast timbre of the speech audio is the timbre of the target speaker; the target corpus includes corpus in the form of speech uttered by the target speaker, and the corpus content of the target corpus includes at least one of a predetermined phrase, a sentence, a short story, and a melody.
[0006] The present invention also provides a voice audio processing method, comprising: determining a target speaker corresponding to a received voice instruction based on the voiceprint features of the received voice instruction; searching for a target audio resource package corresponding to the target speaker when the network status meets specific conditions; determining a target voice audio from the target audio resource package based on a processing result of the voice instruction, synthesizing a response voice based on the processing result and the target voice audio, and broadcasting the response voice.
[0007] The present invention also provides a speech and audio processing device, comprising: a corpus acquisition module, used to acquire a target corpus of a target speaker; a resource package generation module, used to generate a target audio resource package corresponding to the target speaker based on the target corpus; a resource package sending module, used to send the target audio resource package to the first electronic device upon receiving a download request sent by a first electronic device indicating a request to download the target audio resource package, wherein the target audio resource package is used to generate a response voice corresponding to the target speaker; wherein the target audio resource package includes at least one speech audio, the broadcast timbre of the speech audio is the timbre of the target speaker; the target corpus includes corpus in the form of speech uttered by the target speaker, and the corpus content of the target corpus includes at least one of a predetermined phrase, sentence, short story and melody.
[0008] The present invention also provides a voice and audio processing device, comprising: a communication module, used to determine the target speaker corresponding to the voice instruction based on the voiceprint characteristics of the received voice instruction; a query module, used to search for a target audio resource package corresponding to the target speaker when the network status meets specific conditions; a broadcast module, used to determine the target voice audio from the target audio resource package based on the processing result of the voice instruction, and synthesize a response voice based on the processing result and the target voice audio, and broadcast the response voice.
[0009] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-described speech and audio processing methods are implemented.
[0010] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above-mentioned speech and audio processing methods when executed by a processor.
[0011] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned speech and audio processing methods are implemented.
[0012] The speech and audio processing method, device and electronic device provided by the present invention generate a target audio resource package corresponding to a target speaker based on a target corpus of the target speaker by a server, and upon receiving a download request sent by a first electronic device indicating a request to download the target audio resource package, the server sends the target audio resource package to the first electronic device. The first electronic device generates a response speech corresponding to the target speaker based on the target audio resource package, wherein the corpus content of the target corpus includes at least one of predetermined phrases, sentences, short stories and tunes. The speech and audio processing process can be simplified, the time cost of speech and audio processing can be reduced, and the efficiency of speech and audio processing can be improved. After the electronic device downloads the target audio resource package, even if the electronic device is not connected to the mobile Internet or the network communication quality is poor, the speech broadcast can still be performed in the timbre of the target speaker, ensuring the consistency of the speech broadcast in the timbre of the target speaker by the electronic device in various situations, and improving user perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 This is one of the flow charts of the voice and audio processing method provided by the present invention;
[0015] Figure 2 This is the second flow chart of the voice and audio processing method provided by the present invention;
[0016] Figure 3 This is the third flow chart of the voice and audio processing method provided by the present invention;
[0017] Figure 4 This is one of the structural diagrams of the speech and audio processing device provided by the present invention;
[0018] Figure 5 This is the second structural diagram of the speech and audio processing device provided by the present invention;
[0019] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0021] Figure 1 This is one of the flow charts of the speech audio processing method provided by the present invention. Figure 1 The speech audio processing method of the present invention is described. Figure 1 As shown, the method includes: step 101, obtaining target corpus of a target speaker.
[0022] It should be noted that the execution subject of the embodiment of the present invention is the server, which can be a cloud server.
[0023] The first electronic device is an electronic device with a voice interaction function used by a user. When the user has customized requirements for the timbre of the voice broadcast of the first electronic device, for example, if the user wants the first electronic device to broadcast voice in the voice of a certain person, the timbre that the user wants to set is the target timbre.
[0024] It should be noted that the electronic devices in the embodiments of the present invention may include, but are not limited to, mobile communication terminals, cars, refrigerators, washing machines, televisions, speakers, and sweeping robots.
[0025] When the first electronic device is configured to perform voice broadcasting in the timbre of a target speaker, the server may obtain a target corpus of the target speaker and generate a target audio resource package corresponding to the target speaker based on the target corpus of the target speaker. The timbre of the target speaker may be referred to as a target timbre.
[0026] Corpus refers to language materials that have actually appeared in actual use. By analyzing and processing the corpus, the linguistic features and pronunciation features of the corpus can be obtained. The corpus of the target speaker can refer to the speech materials that the target speaker has used in communication, reading, singing, etc. The target corpus includes the corpus in the form of speech produced by the target speaker. In the embodiment of the present invention, the corpus of the target speaker with the target content can be used as the target corpus of the target speaker.
[0027] It should be noted that the target content may be content predetermined based on linguistic analysis that can better reflect the pronunciation characteristics of the target timbre. The target content may include but is not limited to at least one of a predetermined phrase, sentence, short story, and melody.
[0028] Normally, in order to avoid timbre distortion when the first electronic device performs voice broadcasting with the target timbre, it is necessary to collect a large amount of voice audio of the target timbre, and based on the above-mentioned large amount of voice audio of the target timbre, after a complex model training process, obtain the pronunciation characteristics of the target timbre, and generate a target audio resource package corresponding to the target speaker based on the above-mentioned pronunciation characteristics. The time cost invested in the above-mentioned data collection work and model training process is relatively high, and the efficiency of voice timbre customization is relatively low. In the embodiment of the present invention, there is no need to collect a large amount of voice audio of the target timbre, and only a smaller time cost is required to acquire the target corpus of the target speaker, so that the pronunciation characteristics of the target timbre can be obtained more efficiently and accurately based on the target corpus of the target speaker.
[0029] In an embodiment of the present invention, the target corpus of the target speaker can be obtained in a variety of ways. For example: an electronic device can be used to record the audio of the target speaker reading the text including the target content, the audio of the target speaker chanting the music score including the target content, or the audio of the target speaker repeating the voice prompt information including the target content. The electronic device can send the recorded audio to the server, and after the server receives the above audio, the target corpus of the target speaker can be obtained based on the above audio; or, the corpus of the target speaker can be extracted from the stored video or audio, and based on the corpus of the target speaker, the target corpus of the target speaker can be further extracted, so that the acquisition of the target corpus of the target speaker can be realized online, which can further simplify the data collection work and reduce the time cost investment. Wherein, the above-mentioned stored video or audio is not limited to the corpus including the target speaker, but can also include the corpus of other speakers.
[0030] Step 102: Generate a target audio resource package corresponding to the target speaker based on the target corpus.
[0031] Specifically, after the server obtains the target corpus of the target speaker, it can generate a target audio resource package corresponding to the target speaker through deep learning technology based on the target corpus of the target speaker.
[0032] It should be noted that the target audio resource package corresponding to the target speaker may include at least one voice audio, and the timbre of any of the voice audios is the target timbre. The target audio resource package may also include acoustic feature parameters of the target timbre.
[0033] Step 103: upon receiving a download request from the first electronic device indicating a request to download the target audio resource package, the target audio resource package is sent to the first electronic device, where the target audio resource package is used to generate a response voice corresponding to the target speaker.
[0034] After the server generates the target audio resource package corresponding to the target speaker, it can actively send the above target audio resource package to the first electronic device, or after receiving a download request sent by the first electronic device indicating a request to download the above target audio resource package, it can send the above target audio resource package to the first electronic device in response to the download request.
[0035] Normally, when an electronic device is not connected to the mobile Internet or when the electronic device is connected to the mobile Internet but the network environment is poor, the electronic device can usually only perform voice broadcasts using several system-built-in broadcast tones, making it difficult to smoothly and naturally achieve voice broadcasts using the target tone.
[0036] In the embodiment of the present invention, after receiving the target audio resource package corresponding to the target speaker sent by the server, the first electronic device parses the target audio resource package to obtain and store at least one voice audio.
[0037] The first electronic device may set the timbre of the voice announcement as the target timbre in response to the user's timbre setting operation.
[0038] After the first electronic device sets the timbre of the voice broadcast to the target timbre, it can also respond to the user's input, use at least one voice audio in the target audio resource package as the response voice corresponding to the target speaker, and play the above response voice, so that the first electronic device can perform voice broadcast in the target timbre regardless of whether the first electronic device is connected to or not connected to the mobile Internet.
[0039] After the first electronic device receives the target audio resource package corresponding to the target speaker sent by the server, it parses the above target audio resource package and can also obtain the acoustic feature parameters of the target timbre. In response to the user's input, it can generate response data that matches the above input. Based on the above response data and the acoustic feature parameters of the target timbre, it can generate and play the response voice corresponding to the target speaker through deep learning technology, so that the first electronic device can perform voice broadcasting with the target timbre regardless of whether the first electronic device is connected to the mobile Internet or not.
[0040] It should be noted that the user's input can be expressed as a touch output to the electronic device, and the above-mentioned touch output may include but is not limited to click input, sliding input, and pressing input. The user's input can also be expressed as a physical button input corresponding to the electronic device. The user's input can also be expressed as a user's voice input. It is understood that each of the inputs listed above is an exemplary list, that is, the embodiments of the present application include but are not limited to each of the inputs listed above.
[0041] The embodiments of the present invention can simplify the process of voice and audio processing, reduce the time cost of voice and audio processing, and improve the efficiency of voice and audio processing. After the electronic device downloads the target audio resource package, even if the electronic device is not connected to the mobile Internet or the network communication quality is poor, the voice broadcast can still be performed in the timbre of the target speaker. This can ensure the consistency of the electronic device in performing voice broadcasts in the timbre of the target speaker in various situations, and can improve user perception.
[0042] Based on the contents of the above embodiments, a target audio resource package corresponding to a target speaker is generated based on a target corpus, including: inputting the target corpus into an audio synthesis model, and obtaining a target audio resource package output by the audio synthesis model; wherein the audio synthesis model is pre-constructed using a sample corpus of a sample speaker as a sample and a sample audio resource package corresponding to the sample speaker as a sample label; the sample audio resource package includes at least one sample speech audio, and the broadcast timbre of the sample speech audio is the timbre of the sample speaker; the sample corpus includes corpus in the form of speech emitted by the sample speaker, and the corpus content of the target corpus is the same as the corpus content of the sample corpus.
[0043] Specifically, a sample corpus of a sample speaker can be obtained as a sample, a sample audio resource package corresponding to the sample speaker can be obtained as a sample label, and an audio synthesis model can be trained to obtain a trained audio synthesis model. The timbre of the sample speaker can be called a sample timbre.
[0044] It should be noted that there can be multiple sample speakers. The sample audio resource package corresponding to the sample speaker can include at least one sample voice audio, and the timbre of any sample voice audio is the sample timbre. The sample audio resource package corresponding to the sample speaker can also include acoustic feature parameters of the sample timbre.
[0045] It can be understood that the corpus content of the sample corpus is the same as the corpus content of the target corpus, and both are the aforementioned predetermined target content.
[0046] After obtaining the target corpus of the target speaker, the target corpus of the target speaker can be input into a trained audio synthesis model, and the above audio synthesis model can output a target audio resource package corresponding to the target speaker.
[0047] The embodiment of the present invention obtains the target audio resource package corresponding to the target speaker according to the target corpus of the target speaker based on the audio synthesis model. Compared with the usual process of training the model based on a large amount of audio data of the target speaker and obtaining the target audio resource package corresponding to the target speaker, the embodiment of the present invention can obtain the target audio resource package corresponding to the target speaker more accurately and efficiently based on the target corpus of the target speaker with a smaller amount of data, thereby improving the efficiency of speech audio processing.
[0048] Based on the contents of the above embodiments, after sending the target audio resource package to the first electronic device, the method further includes: receiving status information of the response voice sent by the first electronic device. The status information includes at least one of the start time and end time of the response voice broadcast, identification information of the response voice, and feedback information based on the response voice.
[0049] Specifically, after receiving and parsing the target audio resource package, the first electronic device may generate a response voice corresponding to the target speaker based on the target audio resource package, and play the response voice.
[0050] The first electronic device may obtain identification information of the response voice, and the start time and end time of the response voice broadcast as status information of the response voice.
[0051] The first electronic device may also obtain user feedback information based on the response voice as status information of the voice broadcast. The feedback information may be user feedback information inputted regarding the target timbre or user feedback information regarding the content of the response voice.
[0052] It can be understood that based on the user's feedback information on the content of the response voice, the status information of the response voice is obtained, so as to judge whether the target timbre of the response voice is distorted or whether the content of the above response voice is accurate, etc. For example: if the content of the response voice is to prompt the user to click on a certain control and input a certain instruction, and the first electronic device obtains another instruction input by the user by clicking on another control within the target time period after the above response voice is broadcast, it can be inferred that the status information of the broadcast of the above response voice is that the target timbre may be distorted, or the content of the above response voice is inaccurate, etc. After the first electronic device obtains the status information of the above response voice, it can send the status information of the above response voice to the server.
[0053] The server may receive the status information of the response voice sent by the first electronic device.
[0054] Based on the state information, the audio synthesis model is updated.
[0055] Specifically, based on the status information of the above-mentioned response voice sent by the first electronic device, the trained audio synthesis model can be updated, for example: based on the start time and end time of the above-mentioned response voice broadcast by the first electronic device in the above-mentioned status information, the time for the next training of the above-mentioned audio synthesis model is updated; or, based on the feedback information based on the above-mentioned response voice in the above-mentioned status information, the training parameters of the above-mentioned audio synthesis model are updated.
[0056] The embodiment of the present invention can more accurately obtain a target audio resource package corresponding to a target speaker by updating an audio synthesis model based on the status information of the response voice sent by the first electronic device.
[0057] Based on the contents of the above embodiments, after generating a target audio resource package corresponding to a target speaker based on the target corpus, the method further includes: upon receiving a voice instruction sent by a first electronic device, determining the target speaker corresponding to the voice instruction based on the voiceprint features of the voice instruction.
[0058] It should be noted that, under certain conditions where the network status is good, the speech and audio processing method provided by the present invention can also support online synthesis of a response speech corresponding to the target speaker.
[0059] Specifically, the first electronic device may receive a voice instruction input by a user, and may send the voice instruction to the server.
[0060] After receiving the voice command sent by the first electronic device, the server can determine the speaker corresponding to the voice command based on the voiceprint characteristics of the voice command; it can also process the voice command to obtain the processing result of the voice command.
[0061] When the network status meets specific conditions, the target voice audio is determined from the target audio resource package according to the processing result of the voice instruction, and the response voice corresponding to the target speaker is synthesized based on the processing result and the target voice audio.
[0062] Specifically, after the server determines the speaker corresponding to the voice instruction, it can determine whether the speaker corresponding to the voice instruction is the target speaker.
[0063] When it is determined that the speaker corresponding to the above voice instruction is the target speaker, at least one voice audio in the target audio resource package corresponding to the target speaker can be determined as the target voice audio based on the processing result of the above voice instruction.
[0064] Based on the processing results of the target voice audio and the voice command, a response voice corresponding to the target speaker can be synthesized online, and the response voice is sent to the first electronic device, so that the first electronic device plays the response voice after receiving the response voice.
[0065] Specifically, after the server synthesizes the response voice corresponding to the target speaker online, the server can send the response voice to the first electronic device. After receiving the response voice, the first electronic device can broadcast the response voice.
[0066] The embodiment of the present invention determines the speaker corresponding to the voice instruction based on the voiceprint features of the voice instruction sent by the first electronic device when the network environment is good under specific conditions. When the speaker corresponding to the voice instruction is the target speaker, the target voice audio is determined from the target audio resource package corresponding to the target speaker according to the processing result of the voice instruction. The response voice corresponding to the target speaker is synthesized online based on the processing result and the target voice audio, and the response voice is sent to the first electronic device. After receiving the response voice, the first electronic device broadcasts the response voice. The response voice can be synthesized online when the network environment is good, which can reduce the occupation of local computing space and improve the local computing speed.
[0067] Based on the contents of the above embodiments, obtaining the target corpus of the target speaker includes: receiving target audio data of the target timbre sent by the second electronic device; wherein the target audio data is obtained by the second electronic device in response to an audio recording operation based on the target content.
[0068] Specifically, the second electronic device may be an electronic device used by the target speaker and may be the same as or different from the first electronic device.
[0069] The target speaker can read, chant, or repeat the predetermined target content in a target timbre, and can perform an audio recording operation on a second electronic device. In response to the audio recording operation, the second electronic device can collect audio data of the target speaker reading, chanting, or repeating the predetermined target content in a target timbre as the target audio data of the target speaker.
[0070] After the second electronic device obtains the target audio data, it can send the target audio data to the server. The server can receive the target audio data sent by the second electronic device. Based on the target audio data, the server obtains the target corpus. After receiving the target audio data, the server can perform corpus extraction based on the target content of the target audio data to obtain the target corpus of the target speaker.
[0071] It should be noted that after receiving the above-mentioned target audio data, the server can first perform data processing and data verification on the above-mentioned target audio data, and based on the target audio data after data processing and data verification, obtain the target corpus of the target speaker more accurately and efficiently.
[0072] In an embodiment of the present invention, a second electronic device obtains target audio data of a target speaker in response to an audio recording operation based on target content, and sends the target audio data to a server. The server obtains the target corpus of the target speaker based on the target audio data, and can obtain the target corpus of the target speaker more accurately and efficiently, and can provide a data basis for speech audio processing.
[0073] Based on the contents of the above embodiments, before obtaining the target corpus of the target speaker, the above method further includes: obtaining target content.
[0074] Specifically, linguistic analysis can be performed based on pre-acquired corpus to obtain corpus that can better reflect pronunciation characteristics.
[0075] By using a big data analysis method, one or more of phrases, sentences, short stories, or melodies can be extracted from the corpus that can better reflect the pronunciation characteristics as target content, and the target content is sent to the second electronic device.
[0076] Specifically, after the server obtains the target content, it can send the above target content to the second electronic device, so that after receiving the above target content, the second electronic device can display the above target content and / or play the above target content on the user interaction interface of the second electronic device, and obtain the target audio data of the target speaker in response to the audio recording operation of the target speaker based on the above target content.
[0077] The server may further extract the target corpus of the target speaker from the corpus of the target speaker based on the above target content after extracting the corpus of the target speaker from the stored video or audio.
[0078] The embodiment of the present invention can reduce the cost of data collection, simplify the process of voice timbre customization, and improve the efficiency of voice timbre customization by performing linguistic analysis on the pre-acquired corpus before obtaining the target corpus of the target speaker and obtaining the target content from the above corpus.
[0079] Figure 2 This is the second flow chart of the voice audio processing method provided by the present invention. Figure 2 The speech audio processing method of the present invention is described. Figure 2 As shown, the method includes: step 201, determining the target speaker corresponding to the voice instruction according to the voiceprint characteristics of the received voice instruction.
[0080] It should be noted that the execution subject of the embodiment of the present invention is the first electronic device.
[0081] The first electronic device is an electronic device with a voice interaction function used by a user. If the user has customized requirements for the timbre of the voice announcements made by the first electronic device, for example, if the user wants the electronic device to make announcements in the voice of a certain person, the timbre that the user wishes to set is the target timbre. Based on the voice and audio processing method provided by the present invention, the first electronic device can make announcements in the target timbre.
[0082] It should be noted that the electronic devices in the embodiments of the present invention may include, but are not limited to, mobile communication terminals, cars, refrigerators, washing machines, televisions, speakers, and sweeping robots.
[0083] The first electronic device can receive a voice command input by a user, and can determine a target speaker corresponding to the voice command based on the voiceprint features of the voice command; and can also process the voice command to obtain a processing result of the voice command.
[0084] Step 202: When the network status meets a specific condition, search for a target audio resource package corresponding to the target speaker.
[0085] It should be noted that, under specific conditions where the first electronic device is not connected to the mobile Internet or the network communication quality is poor, the voice and audio processing method provided by the present invention can also support local generation of a response voice corresponding to the target speaker.
[0086] Specifically, the first electronic device may locally search for a target audio resource package corresponding to the target speaker.
[0087] Step 203: According to the processing result of the voice instruction, the target voice audio is determined from the target audio resource package, and a response voice is synthesized based on the processing result and the target voice audio, and the response voice is broadcast.
[0088] Specifically, when the first electronic device finds the target audio resource package corresponding to the target speaker, one or more voice audios in the target audio resource package can be determined as target voice audios based on the processing results of the above-mentioned voice instructions, and a response voice can be locally synthesized based on the above-mentioned processing results and the above-mentioned target voice audios, and the above-mentioned response voice can be broadcast.
[0089] In an embodiment of the present invention, when the first electronic device is not connected to the mobile Internet or the network communication quality is poor, the first electronic device determines that the above-mentioned voice instruction corresponds to the target speaker based on the voiceprint characteristics of the received voice instruction, and when the target audio resource package corresponding to the target speaker is found locally, the target voice audio is determined from the target audio resource package according to the processing result of the above-mentioned voice instruction, and the response voice corresponding to the target speaker is synthesized and broadcast based on the above-mentioned processing result and the above-mentioned target voice audio. In the case where the electronic device is not connected to the mobile Internet or the network communication quality is poor, voice broadcasting with the timbre of the target speaker can still be achieved, which can ensure the consistency of the voice broadcasting of the electronic device with the timbre of the target speaker in various situations and improve user perception.
[0090] Based on the contents of the above embodiments, the method further includes: sending a download request to the server, indicating a request to download a target audio resource package corresponding to the target speaker.
[0091] After the server generates the target audio resource package corresponding to the target speaker, it can actively send the above-mentioned target audio resource package to the first electronic device, or after receiving the download request of the above-mentioned target audio resource package sent by the first electronic device, it can respond to the above-mentioned download request and send the above-mentioned target audio resource package to the first electronic device.
[0092] Receives and stores target audio resource packages sent by the service client.
[0093] The first electronic device can receive and store a target audio resource package corresponding to a target speaker sent by the service client. The target audio resource package includes at least one audio voice, the voice audio having the timbre of the target speaker; the target audio resource package is generated based on a target corpus of the target speaker; the target corpus includes corpus in the form of speech produced by the target speaker, and the corpus content of the target corpus includes at least one of a predetermined phrase, sentence, short story, and melody.
[0094] In an embodiment of the present invention, a server generates a target audio resource package corresponding to a target speaker based on a target corpus of the target speaker, and upon receiving a download request sent by a first electronic device indicating a request to download the target audio resource package, the server sends the target audio resource package to the first electronic device. The first electronic device generates a response voice corresponding to the target speaker based on the target audio resource package, thereby ensuring the consistency of the electronic device's voice broadcast in the timbre of the target speaker in various situations and improving user perception.
[0095] Based on the contents of the above embodiments, after performing a voice announcement using the target audio resource package and the target timbre, the method further includes obtaining status information of the response voice. The status information includes at least one of the start and end times of the response voice announcement, identification information of the response voice, and feedback information based on the response voice.
[0096] Specifically, after receiving and parsing the target audio resource package, the first electronic device may generate a response voice corresponding to the target speaker based on the target audio resource package, and play the response voice.
[0097] The first electronic device may obtain the identification information of the response voice, and the start and end times of the response voice announcement as status information of the response voice. The first electronic device may also obtain user feedback information based on the response voice as status information of the voice announcement. The feedback information may be user-input feedback regarding a target timbre, or user feedback regarding the content of the response voice.
[0098] It is understood that based on the user's feedback information regarding the content of the response voice, the status information of the response voice is obtained, thereby determining whether the target timbre of the response voice is distorted or whether the content of the response voice is accurate. For example, if the content of the response voice prompts the user to click on a certain control to enter a certain instruction, and the first electronic device obtains another instruction entered by the user by clicking on another control within a target period after the broadcast of the response voice ends, it can be inferred that the status information of the broadcast of the response voice indicates that the target timbre may be distorted or the content of the response voice is inaccurate.
[0099] After obtaining the state information of the response voice, the first electronic device may send the state information of the response voice to the server.
[0100] Specifically, the server can update the audio synthesis model based on the status information of the above-mentioned response voice sent by the first electronic device, for example: based on the start time and end time of the above-mentioned response voice broadcast by the first electronic device in the above-mentioned status information, update the time for the next training of the above-mentioned audio synthesis model; or, based on the feedback information based on the above-mentioned response voice in the above-mentioned status information, update the training parameters of the above-mentioned audio synthesis model.
[0101] The embodiment of the present invention can more accurately obtain a target audio resource package corresponding to a target speaker by updating an audio synthesis model based on the status information of the response voice sent by the first electronic device.
[0102] Figure 3 This is the third flow chart of the speech audio processing method provided by the present invention. Figure 3 As shown, the second electronic device can guide the target speaker to perform audio recording operations based on the target content on the second electronic device through voice prompts, pop-up prompts, etc. The second electronic device can obtain the target audio data of the target speaker in response to the above audio recording operations.
[0103] The second electronic device can send the acquired target audio data of the target speaker to the server. The server can process and verify the target audio data of the target speaker, and obtain the target corpus of the target speaker based on the processed and verified target audio data.
[0104] Based on the target corpus of the target speaker, the server can obtain the target audio data packet corresponding to the target speaker through deep learning technology.
[0105] The first electronic device can download the target audio data packet from the server and parse the target audio data packet. The first electronic device can set the timbre of the voice broadcast to the target timbre in response to the user's timbre setting operation and generate a response voice based on the target audio data packet.
[0106] Figure 4 This is a structural diagram of the speech audio processing device provided by the present invention. Figure 4 The test device provided by the present invention is described, and the speech and audio processing device described below and the speech and audio processing method provided by the present invention described above can be referred to each other. Figure 4 As shown, the device includes: a corpus acquisition module 401, a resource package generation module 402 and a resource package sending module 403.
[0107] The corpus acquisition module 401 is used to acquire the target corpus of the target speaker.
[0108] The resource package generation module 402 is used to generate a target audio resource package corresponding to a target speaker based on the target corpus.
[0109] The resource package sending module 403 is configured to send the target audio resource package to the first electronic device upon receiving a download request for downloading the target audio resource package sent by the first electronic device. The target audio resource package is used to generate a response voice corresponding to the target speaker.
[0110] Among them, the target audio resource package includes at least one voice audio, the broadcast tone of the voice audio is the tone of the target speaker; the target corpus includes corpus in the form of voice emitted by the target speaker, and the corpus content of the target corpus includes at least one of predetermined phrases, sentences, short stories and tunes.
[0111] Specifically, the corpus acquisition module 401 , the resource package generation module 402 and the resource package sending module 403 are electrically connected.
[0112] It should be noted that the voice and audio processing device in the embodiment of the present invention is a server.
[0113] The resource package generation module 402 can generate a target audio resource package corresponding to the target speaker based on the target corpus of the target speaker through deep learning technology.
[0114] The resource package sending module 403 can actively send the above-mentioned target audio resource package to the first electronic device, or after receiving a download request sent by the first electronic device indicating a request to download the above-mentioned target audio resource package, send the above-mentioned target audio resource package to the first electronic device in response to the above-mentioned download request.
[0115] Optionally, the resource package generation module 402 may further include a resource package generation submodule.
[0116] The resource package generation submodule can be used to input the target corpus into the audio synthesis model and obtain the target audio resource package output by the audio synthesis model; wherein, the audio synthesis model is pre-constructed with the sample corpus of the sample speaker as the sample and the sample audio resource package corresponding to the sample speaker as the sample label; the sample audio resource package includes at least one sample speech audio, and the broadcast timbre of the sample speech audio is the timbre of the sample speaker; the sample corpus includes the corpus in the form of speech emitted by the sample speaker, and the corpus content of the target corpus is the same as the corpus content of the sample corpus.
[0117] Optionally, the resource package generation module 402 may further include a model updating submodule.
[0118] The model update submodule can be used to receive the status information of the response voice sent by the first electronic device; based on the status information, update the audio synthesis model; wherein the status information includes at least one of the start time and end time of the response voice broadcast, the identification information of the response voice, and feedback information based on the response voice.
[0119] Optionally, the speech and audio processing device may further include a target content acquisition module.
[0120] The target content acquisition module can be used to acquire target content before the corpus acquisition module 401 acquires the target corpus of the target speaker; and send the target content to the second electronic device.
[0121] After the embodiment of the present invention downloads the target audio resource package through the electronic device, even if the electronic device is not connected to the mobile Internet or the network communication quality is poor, the voice broadcast can still be performed in the timbre of the target speaker. This can ensure the consistency of the electronic device in performing voice broadcasts in the timbre of the target speaker in various situations, and can improve user perception.
[0122] Figure 5 This is a structural diagram of the speech audio processing device provided by the present invention. Figure 5 The test device provided by the present invention is described, and the speech and audio processing device described below and the speech and audio processing method provided by the present invention described above can be referred to each other. Figure 5 As shown, the device includes: a communication module 501, a query model 502 and a broadcast module 503.
[0123] The communication module 501 is used to determine the target speaker corresponding to the voice instruction according to the voiceprint characteristics of the received voice instruction.
[0124] The query model 502 is used to search for a target audio resource package corresponding to a target speaker when the network status meets a specific condition.
[0125] The broadcast module 503 is used to determine the target voice audio from the target audio resource package according to the processing result of the voice instruction, synthesize the response voice based on the processing result and the target voice audio, and broadcast the response voice.
[0126] Specifically, the communication module 501, the query model 502 and the broadcast module 503 are electrically connected.
[0127] It should be noted that the speech and audio processing device in the embodiment of the present invention is an electronic device. The above electronic device is an electronic device with a speech interaction function used by a user.
[0128] The communication module 501 can receive the voice command input by the user, and can determine the target speaker corresponding to the voice command based on the voiceprint characteristics of the voice command; it can also process the voice command to obtain the processing result of the voice command.
[0129] The query model 502 may locally search for a target audio resource package corresponding to the target speaker.
[0130] The broadcast module 503 can determine one or more voice audios in the target audio resource package as target voice audios based on the processing results of the above-mentioned voice instructions, and can locally synthesize the response voice based on the above-mentioned processing results and the above-mentioned target voice audios, and broadcast the above-mentioned response voice.
[0131] Optionally, the speech and audio processing device may further include a resource package downloading module and a status acquiring module.
[0132] The resource package download module can be used to send a download request to a server requesting a target audio resource package corresponding to a target speaker; receive and store the target audio resource package sent by the server. The target audio resource package includes at least one audio voice, the voice audio being broadcast in the timbre of the target speaker; the target audio resource package is generated based on a target corpus for the target speaker; the target corpus includes the target speaker's speech in the form of speech, and the target corpus includes at least one of predetermined phrases, sentences, short stories, and melodies.
[0133] The status acquisition module can be used to obtain the status information of the response voice; send the status information of the response voice to the server, so that the server updates the audio synthesis model used to obtain the target audio resource package based on the status information; wherein the status information includes the start time and end time of the response voice broadcast, the identification information of the response voice and at least one of the feedback information based on the response voice.
[0134] The embodiment of the present invention determines the target voice audio from the target audio resource package according to the processing result of the above-mentioned voice instruction when the first electronic device is not connected to the mobile Internet or under specific conditions where the network communication quality is poor, and synthesizes and broadcasts the response voice corresponding to the target speaker based on the above-mentioned processing result and the above-mentioned target voice audio, thereby ensuring the consistency of the electronic device's voice broadcast with the timbre of the target speaker in various situations.
[0135] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the voice and audio processing method described in any of the above embodiments.
[0136] In addition, the logic instructions in the memory 630 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0137] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech audio processing method described in any of the above embodiments.
[0138] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the speech audio processing methods provided by the above methods.
[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for processing speech and audio, characterized in that: include: Obtain the target corpus of the target speaker; Based on the target corpus, generating a target audio resource package corresponding to the target speaker; Upon receiving a download request from the first electronic device requesting to download the target audio resource package, sending the target audio resource package to the first electronic device, where the target audio resource package is used to generate a response voice corresponding to the target speaker; The target audio resource package includes at least one voice audio, the broadcast tone of the voice audio is the tone of the target speaker; the target corpus includes the corpus in the form of voice emitted by the target speaker, and the corpus content of the target corpus includes at least one of predetermined phrases, sentences, short stories and tunes.
2. The method for processing speech and audio according to claim 1, wherein: Generating a target audio resource package corresponding to the target speaker based on the target corpus includes: Inputting the target corpus into an audio synthesis model, and obtaining the target audio resource package output by the audio synthesis model; Among them, the audio synthesis model is pre-constructed with the sample corpus of the sample speaker as the sample and the sample audio resource package corresponding to the sample speaker as the sample label; the sample audio resource package includes at least one sample voice audio, and the broadcast timbre of the sample voice audio is the timbre of the sample speaker; the sample corpus includes the corpus in the form of speech emitted by the sample speaker, and the corpus content of the target corpus is the same as the corpus content of the sample corpus.
3. The speech audio processing method according to claim 1 or 2, characterized in that: After sending the target audio resource package to the first electronic device, the method further includes: receiving status information of the response voice sent by the first electronic device; Based on the state information, updating the audio synthesis model; The status information includes at least one of the start time and end time of the response voice broadcast, identification information of the response voice, and feedback information based on the response voice.
4. The method for processing speech and audio according to claim 1, wherein: After generating the target audio resource package corresponding to the target speaker based on the target corpus, the method further includes: Upon receiving a voice command sent by the first electronic device, determining a target speaker corresponding to the voice command based on a voiceprint feature of the voice command; Determining a target voice audio from the target audio resource package according to a processing result of the voice instruction, and synthesizing a response voice corresponding to the target speaker based on the processing result and the target voice audio; The response voice is sent to the first electronic device, so that the first electronic device broadcasts the response voice after receiving the response voice.
5. A method for processing speech and audio, characterized in that: include: Determining a target speaker corresponding to the received voice command based on the voiceprint features of the received voice command; When a network state satisfies a specific condition, searching for a target audio resource package corresponding to the target speaker, the target audio resource package including at least one voice audio, the voice audio having a timbre corresponding to the target speaker; the target audio resource package being generated based on a target corpus of the target speaker; the target corpus including corpus in the form of speech uttered by the target speaker, the corpus content of the target corpus including at least one of predetermined phrases, sentences, short stories, and tunes; According to the processing result of the voice instruction, the target voice audio is determined from the target audio resource package, and a response voice is synthesized based on the processing result and the target voice audio, and the response voice is broadcast.
6. The method for processing speech and audio according to claim 5, wherein: Before searching for a target audio resource package corresponding to the target speaker when the network status satisfies a specific condition, the method further includes: Send a download request to the server requesting to download the target audio resource package corresponding to the target speaker; Receive and store the target audio resource package sent by the server.
7. The method for processing speech and audio according to claim 5, wherein: After synthesizing a response voice based on the processing result and the target voice audio, and broadcasting the response voice, the method further includes: Obtaining status information of the response voice; Sending the status information of the response voice to the server, so that the server updates the audio synthesis model used to obtain the target audio resource package based on the status information; The status information includes at least one of the start time and end time of the response voice broadcast, identification information of the response voice, and feedback information based on the response voice.
8. A speech audio processing device, characterized in that: include: Corpus acquisition module, used to obtain target corpus of target speaker; A resource package generation module, configured to generate a target audio resource package corresponding to the target speaker based on the target corpus; a resource package sending module, configured to, upon receiving a download request from a first electronic device indicating a request to download the target audio resource package, send the target audio resource package to the first electronic device, wherein the target audio resource package is used to generate a response voice corresponding to the target speaker; The target audio resource package includes at least one voice audio, the broadcast tone of the voice audio is the tone of the target speaker; the target corpus includes the corpus in the form of voice emitted by the target speaker, and the corpus content of the target corpus includes at least one of predetermined phrases, sentences, short stories and tunes.
9. A speech audio processing device, characterized in that: include: A communication module, configured to determine a target speaker corresponding to a received voice command based on the voiceprint features of the received voice command; A query module is configured to search for a target audio resource package corresponding to the target speaker when a network state satisfies a specific condition, wherein the target audio resource package includes at least one voice audio, the broadcast timbre of the voice audio being the timbre of the target speaker; the target audio resource package is generated based on a target corpus of the target speaker; the target corpus includes corpus in the form of speech uttered by the target speaker, and the corpus content of the target corpus includes at least one of a predetermined phrase, sentence, short story, and melody; The broadcast module is used to determine the target voice audio from the target audio resource package according to the processing result of the voice instruction, synthesize the response voice based on the processing result and the target voice audio, and broadcast the response voice.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the speech audio processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for synthesizing voice data
CN112289303A
Electronic device and method of obtaining emotion information
CN112788990A