Speech processing method, apparatus, and electronic device

By acquiring the current voice and mixed audio, and analyzing the target scenario and voiceprint information, the smart speaker can accurately identify and process the target application, solving the problem of improving voice processing capabilities during mixed playback in singing scenarios.

CN118969018BActive Publication Date: 2025-10-10SHANGHAI XIAODU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410969333.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-10-10
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

In the smart speaker singing scenario, how to accurately identify which voice-related applications need to process the current voice while mixing and playing, so as to improve voice processing capabilities.

Method used

Acquire the current voice and mixed audio, determine the target application based on the target scenario, and process the current voice through the target application, including scene matching and voiceprint information analysis to identify sub-voices, ensuring the accuracy of voice processing.

Benefits of technology

In the singing scenario, it can accurately determine and process the target application, improve the voice processing capabilities of the smart speaker, and ensure that the mixed playback does not affect the processing of other voice applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118969018B_ABST
    Figure CN118969018B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice processing method and device and electronic equipment, relates to the technical field of computers, and particularly relates to the technical field of voice processing, intelligent sound boxes, and the like. A specific implementation scheme is as follows: a current voice and a current mixed audio are acquired, and the current mixed audio is played, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song; in a case where the current voice matches one or more target scenes, one or more target applications are determined based on the one or more target scenes, wherein different target scenes in the one or more target scenes are associated with different target applications; and the current voice is processed based on the one or more target applications, so as to obtain one or more processing results of the current voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as voice processing and smart speakers. Background Art

[0002] In singing scenarios, smart speakers need to mix and play the voice in real time. However, if other voice-related applications are also running on the smart speaker, it may not be able to distinguish whether the current voice needs to be transmitted to other voice-related applications for corresponding processing. Therefore, how to play the mixed voice in singing scenarios while also intelligently distinguishing which other voice-related applications need to process the current voice, thereby improving the voice processing capabilities of smart speakers in singing scenarios, has become a technical problem to be solved. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for speech processing.

[0004] According to one aspect of the present disclosure, there is provided a speech processing method, comprising:

[0005] Acquire a current voice and a current mixed audio, and play the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song;

[0006] In a case where the current voice has one or more matching target scenes, determining one or more target applications based on the one or more target scenes, wherein different target scenes in the one or more target scenes are associated with different target applications;

[0007] The current speech is processed based on the one or more target applications to obtain one or more processing results of the current speech.

[0008] According to another aspect of the present disclosure, there is provided a speech processing apparatus, comprising:

[0009] A voice acquisition module, configured to acquire a current voice and a current mixed audio, and play the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song;

[0010] a target application determination module, configured to, if the current voice has one or more matching target scenarios, determine one or more target applications based on the one or more target scenarios, wherein different target scenarios in the one or more target scenarios are associated with different target applications;

[0011] The speech processing module is used to process the current speech based on the one or more target applications to obtain one or more processing results of the current speech.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method of any embodiment of the present disclosure.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method according to any embodiment of the present disclosure when executed by a processor.

[0018] By adopting the above-mentioned implementation method, the current voice and the current mixed audio are obtained, the current mixed audio is played, and when there are one or more target scenes that match the current voice, one or more target applications are determined based on the one or more target scenes, and the current voice is processed based on the one or more target applications to obtain a processing result. In this way, the current mixed audio can be obtained and played, and the obtained current voice can also be scene-matched to determine the corresponding target application, and then the current voice is processed by the target application, thereby ensuring that the current mixed audio is played in the singing scene while the target application for corresponding processing of the current voice can be accurately determined, thereby improving the voice processing capability in the singing scene.

[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0021] Figure 1 is a flowchart of a speech processing method according to an embodiment of the present disclosure;

[0022] Figure 2 is a scenario diagram of a speech processing method according to an embodiment of the present disclosure;

[0023] Figure 3 is a schematic block diagram of a speech processing device according to an embodiment of the present disclosure;

[0024] Figure 4 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] Figure 1 : is a schematic flow chart of the audio processing method proposed in an embodiment of the present disclosure, including:

[0027] S110, obtaining a current voice and a current mixed audio, and playing the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of the current song;

[0028] S120: If the current voice has one or more matching target scenarios, determine one or more target applications based on the one or more target scenarios, wherein different target scenarios in the one or more target scenarios are associated with different target applications;

[0029] S130: Process the current speech based on the one or more target applications to obtain one or more processing results of the current speech.

[0030] The speech processing method of the embodiments of the present disclosure can be executed by an electronic device. The electronic device can be one of a smart speaker, a smart TV, and a smart display. It should be understood that the above is only an exemplary description of an electronic device, and the actual processing is not limited to the devices mentioned in the above examples. As long as the electronic device can execute the speech processing method provided by this embodiment, it is within the scope of protection of this embodiment.

[0031] By adopting the above implementation, the current voice and the current mixed audio are acquired, the current mixed audio is played, in the case that the current voice matches one or more target scenes, one or more target applications are determined based on the one or more target scenes, and the current voice is processed based on the one or more target applications to obtain a processing result. In this way, the current mixed audio can be acquired and played, and the acquired current voice can also be matched with a scene to determine a corresponding target application, and then the current voice is processed by the target application, so as to ensure that the current mixed audio is played in the singing scene, and the target application for processing the current voice can also be accurately determined, and the voice processing capability in the singing scene is improved.

[0032] Optionally, before the acquiring the current voice and the current mixed audio and playing the current mixed audio, the method can further include: in response to a voice instruction for starting the singing application, starting the singing application.

[0033] Optionally, before the acquiring the current voice and the current mixed audio and playing the current mixed audio, the method can further include: in response to a selection operation of the singing application on an interaction interface of the electronic device, starting the singing application.

[0034] Optionally, the acquiring the current voice and the current mixed audio and playing the current mixed audio can include: acquiring the current voice; generating the current mixed audio based on the current voice and a current accompaniment of the current song, and playing the current mixed audio.

[0035] The acquiring the current voice can include: acquiring the current voice based on a first audio acquisition device. The first audio acquisition device can be a microphone, which can also be referred to as a loudspeaker. The first audio acquisition device can be built-in in the electronic device or connected to the electronic device through a specified connection mode, which can be a USB (Universal Serial Bus), Bluetooth, or the like, which is not limited or exhaustive here.

[0036] The generating of current mixed audio based on the current voice and the current accompaniment of the current song and playing the current mixed audio include: processing the current voice and the current accompaniment of the current song based on a specified processing method to obtain pre-mixed audio and playing the current mixed audio. The specified processing method includes at least one of the following: volume balance, frequency balance, adding effects, etc. For example, volume balance can be adjusting the volume of the current voice and the current accompaniment of the current song to ensure that the two are proportionally coordinated in the overall mixed audio; frequency balance can be adjusting the frequency of the current voice and the current accompaniment of the current song to avoid frequency conflicts between each other; adding effects can be adding the same or different effects (such as reverberation) to the current voice and the current accompaniment of the current song. This is only an exemplary explanation. Other methods can also be used in actual processing, which will not be described here.

[0037] Among them, the processing of the current voice and the current accompaniment of the current song based on the specified processing method to obtain the pre-mixed audio can be performed by a first audio processing device in the electronic device, and the first audio processing device can be a DSP (Digital Signal Processing) or a CPU (Central Processing Unit).

[0038] Playing the current mixed audio may include: playing the current mixed audio through a playback device, wherein the playback device may be one of the following: a speaker, an earphone, a loudspeaker, and the like.

[0039] Optionally, obtaining the current voice and the current mixed audio, and playing the current mixed audio, includes: sending the current accompaniment of the current song to an audio processing device, receiving the current voice and the current mixed audio sent by the audio processing device, and playing the current mixed audio.

[0040] On the audio processing device side, it can include: receiving the current accompaniment of the current song sent by the electronic device, and collecting the current voice; processing the current voice and the current accompaniment of the current song based on a specified processing method to obtain a pre-mixed audio; sending the current voice and the current mixed audio to the electronic device, wherein the processing of the current voice and the current accompaniment of the current song based on the specified processing method to obtain the pre-mixed audio is the same as the above-mentioned implementation method and will not be repeated here.

[0041] The collecting of the current voice may be: collecting the current voice based on a second audio collecting device in the audio processing device.

[0042] The audio processing device may further include a second audio processing device, and the processing of the current voice and the current accompaniment of the current song based on a specified processing method to obtain the pre-mixed audio may be performed by the second audio processing device, and the second audio processing device may be a DSP.

[0043] By adopting the above embodiment, the current accompaniment of the current song is sent to the audio processing device, and the current voice and the current mixed audio sent by the audio processing device are received. In this way, the electronic device does not need to perform the processing of generating the current mixed audio, which can save the processing resources of the electronic device and avoid the electronic device from freezing.

[0044] In some possible implementations, when there are one or more target scenes that match the current voice, determining one or more target applications based on the one or more target scenes includes: determining a scene recognition result of the current voice based on the current voice and the voice attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein different candidate scenes in the one or more candidate scenes are associated with different candidate applications, and the scene recognition result of the current voice is used to indicate whether the current voice matches each candidate scene; and determining one or more target applications based on the one or more target scenes when it is determined that there are one or more target scenes that match the current voice in the one or more candidate scenes based on the scene recognition result of the current voice.

[0045] The candidate applications are user-specified or pre-configured voice-related applications that need to be executed in a singing scenario, wherein the number of the candidate applications can be one or more.

[0046] Each candidate application in the one or more candidate applications is associated with a respective candidate scenario, wherein the association between each candidate application and the respective candidate scenario is preconfigured.

[0047] The speech attribute requirements corresponding to any candidate scene are used to represent the conditions that the speech required for the candidate scene should meet. For example, it may include at least one of the following: the speech contains specified keywords, the speech is singing speech, any type of speech, etc.; wherein, the speech is singing speech, which may include: the frequency of the speech matches the accompaniment, and / or the text information of the speech matches the lyrics.

[0048] The voice attribute requirements corresponding to any candidate scenario can be determined by generating the voice attribute requirements corresponding to the candidate scenario based on the voice requirements of the candidate applications associated with the candidate scenario. For example, the voice requirements of the candidate applications associated with the candidate scenario can be directly used as the voice attribute requirements corresponding to the candidate scenario. In another example, the voice requirements of the candidate applications associated with the candidate scenario can be processed (e.g., a portion of the voice requirements can be extracted) to obtain the voice attribute requirements corresponding to the candidate scenario.

[0049] Determining the scene recognition result of the current voice based on the current voice and the voice attribute requirements corresponding to each candidate scene in one or more candidate scenes may include: analyzing whether the current voice matches the voice attribute requirements corresponding to the g-th candidate scene based on the voice attribute requirements corresponding to the g-th candidate scene, wherein g is a positive integer and the g-th candidate scene is any one of the one or more candidate scenes; when the current voice matches the voice attribute requirements corresponding to the g-th candidate scene, adding the recognition result of the current voice matching the g-th candidate scene to the scene recognition result of the current voice; when the current voice does not match the voice attribute requirements corresponding to the g-th candidate scene, adding the recognition result of the current voice not matching the g-th candidate scene to the scene recognition result of the current voice.

[0050] The method for determining whether the current voice matches each candidate scene is the same as the method for determining whether the current voice matches the g-th candidate scene. The recognition results of whether the current voice matches each candidate scene are added to the scene recognition results of the current voice in the same way. Finally, the scene recognition results of the current voice that include the recognition results of whether the current voice matches all candidate scenes can be obtained. No further details will be given here.

[0051] The analyzing, based on the speech attribute requirements corresponding to the g-th candidate scene, whether the current speech matches the speech attribute requirements corresponding to the g-th candidate scene may include at least one of the following:

[0052] When the speech attribute requirement corresponding to the g-th candidate scene is that the speech contains a specified keyword, extracting text information of the current speech, determining whether the text information of the current speech contains the specified keyword, and if the text information of the current speech includes the specified keyword, determining that the current speech matches the speech attribute requirement corresponding to the g-th candidate scene; and if the text information of the current speech does not include the specified keyword, determining that the current speech does not match the speech attribute requirement corresponding to the g-th candidate scene;

[0053] When the speech attribute requirement corresponding to the g-th candidate scene is that the speech is a singing speech, the current speech is analyzed to obtain the frequency of the current speech, and whether the frequency of the current speech matches the current accompaniment of the current song is determined; if the frequency of the current speech matches the current accompaniment of the current song, it is determined that the current speech matches the speech attribute requirement corresponding to the g-th candidate scene; if the frequency of the current speech does not match the current accompaniment of the current song, it is determined that the current speech does not match the speech attribute requirement corresponding to the g-th candidate scene;

[0054] When the speech attribute requirement corresponding to the g-th candidate scene is that the speech is a singing speech, extracting text information of the current speech, determining whether the text information of the current speech matches the current lyrics of the current song; when the text information of the current speech matches the current lyrics of the current song, determining that the current speech matches the speech attribute requirement corresponding to the g-th candidate scene; when the text information of the current speech does not match the current lyrics of the current song, determining that the current speech does not match the speech attribute requirement corresponding to the g-th candidate scene;

[0055] When the speech attribute requirement corresponding to the g-th candidate scene is that the speech is of any type, it is determined that the current speech matches the speech attribute requirement corresponding to the g-th candidate scene.

[0056] It should be pointed out that the above is only an exemplary description. In actual processing, the speech attribute requirements corresponding to the g-th candidate scene may include the above multiple requirements or conditions. For example, the speech attribute requirements corresponding to the g-th candidate scene include that the speech contains specified keywords and the speech is a singing speech. Then, the above processing can be performed to determine whether the text information of the current speech contains specified keywords, whether the frequency of the current speech matches the current accompaniment of the current song, and whether the text information of the current speech matches the current lyrics of the current song. Only when the text information of the current speech includes specified keywords, the frequency of the current speech matches the current accompaniment of the current song, and the text information of the current speech matches the current lyrics of the current song, can it be determined that the current speech matches the speech attribute requirements corresponding to the g-th candidate scene. This is also only an exemplary description, and all possible combinations will not be repeated here.

[0057] The determining of one or more target applications based on the one or more target scenes, in a case where it is determined based on the scene recognition result of the current voice that the one or more candidate scenes contain the one or more target scenes that match the current voice, may include: judging based on the scene recognition result of the current voice whether the one or more target scenes that match the current voice exist in the one or more candidate scenes, and in a case where it is determined that the one or more target scenes that match the current voice exist in the one or more candidate scenes, determining the candidate applications associated with each target scene in the one or more target scenes, and using the candidate applications associated with each target scene as the target application.

[0058] After determining whether there are one or more target scenes matching the current voice in the one or more candidate scenes based on the scene recognition result of the current voice, the method may also include: if it is determined that none of the one or more candidate scenes matches the current voice, not processing the current voice.

[0059] By adopting the above-mentioned implementation method, based on the current voice and the voice attribute requirements corresponding to each candidate scene in one or more candidate scenes, it is determined whether the current voice matches each candidate scene; based on the one or more target scenes determined by the matching results, one or more target applications are determined, wherein different candidate scenes in the one or more candidate scenes are associated with different candidate applications. In this way, based on the matching results of the current voice with the voice attribute requirements corresponding to the candidate scenes associated with the candidate applications, it is possible to accurately determine whether the current voice is the voice required by the candidate applications, thereby improving the accuracy of the matching between the current voice and the candidate applications.

[0060] In one possible implementation, the processing of the current voice based on the one or more target applications to obtain one or more processing results of the current voice includes: sending the current voice to the kth target application through a thread, so that the kth target application processes the current voice, and obtaining the processing result of the current voice returned by the kth target application, wherein k is a positive integer, the kth application is one of the one or more target applications, and the thread can be the smallest unit of operation scheduling.

[0061] For example, the kth target application is a wake-up application, and the processing result of the current voice returned by the kth target application can be to wake up the electronic device. For example, the kth target application is a singing scoring application, and the processing result of the current voice returned by the kth target application can be a scoring result. It should be understood that the above is only an exemplary description of the target application, and in actual processing, there can be more types of applications capable of processing based on the voice as the target application, and the specific processing result that each target application can obtain is also related to the processing logic of the target application itself, which is not limited or exhausted here.

[0062] In some possible implementation manners, the determining, in the case that the current voice exists matching one or more target scenes, one or more target applications based on the one or more target scenes comprises: in the case that a plurality of voiceprint information is analyzed from the current voice, extracting a plurality of first sub-voices from the current voice, wherein different first sub-voices in the plurality of first sub-voices correspond to different voiceprint information in the plurality of voiceprint information; in the case that one or more target scenes exist in each second sub-voice of one or more second sub-voices in the plurality of first sub-voices, determining one or more target applications corresponding to each second sub-voice based on the one or more target scenes matched by each second sub-voice.

[0063] The extracting, in the case that a plurality of voiceprint information is analyzed from the current voice, a plurality of first sub-voices from the current voice comprises: analyzing the current voice to obtain the number of voiceprint information contained in the current voice, and in the case that a plurality of voiceprint information is analyzed from the current voice, extracting, based on each voiceprint information in the plurality of voiceprint information, a first sub-voice corresponding to each voiceprint information from the current voice.

[0064] The analyzing the current voice to obtain the number of voiceprint information contained in the current voice can comprise: extracting a feature of the current voice, and analyzing the feature of the current voice by a specified analysis manner to obtain the number of voiceprint information contained in the current voice, wherein the specified analysis manner can comprise any one of MFCC (Mel Frequency Cepstral Coefficient), LPC (Linear Prediction Coding), PLP (Perceptual Linear Predictive), and the like.

[0065] After analyzing the current speech to obtain the number of voiceprint information items contained in the current speech, the method further includes: if only one voiceprint item is obtained from the analysis of the current speech, determining a scene recognition result for the current speech based on the current speech and the corresponding voice attribute requirements of each of the one or more candidate scenes. The method of determining a scene recognition result for the current speech based on the current speech and the corresponding voice attribute requirements of each of the one or more candidate scenes is the same as that in the above embodiment and will not be further described here.

[0066] The method of determining the one or more target applications corresponding to each second sub-voice based on the one or more target scenarios matched by each second sub-voice in the multiple first sub-voices includes: judging whether there is one or more target scenarios matched by each first sub-voice in the multiple first sub-voices, and determining the one or more target applications corresponding to each second sub-voice based on the one or more target scenarios matched by each second sub-voice in the multiple first sub-voices.

[0067] After determining whether each of the plurality of first sub-voices has one or more matching target scenes, the method further includes: if one or more third sub-voices in the plurality of first sub-voices do not have a matching target scene, not processing the one or more third sub-voices.

[0068] By adopting the above embodiment, when multiple different first sub-speech corresponding to multiple different voiceprint information is extracted from the current speech, and a second sub-speech among the multiple different first sub-speech matches one or more target scenarios, one or more target applications corresponding to the second sub-speech are determined based on the one or more target scenarios matched by the second sub-speech. In this way, a more fine-grained analysis of the current speech can be performed to obtain multiple sub-speech corresponding to multiple voiceprints, and the multiple sub-speech can be analyzed to determine whether there are corresponding target scenarios, thereby ensuring the accuracy of the determination of whether each sub-speech matches the target application.

[0069] In a possible implementation, wherein, when each second sub-voice in one or more second sub-voices in the plurality of first sub-voices has one or more matching target scenes, determining one or more target applications corresponding to each second sub-voice based on the one or more target scenes matched by each second sub-voice includes: determining a scene recognition result of the jth first sub-voice based on the jth first sub-voice and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein j is a positive integer, and the scene recognition result of the jth first sub-voice is used to represent the jth first sub-voice and each candidate scene. Whether there is a match; in the case where it is determined based on the scene recognition result of the j-th first sub-voice that there are one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes, the j-th first sub-voice is used as the ith second sub-voice, and the one or more target scenes matching the j-th first sub-voice are used as the one or more target scenes matching the ith second sub-voice, wherein i is a positive integer, and the ith second sub-voice is one of the one or more second sub-voices; based on the one or more target scenes of the ith second sub-voice, determine the one or more target applications corresponding to the ith second sub-voice.

[0070] The method of determining the scene recognition result of the jth first sub-speech based on the jth first sub-speech and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes is the same as the processing method of determining the scene recognition result of the current speech based on the current speech and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes in the above-mentioned implementation method, and will not be repeated here.

[0071] The method of determining, based on the scene recognition result of the j-th first sub-voice, whether there are the one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes, taking the j-th first sub-voice as the ith second sub-voice, and taking the one or more target scenes matching the j-th first sub-voice as the one or more target scenes matching the ith second sub-voice, may include: judging, based on the scene recognition result of the j-th first sub-voice, whether there are the one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes, and, if there are the one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes, taking the j-th first sub-voice as the ith second sub-voice, and taking the one or more target scenes matching the j-th first sub-voice as the one or more target scenes matching the ith second sub-voice.

[0072] After determining whether there are one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes based on the scene recognition result of the j-th first sub-voice, the method may further include: if none of the one or more candidate scenes matches the j-th first sub-voice, not processing the j-th first sub-voice.

[0073] Determining one or more target applications corresponding to the i-th second sub-voice based on one or more target scenes of the i-th second sub-voice may include: determining the candidate applications associated with each target scene in the one or more target scenes of the i-th second sub-voice, and using the candidate applications associated with each target scene as each target scene of the i-th second sub-voice.

[0074] In this embodiment, the method of determining each second sub-voice, the method of determining one or more target scenes for each second sub-voice, and the method of determining one or more target applications corresponding to each second sub-voice are the same as the method of determining the i-th second sub-voice, the method of determining one or more target scenes for the i-th second sub-voice, and the method of determining one or more target applications corresponding to the i-th second sub-voice, and are not repeated here.

[0075] By adopting the above-mentioned implementation method, based on the j-th first sub-speech and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, it is determined whether the j-th first sub-speech matches each candidate scene; in the case that there is a match for one or more target scenes, the j-th first sub-speech is used as the i-th second sub-speech, and based on the one or more target scenes of the i-th second sub-speech, one or more target applications corresponding to the i-th second sub-speech are determined. In this way, based on the matching result between the speech attribute requirements corresponding to the candidate scenes associated with the candidate applications and the first sub-speech, it is determined whether the first sub-speech is the speech required by the candidate application. In this way, a more fine-grained matching of the first sub-speech and the candidate applications can be performed, thereby ensuring the accuracy of the judgment of whether each sub-speech matches the target application.

[0076] In one possible implementation, the processing of the current speech based on the one or more target applications to obtain one or more processing results of the current speech includes: processing the i-th second sub-speech based on the one or more target applications corresponding to the i-th second sub-speech to obtain one or more processing results returned by the one or more target applications corresponding to the i-th second sub-speech, wherein i is a positive integer and the i-th second sub-speech is one of the multiple first sub-speech extracted from the current speech.

[0077] Regarding the specific processing method of processing the i-th second sub-voice based on one or more target applications corresponding to the i-th second sub-voice and obtaining one or more processing results returned by the one or more target applications corresponding to the i-th second sub-voice, the specific processing method should be the same as the specific processing method of processing the current voice based on the one or more target applications and obtaining one or more processing results of the current voice as described in the aforementioned embodiment, so it will not be repeated here.

[0078] In the case where the electronic device is a smart speaker, Figure 2 The above-mentioned embodiments are exemplified as follows:

[0079] Step 1, on the smart speaker side: the singing application 201 calls the HAL (Hardware Abstraction Layer) through the Advanced Linux Sound Architecture Interaction Minilibrary 202, and sends the current accompaniment of the current song to the audio processing device through the HAL. The Advanced Linux Sound Architecture Interaction Minilibrary can also be called libtinyalsa.

[0080] Step 2, audio processing device side: the second audio processing device 203 receives the current accompaniment of the current song through the playback node 204 of the USB sound card, wherein the second audio processing device can be a DSP; the second audio processing device obtains the current voice (i.e., dry sound) through the second audio acquisition device 205; the second audio processing device generates a current mixed audio (i.e., wet sound) based on the current voice and the current accompaniment of the current song; the second audio processing device sends the current mixed audio (i.e., wet sound) and the current voice (i.e., dry sound) to the smart speaker through the recording node 206 of the USB sound card.

[0081] Step three, smart speaker side: call HAL through the advanced Linux sound architecture interactive mini-library 202 to obtain the current voice (i.e. dry sound) and the current mixed audio (i.e. wet sound), and play the current mixed audio. Specifically, the playing of the current mixed audio can be to send the current mixed audio (i.e. wet sound) to the playback device 207 through a thread, and the playback device 207 plays the current mixed audio.

[0082] Step four, smart speaker side: optionally, the scene recognition algorithm 208 determines whether the current voice matches one or more target scenes, and in the case that the current voice matches one or more target scenes, determines one or more target applications based on the one or more target scenes; in the case that the current voice does not match a target scene, does not process the current voice. Optionally, the scene recognition algorithm 208 analyzes the current voice to obtain the number of voiceprint information contained in the current voice, and in the case that multiple voiceprint information is analyzed from the current voice, extracts a first sub-voice corresponding to each voiceprint information from the current voice based on each voiceprint information in the multiple voiceprint information; determines whether each of the multiple first sub-voices matches one or more target scenes, and in the case that one or more second sub-voices of the multiple first sub-voices each match one or more target scenes, determines one or more target applications corresponding to each second sub-voice based on the one or more target scenes matched by each second sub-voice; in the case that only one voiceprint information is analyzed from the current voice, performs processing to determine the scene recognition result of the current voice based on the current voice and the voice attribute requirement of each candidate scene in the one or more candidate scenes; in the case that one or more third sub-voices of the multiple first sub-voices do not match a target scene, does not process the one or more third sub-voices.

[0083] Step five, smart speaker side: optionally, processes the current voice based on the one or more target applications to obtain one or more processing results of the current voice, such as Figure 2 The one or more target scenes can be target application (1) 209 and target application (2) 210. Optionally, processes the ith second sub-voice based on the one or more target applications corresponding to the ith second sub-voice to obtain one or more processing results returned by the one or more target applications corresponding to the ith second sub-voice. After step five is performed, step one is continued to be performed until the singing application is closed.

[0084] Figure 3 A schematic block diagram of a voice processing device provided by an embodiment of the present disclosure is shown. As shown in Figure 3 includes:

[0085] The voice acquisition module 301 is configured to acquire a current voice and a current mixed audio, and play the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song.

[0086] a target application determination module 302 for determining, when the current speech has one or more matching target scenarios, one or more target applications based on the one or more target scenarios, wherein different target scenarios in the one or more target scenarios are associated with different target applications;

[0087] The speech processing module 303 is configured to process the current speech based on the one or more target applications to obtain one or more processing results of the current speech.

[0088] The target application determination module is used to determine the scene recognition result of the current voice based on the current voice and the voice attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein different candidate scenes in the one or more candidate scenes are associated with different candidate applications, and the scene recognition result of the current voice is used to indicate whether the current voice matches each candidate scene; when it is determined based on the scene recognition result of the current voice that there are one or more target scenes that match the current voice in the one or more candidate scenes, one or more target applications are determined based on the one or more target scenes.

[0089] The target application determination module is configured to extract multiple first sub-voices from the current voice when multiple voiceprint information is obtained by analyzing the current voice, wherein different first sub-voices in the multiple first sub-voices correspond to different voiceprint information in the multiple voiceprint information; and determine one or more target applications corresponding to each second sub-voice in one or more second sub-voices in the multiple first sub-voices based on the one or more target scenarios matched by each second sub-voice, when each second sub-voice in the multiple first sub-voices has one or more matching target scenarios.

[0090] The target application determination module is configured to determine a scene recognition result of the jth first sub-voice based on the jth first sub-voice and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein j is a positive integer and the scene recognition result of the jth first sub-voice is used to indicate whether the jth first sub-voice matches each candidate scene; in the case where it is determined based on the scene recognition result of the jth first sub-voice that there are one or more target scenes matching the jth first sub-voice in the one or more candidate scenes, the jth first sub-voice is used as the ith second sub-voice, and the one or more target scenes matching the jth first sub-voice are used as the one or more target scenes matching the ith second sub-voice, wherein i is a positive integer and the ith second sub-voice is one of the one or more second sub-voices; and based on the one or more target scenes of the ith second sub-voice, determine one or more target applications corresponding to the ith second sub-voice.

[0091] The voice acquisition module is used to send the current accompaniment of the current song to the audio processing device, receive the current voice and the current mixed audio sent by the audio processing device, and play the current mixed audio.

[0092] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0093] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0094] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0095] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0096] like Figure 4As shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the device 400 can also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0097] Various components in the device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc., an output unit 407, such as various types of displays, speakers, etc., a storage unit 408, such as a magnetic disk, an optical disk, etc., and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0098] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above. For example, in some embodiments, the above-described methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, at least one step of the above-described methods can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the above-described methods by any other appropriate means, such as by means of firmware.

[0099] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0100] Program code for carrying out methods of the present disclosure can be written in any combination of at least one programming language. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a remote machine or entirely on a remote machine or server.

[0101] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include at least one line of electrical wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0103] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0104] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0105] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0106] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A speech processing method, comprising: Acquire a current voice and a current mixed audio, and play the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song; In a case where the current voice has one or more matching target scenes, determining one or more target applications based on the one or more target scenes, wherein different target scenes in the one or more target scenes are associated with different target applications; The current speech is processed based on the one or more target applications to obtain one or more processing results of the current speech.

2. The method according to claim 1, wherein In the case where the current voice has one or more matching target scenarios, determining one or more target applications based on the one or more target scenarios includes: Determining a scene recognition result for the current speech based on the current speech and speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein different candidate scenes in the one or more candidate scenes are associated with different candidate applications, and the scene recognition result for the current speech is used to indicate whether the current speech matches each candidate scene; In a case where it is determined based on the scene recognition result of the current voice that the one or more candidate scenes include the one or more target scenes matching the current voice, one or more target applications are determined based on the one or more target scenes.

3. The method according to claim 1, wherein In the case where the current voice has one or more matching target scenarios, determining one or more target applications based on the one or more target scenarios includes: In a case where multiple voiceprint information is obtained by analyzing the current speech, extracting multiple first sub-voices from the current speech, wherein different first sub-voices in the multiple first sub-voices correspond to different voiceprint information in the multiple voiceprint information; In a case where each second sub-voice in one or more second sub-voices in the plurality of first sub-voices has one or more matching target scenarios, one or more target applications corresponding to each second sub-voice are determined based on the one or more matching target scenarios of each second sub-voice.

4. The method according to claim 3, wherein: The step of determining, when each of the one or more second sub-voices in the plurality of first sub-voices has one or more matching target scenarios, one or more target applications corresponding to each second sub-voice based on the one or more matching target scenarios of each second sub-voice, includes: Determining a scene recognition result for the jth first sub-speech based on the jth first sub-speech and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, where j is a positive integer and the scene recognition result for the jth first sub-speech is used to indicate whether the jth first sub-speech matches each candidate scene; In a case where it is determined based on the scene recognition result of the j-th first sub-voice that there are one or more target scenes matching the j-th first sub-voice in the one or more candidate scenes, the j-th first sub-voice is used as the i-th second sub-voice, and the one or more target scenes matching the j-th first sub-voice are used as the one or more target scenes matching the i-th second sub-voice, wherein i is a positive integer and the i-th second sub-voice is one of the one or more second sub-voices; Based on the one or more target scenarios of the i-th second sub-speech, one or more target applications corresponding to the i-th second sub-speech are determined.

5. The method according to claim 1, wherein The acquiring of the current voice and the current mixed audio, and playing the current mixed audio, includes: The current accompaniment of the current song is sent to an audio processing device, the current voice and the current mixed audio sent by the audio processing device are received, and the current mixed audio is played.

6. A speech processing device, comprising: A voice acquisition module, configured to acquire a current voice and a current mixed audio, and play the current mixed audio, wherein the current mixed audio is generated based on the current voice and a current accompaniment of a current song; a target application determination module, configured to, if the current voice has one or more matching target scenarios, determine one or more target applications based on the one or more target scenarios, wherein different target scenarios in the one or more target scenarios are associated with different target applications; The speech processing module is used to process the current speech based on the one or more target applications to obtain one or more processing results of the current speech.

7. The device according to claim 6, wherein The target application determination module is used to determine the scene recognition result of the current voice based on the current voice and the voice attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein different candidate scenes in the one or more candidate scenes are associated with different candidate applications, and the scene recognition result of the current voice is used to indicate whether the current voice matches each candidate scene; when it is determined based on the scene recognition result of the current voice that there are one or more target scenes that match the current voice in the one or more candidate scenes, one or more target applications are determined based on the one or more target scenes.

8. The device according to claim 6, wherein The target application determination module is configured to extract multiple first sub-voices from the current voice when multiple voiceprint information is obtained by analyzing the current voice, wherein different first sub-voices in the multiple first sub-voices correspond to different voiceprint information in the multiple voiceprint information; and determine one or more target applications corresponding to each second sub-voice in one or more second sub-voices in the multiple first sub-voices based on the one or more target scenarios matched by each second sub-voice, when each second sub-voice in the multiple first sub-voices has one or more matching target scenarios.

9. The device according to claim 8, wherein The target application determination module is configured to determine a scene recognition result of the jth first sub-voice based on the jth first sub-voice and the speech attribute requirements corresponding to each candidate scene in one or more candidate scenes, wherein j is a positive integer and the scene recognition result of the jth first sub-voice is used to indicate whether the jth first sub-voice matches each candidate scene; in the case where it is determined based on the scene recognition result of the jth first sub-voice that there are one or more target scenes matching the jth first sub-voice in the one or more candidate scenes, the jth first sub-voice is used as the ith second sub-voice, and the one or more target scenes matching the jth first sub-voice are used as the one or more target scenes matching the ith second sub-voice, wherein i is a positive integer and the ith second sub-voice is one of the one or more second sub-voices; and based on the one or more target scenes of the ith second sub-voice, determine one or more target applications corresponding to the ith second sub-voice.

10. The device according to claim 6, wherein The voice acquisition module is used to send the current accompaniment of the current song to the audio processing device, receive the current voice and the current mixed audio sent by the audio processing device, and play the current mixed audio.

11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice processing method and device and distributed system

    CN111833857A

  • Audio processing method and device, electronic equipment and readable storage medium

    CN113160782A