Voice processing method, processor, voice processing apparatus, and electronic device

By identifying preset features in the speech data feature sequence and skipping instructions with unchanged intentions, the problem of slow response of electronic devices when user intentions change is solved, faster and timely response is achieved, and user experience and security are improved.

WO2025140340A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/142435
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-12-25
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

When a user sends a command to an electronic device through voice, if the intention changes, the electronic device responds slowly and does not respond in time.

Method used

By obtaining the feature sequence of speech data, identifying that the command with unchanged intentions is skipped when the preset feature is located at the tail, and waiting for the complete speech data to execute the command with the changed intentions.

Benefits of technology

It improves the response speed and timeliness of electronic devices, and enhances the user's voice interaction experience and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024142435_03072025_PF_FP_ABST
    Figure CN2024142435_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A voice processing method, a processor, a voice processing apparatus, and an electronic device, relating to the technical field of communications, and solving the problem of delayed or untimely response in an electronic device when a user changes an intent while sending a command via voice to the electronic device. The specific solution is: providing a voice processing method, which is applied to the processor. The method comprises: a processor (110) acquires first voice data (S301); and the processor (110) determines a feature sequence corresponding to the first voice data, and when the feature sequence comprises a preset feature and the preset feature is located at the tail of the feature sequence, skips execution of a first instruction corresponding to the first voice data (S302).
Need to check novelty before this filing date? Find Prior Art

Description

Voice processing method, processor, voice processing device and electronic equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 27, 2023, with application number 202311840205.6 and application name “A speech processing method, processor, speech processing device and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of communication technology, and in particular to a voice processing method, processor, voice processing device and electronic equipment. Background Art

[0003] As a new generation of user interaction mode following keyboard interaction, mouse interaction, and touch screen interaction, voice interaction has gradually become popular among users due to its convenience and speed, and is widely used in various electronic devices. Currently, when users interact with electronic devices through voice, they send commands to the electronic device through voice. The electronic device converts the voice into executable instructions and performs the corresponding operation according to the instruction. After the electronic device completes the corresponding operation, the user can send the next command to the electronic device through voice again.

[0004] However, when a user sends a command to an electronic device through voice, if the user changes his / her intention, the electronic device may have problems of slow response and untimely response. Summary of the Invention

[0005] The embodiments of the present application provide a voice processing method, processor, voice processing device, and electronic device, which solve the problem that when a user sends a command to an electronic device via voice, if the user changes his intention, the electronic device may have a slow and untimely response.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions:

[0007] According to a first aspect of an embodiment of the present application, a speech processing method is provided, applied to a processor, the method comprising: first, obtaining first speech data; then, determining a feature sequence corresponding to the first speech data; and, if the feature sequence includes a preset feature and the preset feature is at the end of the feature sequence, skipping execution of a first instruction corresponding to the first speech data.

[0008] In a possible embodiment, the above-mentioned preset feature is used to indicate a change in the user's intention.

[0009] Optionally, the preset feature may be a feature corresponding to a transition word such as "wait a minute" or "oh no" that indicates a transitional meaning, and the embodiment of the present application does not limit this.

[0010] Optionally, skipping the execution of the first instruction corresponding to the first voice data may be equivalent to not executing, abandoning execution, or ignoring execution of the first instruction corresponding to the first voice data, which is not limited in this embodiment of the present application.

[0011] Based on this solution, first, the processor obtains the first voice data. Then, the processor determines the feature sequence corresponding to the first voice data. When the feature sequence includes a preset feature and the preset feature is at the end of the feature sequence, the execution of the first instruction corresponding to the first voice data is skipped. From the user's perspective, when the user's intention changes, the processor does not execute the first instruction before the intention changes. The user will feel that the processor responds quickly and promptly. Therefore, when the processor is applied to a voice processing device, and the voice processing device is applied to an electronic device, the user will feel that the electronic device responds quickly and promptly.

[0012] In combination with the first aspect, in one possible implementation, when the first voice data includes transitional voice data corresponding to a transition word, and the transitional voice data is located at the end of the first voice data, the feature sequence includes a preset feature corresponding to the transition word, and the preset feature is located at the end of the feature sequence.

[0013] In conjunction with the first aspect, in one possible implementation, after obtaining the first voice data, the method further includes: obtaining second voice data. When it is determined that the first voice data and the second voice data are complete voice data, executing a second instruction corresponding to the second voice data.

[0014] Based on this solution, the processor obtains the second voice data, and when it determines that the first voice data and the second voice data are complete voice data, it executes the second instruction corresponding to the second voice data. From the user's perspective, the processor does not execute the first instruction before the user's intention changes, but instead executes the second instruction after the user's intention changes. The user will feel that the processor responds quickly and promptly. Secondly, when it is determined that the first voice data and the second voice data are complete voice data, the processor executes the second instruction, which can more accurately judge the user's intention and improve the user's voice interaction experience.

[0015] In combination with the first aspect, in a possible implementation, the above-mentioned execution of the second instruction corresponding to the second voice data includes: after obtaining the second voice data, and when no voice data is obtained within a preset time period, executing the second instruction.

[0016] Based on this solution, after the processor obtains the second voice data and does not obtain voice data within a preset time period, the processor can determine that the user has stopped speaking. The processor can more accurately confirm the user's intention. At this time, the processor executes the second instruction, which can further improve the user's voice interaction experience with the voice processing device 100.

[0017] In combination with the first aspect, in a possible implementation method, the above-mentioned execution of the second instruction corresponding to the second voice data includes: obtaining a first voiceprint feature corresponding to the voice data, and executing the second instruction when the similarity between the first voiceprint feature and the second voiceprint feature is greater than or equal to a preset value, where the second voiceprint feature is the voiceprint feature of the target user.

[0018] Based on this solution, the processor executes the second instruction when the similarity between the first voiceprint feature and the second voiceprint feature is greater than or equal to a preset value. When the processor is applied to a voice processing device, it can improve the security of the user's interaction with the voice processing device.

[0019] In combination with the first aspect, in a possible implementation, the determining of the feature sequence corresponding to the first speech data includes: determining the feature sequence corresponding to the first speech data by using Mel-frequency cepstral coefficients (MFCC).

[0020] According to a second aspect of an embodiment of the present application, a processor is provided, comprising an interface circuit and a computing circuit coupled to each other. The interface circuit is configured to obtain first voice data. The computing circuit is configured to determine a feature sequence corresponding to the first voice data, and if the feature sequence includes a preset feature and the preset feature is at the end of the feature sequence, skipping execution of a first instruction corresponding to the first voice data.

[0021] In combination with the second aspect, in one possible implementation, when the first voice data includes transitional voice data corresponding to a transitional word, and the transitional voice data is located at the end of the first voice data, the feature sequence includes a preset feature corresponding to the transitional word, and the preset feature is located at the end of the feature sequence.

[0022] In conjunction with the second aspect, in one possible implementation, after the interface circuit acquires the first voice data, the interface circuit is further configured to acquire the second voice data. The computing circuit is further configured to, when determining that the first voice data and the second voice data are complete voice data, execute a second instruction corresponding to the second voice data.

[0023] In conjunction with the second aspect, in a possible implementation, the computing circuit is specifically configured to execute the second instruction after acquiring the second voice data and when no voice data is acquired within a preset time period.

[0024] In combination with the second aspect, in one possible implementation, the computing circuit is specifically used to obtain a first voiceprint feature corresponding to the voice data, and when the similarity between the first voiceprint feature and the second voiceprint feature is greater than or equal to a preset value, execute a second instruction, and the second voiceprint feature is the voiceprint feature of the target user.

[0025] In combination with the second aspect, in a possible implementation, the computing circuit is specifically configured to determine a feature sequence corresponding to the first speech data through Mel-frequency cepstral coefficients (MFCCs).

[0026] In a third aspect of an embodiment of the present application, a speech processing device is provided, comprising a sound collector and a processor coupled to each other, wherein the sound collector is used to collect speech data, and the processor is the processor described in the second aspect or any possible implementation of the second aspect.

[0027] In conjunction with the third aspect, in one possible implementation, the speech processing device further includes a speech enhancement circuit coupled between the sound collector and the processor. The speech enhancement circuit is configured to perform speech enhancement processing on the speech data, where the speech enhancement processing includes at least one of acoustic echo cancellation (AEC), sound source localization (DOA), blind source separation (BSS), dereverberation, or noise reduction.

[0028] In combination with the third aspect, in a possible implementation, the processor is further configured to perform noise reduction processing on the speech data after the speech enhancement processing when the signal-to-noise ratio of the speech data after the speech enhancement processing is less than a preset signal-to-noise ratio.

[0029] In a fourth aspect of an embodiment of the present application, an electronic device is provided, which includes a mainboard and a voice processing device fixed to the mainboard, wherein the voice processing device is the voice processing device described in the third aspect or any possible implementation of the third aspect.

[0030] In a fifth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program runs on a processor, the processor executes the speech processing method described in the first aspect or any possible implementation of the first aspect.

[0031] In a sixth aspect of the embodiments of the present application, a computer program product is provided. When a processor executes the computer program product, the processor executes the speech processing method as described in the first aspect or any possible implementation of the first aspect.

[0032] The descriptions of the second to sixth aspects of this application can refer to the detailed description of the first aspect; and the beneficial effects described in the second to sixth aspects can refer to the analysis of the beneficial effects of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1 is a schematic structural diagram of a speech processing device provided in an embodiment of the present application;

[0034] FIG2 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0035] FIG3 is a flow chart of a speech processing method provided in an embodiment of the present application;

[0036] FIG4 is a flow chart of another speech processing method provided in an embodiment of the present application;

[0037] FIG5 is a schematic diagram of a method for recognizing text from speech data provided in an embodiment of the present application;

[0038] FIG6 is a schematic diagram of the structure of a processor provided in an embodiment of the present application;

[0039] FIG7 is a schematic diagram of the structure of another speech processing device provided in an embodiment of the present application;

[0040] FIG8 is a schematic structural diagram of another speech processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The following sections discuss the making and use of various embodiments in detail. However, it should be understood that many applicable inventive concepts provided herein can be implemented in a variety of specific contexts. The specific embodiments discussed are intended merely to illustrate specific ways to implement and use the present description and technology and are not intended to limit the scope of this application.

[0042] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.

[0043] Various circuits or other components may be described or referred to as being "configured to" perform one or more tasks. In this case, "configured to" is used to imply structure by indicating that the circuit / component includes structure (e.g., circuitry) that performs the one or more tasks during operation. Thus, even when a specified circuit / component is not currently operational (e.g., not turned on), the circuit / component may be referred to as being configured to perform the task. Circuits / components used with the phrase "configured to" include hardware, such as circuitry that performs an operation, etc.

[0044] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, a and b, a and c, b and c or a, b and c, where a, b and c can be single or multiple. In addition, in the embodiments of the present application, words such as "first" and "second" do not limit the quantity and order.

[0045] In this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0046] Before introducing the embodiments of the present application, the technical terms and background technologies involved in the present application are first introduced.

[0047] Mel-scale frequency cepstral coefficients (MFCC): a technique widely used in the field of speech processing to extract features of speech data.

[0048] Voiceprint: It is the sound wave spectrum that carries speech information displayed by electroacoustic instruments.

[0049] Voiceprint features: refers to the unique features of each person's voice that distinguish them from others.

[0050] Currently, when a user interacts with an electronic device through voice, the user sends a command to the electronic device through voice, the electronic device converts the voice into executable instructions, and performs corresponding operations according to the instructions. When the electronic device completes the corresponding operation, the user can send the next command to the electronic device again through voice.

[0051] However, when a user sends a command to an electronic device through voice, if the user changes his / her intention, the electronic device may have problems of slow response and untimely response.

[0052] For example, after a user sends a first command to an electronic device via voice, if the user changes their intention and sends a second command to the electronic device again via voice, and the user wants the electronic device to execute the second instruction corresponding to the second command instead of the first instruction corresponding to the first command, the electronic device will execute the first instruction corresponding to the first command instead of the second instruction corresponding to the second command. From the user's perspective, since the electronic device does not execute the second instruction corresponding to the second command in a timely manner, the user will feel that the electronic device has slow and untimely response problems.

[0053] Based on this, an embodiment of the present application provides a voice processing method, which is applied to a processor. The method obtains the user's voice data and determines the feature sequence corresponding to the voice data. When the feature sequence includes a preset feature and the preset feature is located at the end of the feature sequence, the method skips executing the instruction corresponding to the voice data, so that when the user changes his intention, he can respond in time with a faster reaction speed.

[0054] The speech processing method provided in the embodiment of the present application, when applied to a processor, can be applied to a speech processing device. As shown in FIG1 , a structural diagram of a speech processing device 100 provided in the embodiment of the present application is shown. The speech processing device 100 includes a processor 110 and a sound collector 120 coupled to each other.

[0055] The speech processing device 100 provided in the embodiment of the present application can be applied to electronic devices with speech processing functions. Exemplary types of electronic devices may include, but are not limited to: mobile phones, tablet computers, laptop computers, PDAs, mobile internet devices (MIDs), cameras, wearable devices (such as smart watches, smart bracelets, pedometers, etc.), audio equipment, audio and video players, set-top boxes, game consoles, printers, vehicle-mounted devices (such as devices on vehicles such as cars, bicycles, electric vehicles, airplanes, ships, trains, and high-speed trains), virtual reality (VR) devices, augmented reality (AR) devices, smart home devices (such as refrigerators, televisions, air conditioners, and electric meters), intelligent robots, and wireless terminals in self-driving vehicles.

[0056] For example, as shown in FIG2 , which is a structural diagram of an electronic device 200 provided in an embodiment of the present application, the electronic device 200 includes a mainboard 210 and a voice processing device 100 fixed to the mainboard 210 .

[0057] As shown in Figure 3, it is a flow chart of a speech processing method provided in an embodiment of the present application. The method can be applied to the processor 110 in the above-mentioned speech processing device 100. In the following embodiments of the present application, the method is applied to the processor 110 in the above-mentioned speech processing device 100 as an example for exemplary description. The method includes steps S301-S302.

[0058] S301: The processor 110 obtains first voice data.

[0059] S302: The processor 110 determines a feature sequence corresponding to the first voice data, and skips execution of a first instruction corresponding to the first voice data when the feature sequence includes a preset feature and the preset feature is located at the end of the feature sequence.

[0060] In a possible embodiment, the processor 110 is specifically configured to determine a feature sequence corresponding to the first speech data through Mel-frequency cepstral coefficients (MFCCs). The specific process may refer to the prior art and will not be described in detail in this embodiment of the present application.

[0061] The above-mentioned preset features are used to indicate a change in the user's intention.

[0062] Optionally, the preset feature may be a feature corresponding to a transition word such as "wait a minute" or "oh no" that indicates a transitional meaning, and the embodiment of the present application does not limit this.

[0063] For example, take the example of a user saying "Please open the music software, wait a moment" to the voice processing device 100. The sound collector 120 in the voice processing device 100 can collect the first voice data corresponding to "Please open the music software, wait a moment". The processor 110 can obtain the first voice data and determine the feature sequence corresponding to the first voice data. Since the feature sequence includes a preset feature corresponding to the transition word "wait a moment", and the preset feature is located at the end of the feature sequence, the processor 110 skips executing the first instruction corresponding to the first voice data and does not open the music software. From the user's perspective, when the user's intention changes, the processor 110 skips executing the first instruction before the intention changes. The user will feel that the voice processing device 100 reacts quickly and promptly. When the voice processing device 100 is applied to the electronic device 200, the user will feel that the electronic device 200 reacts quickly and promptly.

[0064] Optionally, skipping the execution of the first instruction corresponding to the first voice data may be equivalent to not executing, abandoning execution, or ignoring execution of the first instruction corresponding to the first voice data, which is not limited in this embodiment of the present application.

[0065] In a possible embodiment, when the first voice data includes transitional voice data corresponding to a transitional word, and the transitional voice data is located at the end of the first voice data, the feature sequence corresponding to the first voice data includes a preset feature corresponding to the transitional word, and the preset feature is located at the end of the feature sequence.

[0066] In a possible embodiment, the processor 110 may be a digital signal processing (DSP) or a microcontroller unit (MCU), which is not limited in the embodiment of the present application.

[0067] In the speech processing method provided in the embodiment of the present application, first, the processor 110 obtains the first speech data. Then, the processor 110 determines the feature sequence corresponding to the first speech data, and when the feature sequence includes a preset feature and the preset feature is located at the end of the feature sequence, the execution of the first instruction corresponding to the first speech data is skipped. From the user's perspective, when the user's intention changes, the processor 110 does not execute the first instruction before the intention changes, and the user will feel that the processor 110 reacts quickly and in a timely manner. Thus, when the processor 110 is applied to the speech processing device 100 and the speech processing device 100 is applied to the electronic device 200, the user will feel that the electronic device 200 reacts quickly and in a timely manner.

[0068] As shown in FIG4 , in a possible embodiment, after the processor 110 obtains the first voice data, the method further includes steps S303 - S304 . When executing steps S303 - S304 and step S302 , the order of execution may not be distinguished, for example, they may be executed simultaneously.

[0069] S303: The processor 110 obtains second voice data.

[0070] S304: When the processor 110 determines that the first voice data and the second voice data are complete voice data, execute a second instruction corresponding to the second voice data.

[0071] In one possible embodiment, the processor 110 is specifically configured to determine that the first and second voice data are complete voice data when the text corresponding to the first and second voice data is identified using an acoustic language model and the text completeness recognition model determines that the text is complete. The specific process of the processor 110 identifying the text using the acoustic language model and determining that the text is complete using the completeness recognition model can be referenced in the prior art and will not be further described in detail in this embodiment of the present application.

[0072] Optionally, the acoustic language model and the text completeness recognition model can be the same model, or can be different models. The embodiments of the present application do not limit this. The following embodiments of the present application use the acoustic language model and the text completeness recognition model as the same model as an example for illustrative explanation.

[0073] Optionally, the above-mentioned text completeness recognition model may include at least one of a deep neural network (DNN), a recurrent neural network (RNN), a hidden Markov model (HMM) or a transformer model (encoder-decoder architecture), and may also include existing or future other algorithms or models that can perform phoneme information and text feature extraction and classification, which is not limited in the embodiments of the present application.

[0074] For example, the processor 110 is applied to the speech processing device 100, the speech processing device 100 is applied to the electronic device 200, the electronic device 200 also includes a speaker, the text integrity recognition model is a deep neural network, and the user says to the electronic device 200 "Please open the music software, wait a moment", and then says to the electronic device 200 "Please open the picture software".

[0075] As shown in Figure 5, first, the sound collector 120 can collect first voice data corresponding to the phrase "Please open the music software, wait a moment." The processor 110 can obtain the first voice data and determine a feature sequence corresponding to the first voice data. Because the feature sequence includes a preset feature corresponding to the transition word "wait a moment," and this preset feature is located at the end of the feature sequence, the processor 110 can skip executing the first instruction corresponding to the first voice data and not open the music software. The electronic device 200 can then respond to the user through the speaker with "Hmm." The sound collector 120 can then continue to collect second voice data corresponding to the phrase "Please open the image software."

[0076] Then, the processor 110 can obtain the second voice data and recognize the text "Please open the music software, wait a moment, please open the picture software" corresponding to the first voice data and the second voice data through a deep neural network.

[0077] Finally, the processor 110 can use a deep neural network model to execute the second instruction corresponding to the second voice data to open the picture software based on the interaction between the user and the electronic device 200 when it determines that the text is a complete text. The electronic device 200 can reply to the user through the speaker "OK, the picture software has been opened for you."

[0078] It can be understood that by adopting the voice processing method provided by the embodiment of the present application, after collecting the user's first voice data, if it is recognized that the user's intention has changed, the first instruction corresponding to the first voice data will not be executed, but the user's second voice data will continue to be collected. When the text corresponding to the first voice data and the second voice data is identified and the text is determined to be a complete text through the text integrity recognition model, the second instruction corresponding to the second voice data will be executed. From the user's perspective, the electronic device 200 does not execute the first instruction before the user's intention changes, but executes the second instruction after the user's intention changes. The user will feel that the voice processing device 100 has a fast response speed and responds in a timely manner. Secondly, when it is determined that the first voice data and the second voice data are complete voice data, executing the second instruction can more accurately judge the user's intention and improve the user's voice interaction experience.

[0079] In a possible embodiment, when the processor 110 is applied to the speech processing device 100 and the speech processing device 100 is applied to the electronic device 200, the electronic device 200 may further include a memory, and the above-mentioned text integrity model may be stored in the memory, so that when the electronic device 200 is not connected to the network, the speech processing device 100 can run offline and still support simple voice interaction between the user and the electronic device 200. Moreover, since the text integrity model is stored in the memory, there is no need to transmit data over the network, which can improve the response speed of the speech processing device 100 and reduce the conversation delay between the user and the electronic device 200. Alternatively, the text integrity model can be stored in a server in the cloud. The text integrity model is not limited by computing resources and storage space and can have more powerful functions. When the electronic device 200 is connected to the network, more complex voice interaction can be achieved through the text integrity model stored in the cloud, further improving the user's voice interaction satisfaction. The embodiment of the present application does not limit whether the text integrity model is specifically stored in the memory or stored in the cloud server. For example, a smaller-scale text integrity model can be stored in the memory, and a larger-scale text integrity model can be stored in the cloud server.

[0080] In one possible embodiment, the processor 110 executes the second instruction corresponding to the second voice data, including: after obtaining the second voice data, the processor 110 executes the second instruction when no voice data is obtained within a preset time period. The specific length of the preset time period is not limited in this embodiment of the application.

[0081] In one possible embodiment, when the processor 110 is implemented in the speech processing device 100, the sound collector 120 in the speech processing device 100 may also be configured to detect silence. When the sound collector 310 detects silence for a preset duration, the processor 110 will not acquire any speech data for the preset duration. The fact that the sound collector 310 detects silence for the preset duration includes: the duration of silence detected by the sound collector 120 is equal to the preset duration, or the duration of silence detected by the sound collector 120 is greater than the preset duration. Thus, when the sound collector 310 detects silence for the preset duration and the processor 110 does not acquire any speech data for the preset duration, the processor 110 can determine that the user has stopped speaking, more accurately confirming the user's intent and further improving the user's voice interaction experience with the speech processing device 100.

[0082] For example, as shown in FIG5 , taking the preset duration as 1 second, after the sound collector 120 completes the collection of the second voice data, if the sound collector 120 detects that the duration of silence is equal to 1 second, it can be understood that the sound collector 120 did not collect any voice data within the preset duration of 1 second, and the processor 110 did not obtain any voice data within the preset duration of 1 second. The processor 110 can determine that the user has stopped speaking, execute the second instruction corresponding to the second voice data, and open the image software. The electronic device 200 can respond to the user through the speaker, "OK, the image software has been opened for you."

[0083] In one possible embodiment, the processor 110 executes the second instruction corresponding to the second voice data, including: the processor 110 obtains a first voiceprint feature corresponding to the voice data, and executes the second instruction when the similarity between the first voiceprint feature and the second voiceprint feature is greater than or equal to a preset value, wherein the second voiceprint feature is the voiceprint feature of the target user. The specific value of the preset value is not limited in this embodiment of the application. Therefore, when the processor 110 is applied to the voice processing device 100, it can improve the security of the user's interaction with the voice processing device 100.

[0084] Optionally, when the processor 110 is applied to the above-mentioned voice processing device 100, the above-mentioned first voice data and the second voice data may be data generated by the same user collected by the sound collector 120, or may be data generated by different users collected by the sound collector 120. The embodiment of the present application is not limited to this. The embodiment of the present application takes the first voice data and the second voice data as an example of data generated by the same user collected by the sound collector 120 for illustrative explanation.

[0085] For example, the processor 110 may determine a first voiceprint feature corresponding to the first voice data and the second voice data through voiceprint recognition (VPR). When the similarity between the first voiceprint feature and the second voiceprint feature is equal to a preset value, the processor 110 executes the second instruction.

[0086] In the voice processing method provided in the embodiment of the present application, after processor 110 obtains the first voice data, processor 110 obtains the second voice data. When processor 110 determines that the first voice data and the second voice data are complete voice data, processor 110 executes the second instruction corresponding to the second voice data, thereby more accurately determining the user's intention and improving the user's voice interaction experience.

[0087] Based on this, as shown in FIG6 , an embodiment of the present application further provides a processor 110 , which includes an interface circuit 111 and a computing circuit 112 coupled to each other.

[0088] The interface circuit 111 is used to obtain the first voice data. The calculation circuit 112 is used to determine the feature sequence corresponding to the first voice data, and when the feature sequence includes a preset feature and the preset feature is at the end of the feature sequence, skip the execution of the first instruction corresponding to the first voice data.

[0089] In a possible embodiment, the calculation circuit 112 is specifically configured to determine a feature sequence corresponding to the first speech data through Mel-frequency cepstral coefficients (MFCCs). The specific process may refer to the prior art and will not be described in detail in this embodiment of the present application.

[0090] The above-mentioned preset features are used to indicate a change in the user's intention.

[0091] Optionally, the preset feature may be a feature corresponding to a transition word such as "wait a minute" or "oh no" that indicates a transitional meaning, and the embodiment of the present application does not limit this.

[0092] Optionally, skipping the execution of the first instruction corresponding to the first voice data may be equivalent to not executing, abandoning execution, or ignoring execution of the first instruction corresponding to the first voice data, which is not limited in this embodiment of the present application.

[0093] In a possible embodiment, when the first voice data includes transitional voice data corresponding to the transitional word, and the transitional voice data is located at the end of the first voice data, the feature sequence includes a preset feature corresponding to the transitional word, and the preset feature is located at the end of the feature sequence.

[0094] The processor 110 provided in the embodiment of the present application obtains the first voice data through the interface circuit 111, determines the feature sequence corresponding to the first voice data through the calculation circuit 112, and skips the execution of the first instruction corresponding to the first voice data when the feature sequence includes a preset feature and the preset feature is located at the end of the feature sequence. From the user's perspective, when the user's intention changes, the processor 110 does not execute the first instruction before the intention changes, and the user will feel that the processor 110 reacts quickly and in a timely manner. Thus, when the processor 110 is applied to the voice processing device 100 and the voice processing device 100 is applied to the electronic device 200, the user will feel that the electronic device 200 reacts quickly and in a timely manner.

[0095] In a possible embodiment, after acquiring the first voice data, the interface circuit 111 is further configured to acquire the second voice data. The calculation circuit 112 is further configured to execute the second instruction corresponding to the second voice data when determining that the first voice data and the second voice data are complete voice data.

[0096] In one possible embodiment, the computing circuit 112 is specifically configured to determine that the first and second speech data are complete speech data when the text corresponding to the first and second speech data is identified using an acoustic language model and the text completeness recognition model determines that the text is complete. The specific process of the computing circuit 112 identifying the text using the acoustic language model and determining that the text is complete using the completeness recognition model can be referenced in the prior art and will not be further described in detail in this embodiment of the present application.

[0097] Optionally, the acoustic language model and the text completeness recognition model can be the same model, or can be different models. The embodiments of the present application do not limit this. The following embodiments of the present application use the acoustic language model and the text completeness recognition model as the same model as an example for illustrative explanation.

[0098] Optionally, the above-mentioned text completeness recognition model may include at least one of a deep neural network, a recurrent neural network, a hidden Markov model or a transformer model, and may also include other existing or future algorithms or models that can perform phoneme information and text feature extraction and classification. The embodiments of the present application are not limited to this.

[0099] In one possible embodiment, the computing circuit 112 is specifically configured to execute the second instruction after acquiring the second voice data and if no voice data is acquired within a preset time period. The specific duration of the preset time period is not limited in this embodiment of the present application. Thus, when the processor 110 is applied to the speech processing device 100, it can determine whether the user has stopped speaking, can more accurately confirm the user's intention, and can further improve the user's voice interaction experience with the speech processing device 100.

[0100] In one possible embodiment, the computing circuit 112 is specifically configured to obtain a first voiceprint feature corresponding to the voice data and, when the similarity between the first voiceprint feature and the second voiceprint feature is greater than or equal to a preset value, execute a second instruction. The second voiceprint feature is the voiceprint feature of the target user. The specific value of the preset value is not limited in this embodiment of the present application. Thus, when the processor 110 is implemented in the voice processing device 100, it can improve security when a user interacts with the voice processing device 100.

[0101] Optionally, when the processor 110 is applied to the above-mentioned voice processing device 100, the above-mentioned first voice data and second voice data can be data generated by the same user collected by the sound collector 120, or can be data generated by different users collected by the sound collector 120. This embodiment of the present application is not limited to this.

[0102] After acquiring the first voice data, the processor 110 and interface circuit 111 provided in the embodiment of the present application are further configured to acquire the second voice data. The computing circuit 112 is further configured to execute the second instruction corresponding to the second voice data when determining that the first voice data and the second voice data are complete voice data, thereby more accurately determining the user's intent and improving the user's voice interaction experience.

[0103] As shown in FIG1 , an embodiment of the present application further provides a speech processing device 100 , which includes a processor 110 and a sound collector 120 coupled to each other.

[0104] The sound collector 120 is used to collect voice data. The processor 110 is used to obtain the voice data collected by the sound collector 120. The structure of the processor 110 is the same as that of the processor 110 described in FIG. 6 . For a description of the processor 110 , reference can be made to the voice processing method described in FIG. 3 or FIG. 4 , as well as the description of the processor 110 described in FIG. 6 .

[0105] As shown in Figure 7, in one possible embodiment, the sound collector 120 includes at least one microphone 121 and an always-on microphone activity detection circuit (Always On MAD) 122 coupled to the at least one microphone 121. The present embodiment does not limit the specific number of microphones included in the at least one microphone 121. The following embodiment uses the example of at least one microphone 121 including four microphones as an example for illustrative description.

[0106] Among them, at least one microphone 121 is used to collect the user's voice signal and convert the voice signal into an analog signal. The always-on microphone activity detection circuit 122 is used to convert the analog signal into a digital signal and implement a voice wake-up function and / or a mute detection function based on the digital signal. The digital signal can also be called voice data, for example, it can be the first voice data mentioned above.

[0107] For example, as shown in FIG7 , assuming that at least one microphone 121 includes four microphones, the always-on microphone activity detection circuit 122 may include four analog-to-digital converters (ADCs) 1221, an anti-blocking microphone 1222, and a microphone activity detection circuit MAD 1223. One end of each of the four ADCs 1221 is coupled to each of the four microphones, and the other ends of each of the four ADCs 1221 are coupled to one end of each of the anti-blocking microphones 1222, which in turn are coupled to MAD 1223. The four ADCs 1221 are configured to convert analog signals into digital signals. The anti-blocking microphone 1222 is configured to detect whether any of the four microphones are blocked and select an unblocked microphone to collect the user's voice signal. MAD 1223 is configured to determine whether to perform voice wake-up or silence detection based on the digital signal output by the anti-blocking microphone 1222 and a preset threshold.

[0108] The speech processing device 100 provided in an embodiment of the present application can collect a user's first speech data through the sound collector 120, obtain a feature sequence corresponding to the first speech data through the processor 110, and skip the execution of the first instruction corresponding to the first speech data when the feature sequence includes a preset feature and the preset feature is at the end of the feature sequence. From the user's perspective, when the user's intention changes, the speech processing device 100 does not execute the first instruction before the intention change, and the user will feel that the speech processing device 100 responds quickly and promptly.

[0109] In a possible embodiment, as shown in Fig. 8, the speech processing device 100 further includes a memory 130 coupled to the processor 110. The memory 130 is configured to store the second voiceprint feature of the target user.

[0110] The speech processing device 100 provided in the embodiment of the present application collects the second speech data of the user through the sound collector 120, and the processor 110 executes the second instruction corresponding to the second speech data when it determines that the first speech data and the second speech data are complete speech data. From the user's perspective, the speech processing device 100 does not execute the first instruction before the user's intention changes, but executes the second instruction after the user's intention changes. The user will feel that the speech processing device 100 reacts quickly and promptly. Moreover, when it is determined that the first speech data and the second speech data are complete speech data, the processor 110 executes the second instruction, which can more accurately judge the user's intention and improve the user's voice interaction experience.

[0111] In a possible embodiment, as shown in Figure 8, the speech processing device 100 may further include a speech enhancement circuit 140 coupled between the sound collector 120 and the processor 110, and the speech enhancement circuit 140 is used to perform speech enhancement processing on the speech data, and the speech enhancement processing includes at least one of echo cancellation (acoustic echo canceller, AEC), sound source localization (direction of arrival, DOA), blind source separation (blind source separation, BSS), dereverberation or noise reduction.

[0112] Specifically, the speech enhancement circuit 140 may include an echo cancellation circuit 141, a beamforming circuit 142, a dereverberation circuit 143 and a noise reduction circuit 144 coupled in sequence, and may also include a sound source localization circuit 145 and a blind source separation circuit 146, one end of the sound source localization circuit 145 is coupled to the echo cancellation circuit 141, the other end of the sound source localization circuit 145 is coupled to the beamforming circuit 142, one end of the blind source separation circuit 146 is coupled to the echo cancellation circuit 141, and the other end of the blind source separation circuit 146 is coupled to the noise reduction circuit 144.

[0113] In the example of the speech processing device 100 being applied to the electronic device 200 described above, where the electronic device 200 includes a speaker, the echo cancellation circuit 141 is used to eliminate the sound produced by the speaker when the speaker is playing outward. The sound source localization circuit 145 is used to determine the position and direction of the sound source in space, and the beamforming circuit 142 is used to receive speech data based on the position and direction. Specifically, as shown in FIG8 , taking the example of at least one microphone 121 including four microphones, the sound source localization circuit 145 can determine the position and direction of the sound source in space and send parameters ω1-ω4 corresponding to each microphone to the beamforming circuit 142. The beamforming circuit 142 can adjust the speech data received by each microphone based on the parameters ω1-ω4 and combine the adjusted speech data to achieve speech data that enhances the direction and position of the sound source. The dereverberation circuit 143 is used to eliminate reverberation effects in the speech data. The blind source separation circuit 146 is used to distinguish between different users when multiple users interact with the speech processing device 100 through voice.

[0114] In a possible embodiment, as shown in FIG8 , the voice enhancement circuit 140 and the always-on microphone activity detection circuit 122 in the sound collector 120 may be one circuit, which may be referred to as an audio codec (AUDIO CODEC) 150. The audio codec 150 may also include other circuits, which is not limited in this embodiment of the present application.

[0115] The speech processing device 100 provided in the embodiment of the present application performs speech enhancement processing on speech data through the speech enhancement circuit 140, thereby improving the quality of the speech data, improving the purity of the speech data, and improving the accuracy of speech recognition, thereby improving the user's speech interaction satisfaction.

[0116] In a possible embodiment, the processor 110 is further configured to perform noise reduction processing on the speech data after the speech enhancement processing when the signal-to-noise ratio of the speech data after the speech enhancement processing is less than a preset signal-to-noise ratio. The embodiment of the present application does not limit the specific value of the preset signal-to-noise ratio.

[0117] In a possible embodiment, the processor 110 may perform AI noise reduction processing on the speech data after the speech enhancement processing by using an artificial intelligence (AI) noise reduction model.

[0118] Optionally, the above-mentioned AI denoising model may include at least one of a deep neural network, a convolutional neural network (CNN), a gated recurrent unit (GRU) or a counterfactual recurrent network (CRN), and may also include existing or future other algorithms or models that can perform AI denoising, which is not limited in the embodiments of the present application.

[0119] In the speech processing device 100 provided in the embodiment of the present application, the processor 110 performs noise reduction processing on the speech data after speech enhancement processing when the signal-to-noise ratio of the speech data after speech enhancement processing is less than a preset signal-to-noise ratio. This can further improve the quality of the speech data, increase the purity of the speech data, and enhance the accuracy of speech recognition, thereby further improving the user's satisfaction with the speech interaction.

[0120] Based on this, an embodiment of the present application also provides an electronic device, which may be the electronic device 200 shown in Figure 2 above. The electronic device 200 includes a mainboard 210 and a voice processing device 100 fixed to the mainboard 210. The structure of the voice processing device 100 may be the structure of the voice processing device described in any of the above Figures 1, 7 and 8.

[0121] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program code is stored. When the processor 110 executes the computer program code, the processor 110 executes the speech processing method shown in FIG. 3 or FIG. 4 .

[0122] The embodiment of the present application further provides a computer program product. When the processor 110 executes the computer program product, the processor 110 executes the speech processing method shown in FIG. 3 or FIG. 4 .

[0123] The above detailed description of the voice method and the analysis of beneficial effects can all be referred to the processor 110, the voice processing device 100, the electronic device 200, the computer-readable storage medium and the computer program product, and the embodiments of the present application will not be repeated here.

[0124] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A voice processing method, characterized in that, Applied to a processor, the method includes: Obtain first voice data; Determine a feature sequence corresponding to the first voice data. When the feature sequence includes a preset feature and the preset feature is located at the tail of the feature sequence, skip the execution of a first instruction corresponding to the first voice data.

2. The method according to claim 1, characterized in that When the first voice data includes turning voice data corresponding to a turning word and the turning voice data is located at the tail of the first voice data, the feature sequence includes a preset feature corresponding to the turning word and the preset feature is located at the tail of the feature sequence.

3. The method according to claim 1 or 2, characterized in that, After obtaining the first voice data, the method further includes: Obtain second voice data; When it is determined that the first voice data and the second voice data are complete voice data, execute a second instruction corresponding to the second voice data.

4. The method according to claim 3, wherein The execution of the second instruction corresponding to the second voice data includes: After obtaining the second voice data and when no voice data is obtained within a preset duration, execute the second instruction.

5. The method according to claim 3 or 4, characterized in that, The execution of the second instruction corresponding to the second voice data includes: Obtain a first voiceprint feature corresponding to the voice data. When the similarity between the first voiceprint feature and a second voiceprint feature is greater than or equal to a preset value, execute the second instruction, where the second voiceprint feature is the voiceprint feature of a target user.

6. The method according to any one of claims 1-5, characterized in that The determination of the feature sequence corresponding to the first voice data includes: Determine the feature sequence corresponding to the first voice data through Mel Frequency Cepstral Coefficients (MFCC).

7. A processor, characterized in that, The processor includes an interface circuit and a computing circuit that are coupled to each other; The interface circuit is configured to obtain first voice data; The computing circuit is configured to determine a feature sequence corresponding to the first voice data. When the feature sequence includes a preset feature and the preset feature is located at the tail of the feature sequence, skip the execution of a first instruction corresponding to the first voice data.

8. The processor according to claim 7, characterized in that When the first voice data includes turning voice data corresponding to a turning word and the turning voice data is located at the tail of the first voice data, the feature sequence includes a preset feature corresponding to the turning word and the preset feature is located at the tail of the feature sequence.

9. The processor according to claim 7 or 8, characterized in that, After the interface circuit obtains the first voice data; The interface circuit is further configured to obtain second voice data; The computing circuit is further configured to, when it is determined that the first voice data and the second voice data are complete voice data, execute a second instruction corresponding to the second voice data.

10. The processor according to claim 9, wherein: The computing circuit is specifically configured to, after obtaining the second voice data and when no voice data is obtained within a preset duration, execute the second instruction.

11. The processor according to claim 9 or 10, wherein: The computing circuit is specifically configured to obtain a first voiceprint feature corresponding to the voice data. When the similarity between the first voiceprint feature and a second voiceprint feature is greater than or equal to a preset value, execute the second instruction, where the second voiceprint feature is the voiceprint feature of a target user.

12. The processor according to any one of claims 7-11, characterized in that the computing circuit is specifically configured to determine the feature sequence corresponding to the first voice data through Mel-frequency cepstral coefficients (MFCC).

13. A voice processing device, characterized in that, The device includes a sound collector and a processor that are coupled to each other; the sound collector is configured to collect voice data, and the processor is the processor according to any one of claims 7-12.

14. The device according to claim 13, characterized in that, The device further includes a voice enhancement circuit coupled between the sound collector and the processor; the voice enhancement circuit is configured to perform voice enhancement processing on the voice data, and the voice enhancement processing includes at least one of acoustic echo cancellation (AEC), direction of arrival (DOA) source localization, blind source separation (BSS), dereverberation, or noise reduction.

15. The device according to claim 14, characterized in that the processor is further configured to perform noise reduction processing on the voice data after voice enhancement processing when the signal-to-noise ratio of the voice data after voice enhancement processing is less than a preset signal-to-noise ratio.

16. An electronic device, characterized in that, The electronic device includes a main board and a voice processing device fixed to the main board, and the voice processing device is the voice processing device according to any one of claims 13-15.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, the processor is caused to execute the voice processing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multimedia playing method and device and multimedia providing method and device

    CN110139155A

  • Dialogue processing method and device and computer readable storage medium

    CN110910866A

  • Multi-mode voice awakening and interrupting method and device

    CN113113009A

  • Equipment control method and device based on user intention recognition

    CN116415591A

  • Method and device for carrying out continuous variable adjustment on vehicle equipment and functions through voice control

    CN116434747A