Speech recognition method and apparatus, electronic device, computer readable storage medium

CN115331664BActive Publication Date: 2026-10-09GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110506377.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-10
Publication Date
2026-10-09
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

虽然,通过外加人声检测设的方式降低了语音识别功能的功耗,但是,增加人声检测设备,就造成了成本的增加

Benefits of technology

[0016]Initially, both the voice detection and voice recognition functions in the second chip are in sleep mode, resulting in low power consumption for the electronic device. If the first chip receives a first preset scenario trigger signal, it first wakes up the voice detection function in the second chip. The voice detection function in the second chip then checks if the speech to be recognized contains key information. If key information is detected, the voice recognition function in the second chip is then activated. Finally, voice recognition can be performed based on the voice recognition function in the second chip. This two-stage wake-up process avoids the voice detection and voice recognition functions in the second chip being constantly active, thus reducing system power consumption. Furthermore, since both the first and second chips are existing components in the electronic device, no additional components are needed, thereby reducing both system power consumption and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331664B_ABST
    Figure CN115331664B_ABST
Patent Text Reader

Abstract

The application relates to a speech recognition method and device, an electronic device and a computer readable storage medium, and is applied to an electronic device. The electronic device comprises a first chip and a second chip. The first chip is used for running a first operating system, and the second chip is used for running a second operating system. The method comprises the following steps: if a speech detection function and a speech recognition function in the second chip are in a dormant state, and the first chip receives a first preset scene trigger signal, then the speech detection function in the second chip is awakened through the first chip. If key information is detected from the speech to be recognized based on the speech detection function, then the speech recognition function in the second chip is awakened, and the speech to be recognized is subjected to speech recognition based on the speech recognition function in the second chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a speech recognition method and apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] In electronic devices with dual operating systems, the primary operating system supports complex functions, but consumes more power when running under it. The secondary operating system, on the other hand, supports simpler functions and consumes less power when running under it.

[0003] Speech recognition is a complex function that requires the main operating system to function, resulting in high power consumption in electronic devices. To reduce power consumption, traditional methods typically involve adding an external voice detection device to activate the speech recognition function within the main operating system. This means the speech recognition function is only activated when the device detects a human voice; otherwise, it remains dormant. While this reduces power consumption, it also increases costs. Furthermore, since the voice detection device only detects human voices, its use without voice detection results in wasted resources. Summary of the Invention

[0004] This application provides a speech recognition method and apparatus, an electronic device, and a computer-readable storage medium, which can reduce the power consumption of speech recognition functions without increasing costs.

[0005] A speech recognition method is applied to an electronic device, the electronic device including a first chip and a second chip, the first chip being used to run a first operating system and the second chip being used to run a second operating system, the method comprising:

[0006] When the voice detection and voice recognition functions in the second chip are in a dormant state, if the first chip receives a first preset scene trigger signal, the voice detection function in the second chip will be woken up by the first chip.

[0007] If key information is detected from the speech to be recognized based on the speech detection function, the speech recognition function in the second chip is activated.

[0008] Based on the speech recognition function in the second chip, the speech to be recognized is performed to obtain the speech recognition result.

[0009] A voice recognition device is applied to an electronic device, the electronic device including a first chip and a second chip, the first chip being used to run a first operating system, and the second chip being used to run a second operating system, the device comprising:

[0010] The voice detection function wake-up module is used to wake up the voice detection function in the second chip when the voice detection function and voice recognition function in the second chip are in a dormant state and the first chip receives a first preset scene trigger signal.

[0011] The voice recognition function wake-up module is used to wake up the voice recognition function in the second chip if key information is detected from the voice to be recognized based on the voice detection function.

[0012] The speech recognition module is used to perform speech recognition on the speech to be recognized based on the speech recognition function in the second chip to obtain a speech recognition result.

[0013] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the Bluetooth communication method as described above.

[0014] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the Bluetooth communication method as described above.

[0015] The aforementioned speech recognition method, apparatus, electronic device, and computer-readable storage medium are applied to an electronic device. The electronic device includes a first chip and a second chip. The first chip runs a first operating system, and the second chip runs a second operating system. The method includes: When the speech detection and speech recognition functions in the second chip are in a dormant state, if the first chip receives a first preset scenario trigger signal, the speech detection function in the second chip is woken up by the first chip. If key information is detected from the speech to be recognized based on the speech detection function, the speech recognition function in the second chip is woken up, and based on the speech recognition function of the second chip, speech recognition is performed on the speech to be recognized to obtain a speech recognition result.

[0016] Initially, both the voice detection and voice recognition functions in the second chip are in sleep mode, resulting in low power consumption for the electronic device. If the first chip receives a first preset scenario trigger signal, it first wakes up the voice detection function in the second chip. The voice detection function in the second chip then checks if the speech to be recognized contains key information. If key information is detected, the voice recognition function in the second chip is then activated. Finally, voice recognition can be performed based on the voice recognition function in the second chip. This two-stage wake-up process avoids the voice detection and voice recognition functions in the second chip being constantly active, thus reducing system power consumption. Furthermore, since both the first and second chips are existing components in the electronic device, no additional components are needed, thereby reducing both system power consumption and cost. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a diagram illustrating the application environment of a speech recognition method in one embodiment;

[0019] Figure 2 Here is a flowchart of a speech recognition method in one embodiment;

[0020] Figure 3 This is a schematic diagram of the internal structure of an electronic device in one embodiment;

[0021] Figure 4 Here is a flowchart of a speech recognition method in another embodiment;

[0022] Figure 5 This is a schematic diagram of the communication interfaces on the first chip, the first processor, and the second processor, and the connections between the communication interfaces in one embodiment.

[0023] Figure 6 for Figure 4 A flowchart illustrating the method for configuring communication functions between the second processor and the first chip, the first processor, and the second processor;

[0024] Figure 7 Here is a flowchart of a speech recognition method in another embodiment;

[0025] Figure 8 Here is a flowchart of a speech recognition method in a specific embodiment;

[0026] Figure 9 This is a structural block diagram of a speech recognition device in one embodiment;

[0027] Figure 10 This is a schematic diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0029] Figure 1 This is a diagram illustrating an application scenario of a speech recognition method in one embodiment. For example... Figure 1 As shown, the application environment includes an electronic device 120, which includes a first chip and a second chip. The first chip runs a first operating system, and the second chip runs a second operating system. The voice detection and voice recognition functions in the second chip of the electronic device 120 are in a dormant state. If the first chip receives a first preset scenario trigger signal, it wakes up the voice detection function in the second chip. If key information is detected from the speech to be recognized based on the voice detection function, it wakes up the voice recognition function in the second chip. Based on the voice recognition function in the second chip, it performs voice recognition on the speech to be recognized to obtain the voice recognition result. Here, the electronic device 120 can be any terminal device such as a mobile phone, tablet computer, PDA (Personal Digital Assistant), wearable device (smartwatch, etc.), smart home device, or smart vehicle device.

[0030] Traditionally, voice recognition functionality on the main operating system is activated by adding an external voice detection device. This means the voice recognition function only activates when the device detects a human voice; otherwise, it remains dormant. While this reduces power consumption, it also increases costs. Furthermore, since the device only detects human voices, its resources are wasted when voice detection is not performed.

[0031] Therefore, this application proposes a speech recognition method that reduces system power consumption without requiring additional components, thus avoiding waste of component data. The speech recognition method in this embodiment is designed to run on... Figure 1Taking electronic device 120 as an example, the electronic device includes a first chip and a second chip. The first chip runs a first operating system, and the second chip runs a second operating system. Here, the first chip can be a microcontroller unit (MCU), and the second chip can be a system-on-a-chip (SOC). The first operating system can be a real-time operating system (RTOS), which consumes less power during operation. The second operating system can be an Android system, which consumes more power during operation, and the two systems can switch between each other.

[0032] like Figure 2 The diagram shows a flowchart of a speech recognition method in one embodiment. The method includes steps 220 to 260. Wherein,

[0033] Step 220: If the voice detection function and voice recognition function in the second chip are in a dormant state, and the first chip receives a first preset scene trigger signal, the voice detection function in the second chip will be woken up by the first chip.

[0034] The voice detection function is the process of performing voice recognition to match key information. The voice recognition function is the process of performing voice recognition to obtain semantic information. If a user sends a voice message to be recognized to the electronic device, the electronic device processes that voice message. Specifically, the process of the electronic device processing the voice message can be divided into two stages: the first stage is the voice detection stage, and the second stage is the voice recognition stage.

[0035] Initially, both the voice detection and voice recognition functions in the second chip are in sleep mode, resulting in low power consumption for the electronic device. If the first chip receives a first preset scenario trigger signal, it wakes up the voice detection function in the second chip, thus entering the first stage. The voice detection function in the second chip performs voice detection on the speech to be processed, checking whether it contains key information. The speech to be processed can be the voice signal received by the electronic device from the user. The speech to be processed can be divided into two parts according to the chronological order in which the electronic device receives it: a first part and a second part. The first part, received earlier, generally contains keywords, while the second part may contain user commands. For example, if the user says to the electronic device, "Hey A, what's the weather like today?", the first part can be divided into "Hey A," and the second part into "What's the weather like today?". The first part contains the key information "Hey A."

[0036] The key information here refers to the information that can wake up the speech recognition function in the second chip, mainly in the form of keywords. For example, the wake-up keyword (the name of the speech recognition function) is pre-set for the speech recognition function in the second chip, such as Xiao A, Xiao Li, etc. This application does not limit this.

[0037] Step 240: If key information is detected from the speech to be recognized based on the speech detection function, the speech recognition function in the second chip is activated.

[0038] The speech detection function in the second chip performs speech detection on the speech to be processed, checking whether it contains key information. If key information is detected in the speech, it means the user has invoked the speech recognition function in the second chip, which is then activated to perform speech recognition on the speech. If no key information is detected in the speech, it means the user has not invoked the speech recognition function in the second chip.

[0039] Furthermore, when key information is detected from the speech to be recognized based on the speech detection function, speech detection is mainly performed on the first part of the speech to be recognized, and speech detection is not performed on the second part of the speech, because key information usually appears in the first part of the speech.

[0040] Step 260: Based on the speech recognition function in the second chip, perform speech recognition on the speech to be recognized to obtain the speech recognition result.

[0041] After activating the speech recognition function in the second chip, speech recognition can be performed on the speech to be recognized. Specifically, the second part of the speech to be recognized is recognized using the speech recognition function in the second chip. For example, performing speech recognition on the second part of the speech, "How's the weather today?", yields the speech recognition result: "Today's weather." The electronic device can then use the speech recognition result to query locally or from a server to obtain the query result and provide feedback to the user.

[0042] Specifically, the speech recognition process can be performed locally by the second chip, or the second chip can upload the speech to a server, where the server then performs the speech recognition in the cloud. This application does not limit the specific method used.

[0043] In this embodiment, the voice detection and voice recognition functions in the second chip are initially in a dormant state, resulting in low power consumption for the electronic device. If the first chip receives a first preset scenario trigger signal, it first wakes up the voice detection function in the second chip. The voice detection function in the second chip detects whether the speech to be recognized contains key information. If key information is detected in the speech, the voice recognition function in the second chip is then woken up. Finally, voice recognition can be performed based on the voice recognition function in the second chip. By using a two-stage wake-up method, the voice detection and voice recognition functions in the second chip are prevented from being constantly active, thus reducing system power consumption. Furthermore, since both the first and second chips are existing components in the electronic device, no additional components are needed, thereby reducing system power consumption and simultaneously lowering costs.

[0044] In one embodiment, the first preset scenario trigger signal is a signal sent by the electronic device to the first chip when the sensor detects an operation that triggers the first preset scenario.

[0045] Specifically, when an operation triggering the first preset scenario occurs, the signal received by the sensor changes. This operation can be initiated by the user or triggered passively; this application does not limit this. Based on the detected signal change, the sensor sends a first preset scenario trigger signal to the first chip. These first preset scenarios are pre-defined scenarios where the user will use the voice recognition function. Therefore, when the user performs an operation triggering the first preset scenario, the sensor sends the first preset scenario trigger signal to the first chip. Upon receiving the first preset scenario trigger signal, the first chip wakes up the voice detection function in the second chip. Furthermore, if key information is detected from the speech to be recognized based on the voice detection function, the voice recognition function in the second chip is then activated.

[0046] The first preset scenario includes at least one of the following: a wrist-raising scenario, a screen-on scenario, a scenario involving connection to an external device, a scenario where the user is in sleep mode, and a scenario where the user is within a preset positioning area. That is, when an operation triggers at least one of these scenarios, the sensor sends a first preset scenario trigger signal to the first chip. The wrist-raising scenario is when the user raises their wrist; the screen-on scenario is when the screen is lit. External devices include headphones, other electronic devices, etc. The preset positioning area can be pre-set, for example, a preset danger zone or warning zone can be set as the preset positioning area.

[0047] For example, a user can trigger a wrist-raising scene by raising their wrist. Therefore, when the user raises their wrist, triggering the wrist-raising scene, the signal received by the sensor changes. Based on this signal change, the sensor sends a wrist-raising scene trigger signal to the first chip. Upon receiving the first preset scene trigger signal, the first chip activates the voice detection function in the second chip.

[0048] In this embodiment, when the sensor detects an operation that triggers a first preset scenario, it sends a first preset scenario trigger signal to the first chip. Then, upon receiving the first preset scenario trigger signal, the first chip can wake up the voice detection function in the second chip. Therefore, by using the sensor and the first chip, the voice detection function in the second chip is only woken up when the first preset scenario is triggered. This avoids the voice detection function in the second chip being constantly active, thus reducing system power consumption.

[0049] In one embodiment, the second chip includes a first processor for performing a voice detection function; waking up the voice detection function in the second chip via the first chip includes:

[0050] The first chip wakes up the first processor in the second chip.

[0051] like Figure 3 The diagram shown illustrates the internal structure of an electronic device in one embodiment. The electronic device includes a sensor 320, a first chip 340, a second chip 360, and a microphone 380. The second chip 360 includes a first processor 362. The sensor 320 establishes a one-way communication connection 1 with the first chip 340, through which it can send signals to the first chip 340. The first chip 340 then establishes a one-way communication connection 2 with the first processor 362, through which it can send signals to the first processor 362.

[0052] Because the first processor is used to perform the voice detection function, when the voice detection function in the second chip is woken up by the first chip, the first processor in the second chip is actually woken up by the first chip.

[0053] Specifically, when sensor 320 detects an operation that triggers the first preset scenario, it sends a first preset scenario trigger signal to first chip 340 via one-way communication connection 1. Upon receiving the first preset scenario trigger signal, the first chip sends a wake-up signal to first processor 362 via one-way communication connection 2 to wake up the first processor 362. After being woken up, the first processor 362 receives the voice to be processed from the sound pickup device 380 and performs voice detection on the voice to be processed.

[0054] In this embodiment, because the first processor is used to perform the voice detection function, waking up the voice detection function in the second chip via the first chip actually means waking up the first processor in the second chip via the first chip. By separating the voice detection and voice recognition functions into different processors in the second chip, only the first processor needs to be woken up when the voice detection function is required. This avoids excessive waking of other processors in the second chip, thus reducing system power consumption.

[0055] In one embodiment, the second chip further includes a second processor for performing a speech recognition function; if key information is detected from the speech to be recognized based on the speech detection function, the speech recognition function in the second chip is activated, including:

[0056] If key information is detected from the speech to be recognized based on the speech detection function, the second processor is woken up by the first processor.

[0057] Combination Figure 3 The diagram shown illustrates the internal structure of an electronic device in one embodiment. The second chip 360 includes a first processor 362 and a second processor 364. The sensor 320 establishes a one-way communication connection 1 with the first chip 340, through which it can send signals to the first chip 340. The first chip 340 then establishes a one-way communication connection 2 with the first processor 362, through which it can send signals to the first processor 362. The first processor 362 then establishes a one-way communication connection 3 with the second processor 364, through which it can send signals to the second processor 364.

[0058] A bidirectional communication connection is established between the second processor 364 and the first processor 362 (as shown by the thick arrow in the figure). This bidirectional communication connection 4 enables data transfer between the second processor 364 and the first processor 362. A bidirectional communication connection 5 is also established between the second processor 364 and the first chip 340 (as shown by the thick arrow in the figure). This bidirectional communication connection also enables data transfer between the second processor 364 and the first chip 340. Here, the first processor 362 is an audio digital signal processor (aDSP), and the second processor 364 is an application processor (AP).

[0059] Because the first processor is used to perform the speech detection function, waking up the speech detection function in the second chip via the first chip actually means waking up the first processor in the second chip via the first chip. After the first processor is woken up, it detects whether the speech to be processed contains key information. If key information is detected in the speech to be recognized, the speech recognition function in the second chip is woken up. Specifically, the second processor is woken up by the first processor, and then performs speech recognition on the speech to be recognized by the second processor.

[0060] Specifically, when sensor 320 detects an operation that triggers the first preset scenario, it sends a first preset scenario trigger signal to first chip 340 via one-way communication connection 1. Upon receiving the first preset scenario trigger signal, the first chip sends a wake-up signal to first processor 362 via one-way communication connection 2 to wake up first processor 362. After being woken up, first processor 362 receives the speech to be processed from microphone 380 and performs speech detection. If key information is detected in the speech to be recognized, first processor 362 sends a wake-up signal to second processor 364 via one-way communication connection 3 to wake up second processor 364. After being woken up, second processor 364 can perform speech recognition on the speech to be recognized.

[0061] In this embodiment, because the first processor performs the speech detection function and the second processor performs the speech recognition function, waking up the speech detection function in the second chip via the first chip actually wakes up the first processor in the second chip via the first chip. Similarly, waking up the speech recognition function in the second chip actually wakes up the second processor via the first processor. By separating the speech detection and speech recognition functions into different processors within the second chip, only the first processor needs to be woken up when the speech detection function is required. This avoids excessive waking up of other processors in the second chip. Only when the speech recognition function is further needed is the second processor specifically woken up. This reduces system power consumption.

[0062] In one embodiment, a speech recognition method is provided, further comprising:

[0063] In response to the user's activation of the voice recognition function, the second processor is controlled to configure the communication function of the first chip, the first processor, and the second processor; wherein, the communication function configuration is used to configure the signal flow direction and the initial signal transmitted when the first chip, the first processor, and the second processor transmit signals.

[0064] After configuration is complete, control the second processor to enter hibernation mode.

[0065] Specifically, such as Figure 4 As shown, a speech recognition method is provided, including:

[0066] Step 410: In response to the user's operation to enable the voice recognition function, control the second processor to configure the communication function of the first chip, the first processor and the second processor.

[0067] Specifically, a voice recognition function switch is provided on the electronic device, which users can use to turn the voice recognition function on or off. When the user activates the voice recognition function, the electronic device responds by controlling the second processor to configure the communication function between the first chip, the first processor, and the second processor. Here, the communication function configuration is mainly used to configure the signal flow and the initial signal transmitted between the first chip, the first processor, and the second processor. Specifically, the signal flow is configured so that the first chip sends signals unidirectionally to the first processor, and the first processor sends signals unidirectionally to the second processor. If the initial signal is represented by a voltage level, it can be configured to be high or low. This application does not limit this.

[0068] After the second processor configures the communication functions of the first chip, the first processor, and the second processor, the configured first chip, the first processor, and the second processor can wake up the first processor in the second chip through the first chip, and wake up the second processor in the second chip when the first processor detects key information from the speech to be recognized, and then perform speech recognition on the speech to be recognized through the second processor.

[0069] Step 430: After configuration is complete, control the second processor to enter hibernation mode.

[0070] After the second processor configures the communication functions of the first chip, the first processor, and the second processor, if no other processes are running on the second processor, the second processor can be controlled to enter a sleep state to reduce the power consumption of the electronic device.

[0071] Step 450: When the first processor and the second processor in the second chip are in a sleep state, if the first chip receives a first preset scene trigger signal, the first processor in the second chip will be woken up by the first chip.

[0072] In response to the user's activation of the voice recognition function, after the second processor configures the communication functions of the first chip, the first processor, and the second processor, the first processor and the second processor in the second chip are in a sleep state. The first chip detects whether it has received a first preset scene trigger signal. If the first chip receives the first preset scene trigger signal, it wakes up the first processor in the second chip.

[0073] Step 470: If key information is detected from the speech to be recognized by the first processor, the second processor in the second chip is woken up.

[0074] After the first processor is woken up, it detects whether the speech to be recognized contains key information. If the first processor detects key information from the speech to be recognized, it wakes up the second processor.

[0075] Step 490: Based on the second processor in the second chip, perform speech recognition on the speech to be recognized.

[0076] Once the second processor is activated, it can perform speech recognition on the speech to be recognized.

[0077] In this embodiment, when the user activates the voice recognition function, the electronic device responds by controlling the second processor to configure the communication functions of the first chip, the first processor, and the second processor. This enables the configured first chip, first processor, and second processor to wake up the first processor in the second chip via the first chip, and, if the first processor detects key information in the speech to be recognized, wake up the second processor in the second chip to perform voice recognition. By enabling the voice recognition function, the second processor is only controlled to configure the communication functions of the first chip, the first processor, and the second processor when the user activates the function. When the user does not activate the function, no communication function configuration is required, and the electronic device will not support the voice detection function (i.e., voice recognition). Therefore, the power consumption of the electronic device is reduced from the source.

[0078] Following the previous embodiment, as Figure 5 The diagram shows the communication interfaces on the first chip, the first processor, and the second processor, and the connections between these interfaces. The first chip 340 includes a first communication interface 340a, the first processor 362 includes a second communication interface 362a and a third communication interface 362b, and the second processor 364 includes a fourth communication interface 364a. Specifically, the first chip 340 communicates with the second communication interface 362a on the first processor 362 via its first communication interface 340a, and the third communication interface 362b on the first processor 362 communicates with the fourth communication interface 364a on the second processor 364.

[0079] Among them, such as Figure 6 As shown, step 410 involves controlling the second processor to configure communication functions for the first chip, the first processor, and the second processor, including:

[0080] Step 412: Through the second processor, configure the initial signal of the first communication interface and configure the signal flow between the first communication interface and the second communication interface to flow from the first chip to the first processor;

[0081] Step 414: Through the second processor, configure the initial signal of the third communication interface and configure the signal flow between the third communication interface and the fourth communication interface to flow from the first processor to the second processor.

[0082] Specifically, when the user activates the voice recognition function, the electronic device responds by controlling the second processor to configure the communication functions of the first chip, the first processor, and the second processor. First, the second processor configures the initial signal of the first communication interface and the signal flow between the first and second communication interfaces as from the first chip to the first processor. That is, the first chip can send signals unidirectionally to the first processor. Here, if the initial signal is represented by a voltage level, the initial voltage level of the first communication interface can be configured to be high or low. This application does not limit this. When the voltage level of the signal on the first communication interface is the same as that of the initial signal, the first processor is in a sleep state.

[0083] Because the first chip can wake up the first processor and also put the first processor into sleep mode, it can send two different signals unidirectionally to the first processor. One signal is used to wake up the first processor, and the other signal is used to put the first processor into sleep mode.

[0084] Secondly, the second processor configures the initial signal of the third communication interface and the signal flow between the third and fourth communication interfaces to be from the first processor to the second processor. That is, the first processor can send signals unidirectionally to the second processor. Similarly, if the initial signal is represented by a voltage level, the initial voltage level of the third communication interface can be configured to be high or low. This application does not limit this. When the voltage level of the signal on the third communication interface is the same as that of the initial signal, the second processor is in a sleep state.

[0085] Because the first processor can only wake up the second processor but cannot control the second processor to go into sleep mode, the first processor can only send a one-way signal to the second processor to wake it up.

[0086] In this embodiment, when configuring the communication functions of the first chip, the first processor, and the second processor, the second processor mainly configures four communication interfaces: the first communication interface, the second communication interface, the third communication interface, and the fourth communication interface. Initial signals for the first and third communication interfaces are configured so that these initial signals can be used as a reference for subsequent updates. This allows the first chip to send signals unidirectionally to the first processor, and the first processor to send signals unidirectionally to the second processor. Thus, the first chip can wake up the first processor, and the first processor can wake up the second processor. This two-stage wake-up process avoids the voice detection and voice recognition functions in the second chip from being constantly active, thereby reducing system power consumption.

[0087] In one embodiment, waking up a first processor in a second chip via a first chip includes:

[0088] The first chip is controlled to update the initial signal of the first communication interface to the target signal;

[0089] The first processor is woken up based on the target signal.

[0090] The initial signal of the first communication interface is a pre-configured initial level value by the second processor. When both the first and second processors in the second chip are in a sleep state, if the first chip receives a first preset scenario trigger signal, it will wake up the first processor in the second chip.

[0091] Specifically, the first chip updates the initial signal of the first communication interface to the target signal, and wakes up the first processor based on the target signal. The target signal is the opposite of the initial signal; for example, if the initial signal is low, the target signal is high. At this time, the first chip sends the target signal to the first processor through the first communication interface. After the first processor receives the target signal through its second communication interface, it is woken up based on the target signal. After the first processor is woken up, it detects whether the speech to be recognized contains key information. If the first processor detects key information in the speech, it wakes up the second processor.

[0092] In this embodiment, when the first processor in the second chip is woken up by the first chip, the first chip updates the initial signal of the first communication interface to the target signal, and wakes up the first processor based on the target signal. That is, when the signal of the first communication interface has the same level as the initial signal, the first processor is in a sleep state; when the signal of the first communication interface is the target signal, i.e., has the same level as the target signal, the first processor is woken up and put into operation. Thus, by controlling the signal of the first communication interface through the first chip, the state of the first processor is switched. The first processor in the second chip is woken up only when the first chip receives a first preset scenario trigger signal. This reduces the power consumption of the electronic device.

[0093] Following the previous embodiment, after waking up the first processor based on the target signal, the method further includes:

[0094] The first processor controls the activation of the microphone to collect the speech to be recognized.

[0095] The first processor is controlled to execute a speech detection algorithm to identify whether there is key information in the speech to be recognized.

[0096] Specifically, when the first and second processors in the second chip are in a sleep state, if the first chip receives a first preset scenario trigger signal, it updates the initial signal of the first communication interface to the target signal. The first chip then wakes up the first processor based on the target signal. After the first processor is woken up, it controls the activation of the microphone to collect the speech to be recognized. The first processor then executes a speech detection algorithm to identify whether key information exists in the speech.

[0097] The sound pickup device is a device used to collect sound; for example, it includes at least one of a microphone, headphones, or recording equipment. The speech to be recognized is the speech emitted by the user. The speech detection algorithm can be a deep learning-based algorithm; this application does not limit its scope.

[0098] For example, if the first processor controls the microphone to be turned on to collect the user's voice message "Hey A, how's the weather today?", then the first processor will execute a voice detection algorithm to identify whether the key information "A" exists in the voice message "Hey A, how's the weather today?". This can be done by recognizing the voice message to obtain text, and then comparing the obtained text with the text corresponding to the key information "A" to determine whether the key information exists in the voice message.

[0099] In this embodiment, after waking up the first processor based on the target signal, the first processor controls the activation of the microphone to collect the speech to be recognized. The first processor executes a speech detection algorithm to identify whether key information exists in the speech. If key information is detected in the speech, the second processor is woken up, and the second processor performs speech recognition on the speech. This two-stage wake-up avoids the first and second processors in the second chip being constantly active, thus reducing system power consumption. Furthermore, since both the first and second chips are existing components in the electronic device, no additional components are needed, thus reducing system power consumption and cost.

[0100] In one embodiment, waking up a second processor in a second chip via a first processor includes:

[0101] The first processor is controlled to update the initial signal of the third communication interface to the target signal;

[0102] The second processor is woken up based on the target signal.

[0103] The initial signal of the third communication interface is a pre-configured initial level value by the second processor. When the first and second processors in the second chip are in a sleep state, if the first chip receives a first preset scenario trigger signal, it wakes up the first processor in the second chip. If the first processor detects key information in the speech to be recognized, it wakes up the second processor in the second chip to perform speech recognition.

[0104] Specifically, the first processor updates the initial signal of the third communication interface to the target signal, and wakes up the second processor based on the target signal to perform speech recognition on the speech to be recognized. The target signal is the opposite of the initial signal; for example, if the initial signal is low, the target signal is high. The first processor then sends the target signal to the second processor through the third communication interface. After the fourth communication interface on the second processor receives the target signal, the second processor is woken up. Once the second processor is woken up, speech recognition can be performed on the speech to be recognized.

[0105] In this embodiment, when the first processor wakes up the second processor in the second chip to perform speech recognition on the speech to be recognized, the first processor updates the initial signal of the third communication interface to the target signal. Then, the second processor is woken up based on the target signal to perform speech recognition on the speech to be recognized. That is, when the signal of the third communication interface has the same level value as the initial signal, the second processor is in a sleep state; when the signal of the third communication interface is the target signal, that is, when its level value is the same as the target signal, the second processor is woken up and put into operation. Thus, by controlling the signal of the third communication interface through the first processor, the state of the second processor is switched. Only when the first processor detects key information in the speech to be recognized will the first processor wake up the second processor. This reduces the power consumption of the electronic device.

[0106] In one embodiment, a speech recognition method is provided, further comprising:

[0107] If the voice detection function fails to detect key information from the speech to be recognized, the first processor is controlled to poll the signal level of the second communication interface.

[0108] If the first chip receives the second preset scenario trigger signal, it updates the signal level value of the second communication interface to the preset level value; the preset scenario trigger signal is the signal sent by the electronic device to the first chip when the sensor detects an operation that triggers the second preset scenario.

[0109] If the polled signal level is found to be a preset level, the first processor is controlled to enter a sleep state.

[0110] like Figure 7 As shown, a speech recognition method is provided, including:

[0111] Step 702: In response to the user's operation to enable the voice recognition function, control the second processor to configure the communication function of the first chip, the first processor and the second processor.

[0112] Specifically, the second processor configures the initial signal of the first communication interface and the signal flow between the first and second communication interfaces as flowing from the first chip to the first processor. The second processor also configures the initial signal of the third communication interface and the signal flow between the third and fourth communication interfaces as flowing from the first processor to the second processor.

[0113] Step 704: After configuration is complete, control the second processor to enter sleep mode;

[0114] Step 706: When the first processor and the second processor in the second chip are in a sleep state, if the first chip receives a first preset scene trigger signal, the first processor in the second chip will be woken up by the first chip.

[0115] Step 708: Determine whether the first processor has detected key information from the speech to be recognized;

[0116] Step 710: If key information is detected from the speech to be recognized by the first processor, the second processor is woken up by the first processor to perform speech recognition on the speech to be recognized and obtain the speech recognition result.

[0117] Step 712: If the first processor does not detect key information from the speech to be recognized, then control the first processor to poll the signal level value of the second communication interface.

[0118] Specifically, since the first processor did not detect any key information in the speech to be recognized, it indicates that no key information was present in the collected speech. Therefore, while continuing to collect and detect the speech, the first processor is controlled to poll the signal level of the second communication interface.

[0119] Step 714: If the first chip receives the second preset scenario trigger signal, it updates the signal level value of the second communication interface to the preset level value; the second preset scenario trigger signal is the signal sent by the electronic device to the first chip when the sensor detects the operation that triggers the second preset scenario.

[0120] Specifically, during the speech recognition process of the second processor, when the user triggers an operation that initiates a second preset scenario, the sensor sends a second preset scenario trigger signal to the first chip. Upon receiving the second preset scenario trigger signal, the first chip updates the signal level of the second communication interface to a preset level value. The second preset scenario is the opposite of the first preset scenario. For example, if the first preset scenario is a wrist-raising scenario, then the second preset scenario is a wrist-lowering scenario; if the first preset scenario is a screen-on scenario, then the second preset scenario is a screen-off scenario; if the first preset scenario is a scenario of connecting to an external device, then the second preset scenario is a scenario of disconnecting from the external device; if the first preset scenario is a scenario where the user is in a sleep state, then the second preset scenario is a scenario where the user is not in a sleep state.

[0121] Step 716: If the polled signal level value is a preset level value, control the first processor to enter sleep mode.

[0122] If the signal level of the second communication interface obtained by polling is consistent with the signal level of the initial signal, the first processor is controlled to shut down the voice detection algorithm and enter sleep mode.

[0123] In this embodiment, after the first processor is woken up, if the first processor does not detect key information from the speech to be recognized, it polls the signal level of the second communication interface. When the signal level is a preset value, the first processor switches back to sleep mode. This achieves timely switching back to sleep mode when no key information is detected from the speech and the signal level is a preset value, thereby reducing the power consumption of the electronic device.

[0124] In one embodiment, a speech recognition method is provided, further comprising:

[0125] If key information is detected from the speech to be recognized based on the speech detection function, the first processor is controlled to enter a sleep state.

[0126] In this embodiment, the first processor and the second processor in the second chip are in a sleep state. If the first chip receives a first preset scenario trigger signal, it wakes up the first processor in the second chip. If the first processor detects key information from the speech to be recognized, it controls the first processor to shut down the speech detection algorithm and enter a sleep state, while waking up the second processor in the second chip. Based on the second processor in the second chip, speech recognition is performed on the speech to be recognized. Because key information has been detected, it is no longer necessary to repeatedly detect key information, so the first processor is promptly controlled to shut down the speech detection algorithm and enter a sleep state. This reduces the power consumption of the electronic device.

[0127] In one embodiment, a speech recognition method is provided, further comprising:

[0128] In response to the user's operation to turn off the voice recognition function, the second processor sends a shutdown command to the first processor and the first chip respectively;

[0129] Based on the shutdown command, the first processor is controlled to pause receiving the voice to be recognized, and the first chip is controlled to shut down the first communication interface.

[0130] Specifically, a voice recognition function switch is provided on the electronic device, which users can use to turn the voice recognition function on or off. When the user turns the voice recognition function off, the electronic device responds by sending a shutdown command to the first processor and the first chip via the second processor. Based on the shutdown command, the first processor is controlled to shut down the microphone to pause receiving the voice to be recognized, and the first chip is controlled to shut down the first communication interface.

[0131] Specifically, the first chip shuts down its first communication interface, thus preventing it from receiving signals. Consequently, the second communication interface on the first processor also fails to receive signals, and therefore, the first processor cannot send signals to the second processor's fourth communication interface via the third communication interface. As a result, the first chip cannot wake up the first processor, and the first processor cannot wake up the second processor; both processors are in a sleep state.

[0132] In this embodiment, in response to the user's action to disable the voice recognition function, a second processor sends a shutdown command to both the first processor and the first chip. Based on the shutdown command, the first processor is controlled to pause receiving the voice to be recognized, and the first chip is controlled to shut down the first communication interface. After the first communication interface is shut down, both the first and second processors enter a sleep state. Therefore, after the user disables the voice recognition function, the power consumption of the electronic device is reduced.

[0133] In a specific embodiment, such as Figure 8 As shown, a speech recognition method is provided to run on Figure 1 The method is described using an electronic device 120 as an example. The electronic device includes a first chip and a second chip. The first chip runs a first operating system, and the second chip runs a second operating system. Here, the first chip can be a microcontroller unit (MCU), and the second chip can be a system-on-a-chip (SOC). The second chip includes a first processor and a second processor, wherein the first processor is an aDSP, and the second processor is an AP. This speech recognition method includes:

[0134] Step 802, in response to the user's operation to enable the voice recognition function;

[0135] Step 804: Control the second processor to configure communication functions for the first chip, the first processor, and the second processor;

[0136] Step 806: Start the sensor through the first chip, and continuously detect whether an operation that triggers the first preset scenario has occurred;

[0137] Step 808: If the sensor detects an operation that triggers the first preset scene, then send the first preset scene trigger signal to the first chip;

[0138] Step 810: The first chip receives a first preset scene trigger signal and wakes up the first processor in the second chip.

[0139] Step 812: The first processor controls the activation of the sound pickup device to collect the speech to be recognized;

[0140] Step 814: Control the first processor to execute the speech detection algorithm to identify whether there is key information in the speech to be recognized;

[0141] Step 816: If the first processor detects key information from the speech to be recognized, it will wake up the second processor in the second chip.

[0142] Step 818: Based on the second processor in the second chip, perform speech recognition on the speech to be recognized.

[0143] Step 820, in response to the user's action of turning off the voice recognition function;

[0144] Step 822: The second processor sends shutdown commands to the first processor and the first chip respectively;

[0145] Step 824: Based on the shutdown command, control the first processor to pause receiving the voice to be recognized, and control the first chip to shut down the first communication interface.

[0146] In this embodiment, because the first processor performs the speech detection function and the second processor performs the speech recognition function, waking up the speech detection function in the second chip via the first chip actually wakes up the first processor in the second chip via the first chip. Similarly, waking up the speech recognition function in the second chip actually wakes up the second processor via the first processor. By separating the speech detection and speech recognition functions into different processors within the second chip, only the first processor needs to be woken up when the speech detection function is required. This avoids excessive waking up of other processors in the second chip. Only when the speech recognition function is further needed is the second processor specifically woken up. This reduces system power consumption.

[0147] In one embodiment, such as Figure 9 As shown, a voice recognition device 900 is provided, applied to an electronic device. The electronic device includes a first chip and a second chip. The first chip is used to run a first operating system, and the second chip is used to run a second operating system. The device includes:

[0148] The voice detection function wake-up module 920 is used to wake up the voice detection function in the second chip when the voice detection function and voice recognition function in the second chip are in a dormant state and the first chip receives a first preset scene trigger signal.

[0149] The voice recognition function wake-up module 940 is used to wake up the voice recognition function in the second chip if key information is detected from the voice to be recognized based on the voice detection function.

[0150] The speech recognition module 960 is used to perform speech recognition on the speech to be recognized based on the speech recognition function in the second chip to obtain the speech recognition result.

[0151] In one embodiment, the first preset scenario trigger signal is a signal sent by the electronic device to the first chip when the sensor detects an operation that triggers the first preset scenario.

[0152] In one embodiment, the second chip includes a first processor for performing a voice detection function; and a voice detection function wake-up module 920 for waking up the first processor in the second chip via the first chip.

[0153] In one embodiment, the second chip further includes a second processor for performing a speech recognition function; the speech recognition function wake-up module 940 is used to wake up the second processor through the first processor if key information is detected from the speech to be recognized based on the speech detection function.

[0154] In one embodiment, a voice recognition device 900 is provided, the device further comprising:

[0155] The voice recognition function activation and configuration module is used to respond to the user's activation operation of the voice recognition function and control the second processor to configure the communication function of the first chip, the first processor and the second processor; wherein, the communication function configuration is used to configure the signal flow direction and the initial signal transmitted when transmitting signals between the first chip, the first processor and the second processor; after the configuration is completed, the second processor is controlled to enter the sleep state.

[0156] In one embodiment, the first chip communicates with the second communication interface on the first processor through a first communication interface on the first chip, and the third communication interface on the first processor communicates with the fourth communication interface on the second processor.

[0157] The voice recognition function activation and configuration module is also used to configure the initial signal of the first communication interface and configure the signal flow between the first communication interface and the second communication interface to flow from the first chip to the first processor via the second processor; and to configure the initial signal of the third communication interface and configure the signal flow between the third communication interface and the fourth communication interface to flow from the first processor to the second processor via the second processor.

[0158] In one embodiment, the voice detection function wake-up module 920 is further configured to control the first chip to update the initial signal of the first communication interface to the target signal; and wake up the first processor based on the target signal.

[0159] In one embodiment, a voice recognition device is provided, further comprising:

[0160] The speech acquisition module is used to control the activation of the microphone via the first processor to acquire the speech to be recognized.

[0161] The key information recognition module is used to control the first processor to execute the speech detection algorithm to identify whether there is key information in the speech to be recognized.

[0162] In one embodiment, the voice recognition function wake-up module 940 is further configured to control the first processor to update the initial signal of the third communication interface to the target signal; and wake up the second processor based on the target signal.

[0163] In one embodiment, a voice recognition device is provided, further comprising:

[0164] The second communication interface polling module is used to control the first processor to poll the signal level value of the second communication interface if no key information is detected from the speech to be recognized based on the speech detection function.

[0165] The second preset scene triggering module is used to update the signal level value of the second communication interface to a preset level value if the first chip receives the second preset scene triggering signal; the second preset scene triggering signal is a signal sent to the first chip by the electronic device when the sensor detects an operation that triggers the second preset scene;

[0166] The first processor sleep control module is used to control the first processor to enter sleep mode if the polled signal level value is a preset level value.

[0167] In one embodiment, the voice recognition function wake-up module 940 is further configured to control the first processor to enter a sleep state if key information is detected from the voice to be recognized based on the voice detection function.

[0168] In one embodiment, a voice recognition device is provided, further comprising:

[0169] The voice recognition function shutdown module is used to respond to the user's shutdown operation of the voice recognition function by sending shutdown commands to the first processor and the first chip respectively through the second processor; based on the shutdown command, it controls the first processor to pause receiving the voice to be recognized and controls the first chip to shut down the first communication interface.

[0170] In one embodiment, a voice recognition device is provided, further comprising:

[0171] The operation execution module is used to execute the target operation corresponding to the speech recognition result based on the speech recognition result.

[0172] In one embodiment, the first preset scenario includes at least one of the following: a wrist-raising scenario, a screen-on scenario, a scenario of connecting to an external device, a scenario where the user is in a sleep state, and a scenario where the user is in a preset location area.

[0173] In one embodiment, the first processor is an audio processor and the second processor is an application processor.

[0174] In one embodiment, the first chip is a microcontroller unit (MCU) and the second chip is a system-on-a-chip (SOC).

[0175] It should be understood that although the steps in the flowchart above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the above diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0176] The division of the various modules in the above-described speech recognition device is only for illustrative purposes. In other embodiments, the speech recognition device can be divided into different modules as needed to complete all or part of the functions of the above-described speech recognition device.

[0177] For specific limitations regarding the speech recognition device, please refer to the limitations on the speech recognition method above, which will not be repeated here. Each module in the aforementioned speech recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0178] In one embodiment, an electronic device is also provided, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of a speech recognition method provided in the above embodiments.

[0179] Figure 10 This is a schematic diagram of the internal structure of an electronic device in one embodiment. For example... Figure 10 As shown, the electronic device includes a processor and a memory connected via a system bus. The processor provides computing and control capabilities to support the operation of the entire electronic device. The memory may include non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. These computer programs can be executed by the processor to implement a speech recognition method provided in the various embodiments described above. The internal memory provides a cached runtime environment for the operating system computer programs in the non-volatile storage media. The electronic device can be any terminal device such as a mobile phone, tablet computer, PDA (Personal Digital Assistant), POS (Point of Sales), in-vehicle computer, or wearable device.

[0180] The implementation of each module in the speech recognition device provided in this application embodiment can be in the form of a computer program. This computer program can run on an electronic device or electronic device. The program modules constituted by this computer program can be stored in the electronic device or the memory of the electronic device. When the computer program is executed by a processor, it implements the steps of the method described in the embodiments of this application.

[0181] This application also provides a computer-readable storage medium. One or more non-volatile computer-readable storage media containing computer-executable instructions, which, when executed by one or more processors, cause the processors to perform the steps of a speech recognition method.

[0182] A computer program product containing instructions that, when run on a computer, causes the computer to perform a speech recognition method.

[0183] Any references to memory, storage, database, or other media used in the embodiments of this application may include non-volatile and / or volatile memory. Suitable non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which is used as external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0184] The above speech recognition embodiments are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A speech recognition method, characterized in that, Applied to an electronic device, the electronic device including a first chip and a second chip, the first chip being used to run a first operating system, and the second chip being used to run a second operating system, the method includes: When the voice detection and voice recognition functions in the second chip are in a dormant state, if the first chip receives a first preset scenario trigger signal, it will wake up the voice detection function in the second chip. The voice detection function is a process of performing voice recognition to match key information. The first preset scenario trigger signal is sent to the first chip by the electronic device when it detects an operation that triggers the first preset scenario through a sensor. When the sensor detects an operation that triggers the first preset scenario, the signal received by the sensor changes. The first preset scenario includes one or more of the following: wrist raise scenario, screen light-up scenario, connection to an external device scenario, user in sleep mode scenario, and user in a preset positioning area scenario. If key information is detected from the speech to be recognized based on the speech detection function, the speech recognition function in the second chip is activated. The speech recognition function is the process of performing speech recognition on the speech to obtain semantic information. Based on the speech recognition function in the second chip, the speech to be recognized is performed to obtain the speech recognition result.

2. The speech recognition method according to claim 1, characterized in that, The second chip includes a first processor, which performs a voice detection function; waking up the voice detection function in the second chip via the first chip includes: The first chip wakes up the first processor in the second chip.

3. The speech recognition method according to claim 2, characterized in that, The second chip further includes a second processor, which is used to perform a speech recognition function; the step of activating the speech recognition function in the second chip if key information is detected from the speech to be recognized based on the speech detection function includes: If key information is detected from the speech to be recognized based on the speech detection function, the second processor is woken up by the first processor.

4. The speech recognition method according to claim 3, characterized in that, The method further includes: In response to the user's activation of the voice recognition function, the second processor is controlled to configure the communication function of the first chip, the first processor, and the second processor; wherein, the communication function configuration is used to configure the signal flow direction and the initial signal transmitted when the first chip, the first processor, and the second processor transmit signals. After configuration is complete, control the second processor to enter sleep mode.

5. The speech recognition method according to claim 4, characterized in that, The first chip communicates with the second communication interface on the first processor through the first communication interface on the first chip, and the third communication interface on the first processor communicates with the fourth communication interface on the second processor. The configuration of communication functions between the second processor and the first chip, the first processor, and the second processor includes: The second processor configures the initial signal of the first communication interface and configures the signal flow between the first communication interface and the second communication interface to flow from the first chip to the first processor. The second processor configures the initial signal of the third communication interface and configures the signal flow between the third communication interface and the fourth communication interface to flow from the first processor to the second processor.

6. The speech recognition method according to claim 5, characterized in that, The step of waking up the first processor in the second chip via the first chip includes: Control the first chip to update the initial signal of the first communication interface to the target signal; The first processor is awakened based on the target signal.

7. The speech recognition method according to claim 6, characterized in that, After waking up the first processor based on the target signal, the method further includes: The first processor controls the activation of the sound pickup device to collect the speech to be recognized; The first processor is controlled to execute a speech detection algorithm to identify whether the key information exists in the speech to be identified.

8. The speech recognition method according to claim 5, characterized in that, The step of waking up the second processor in the second chip via the first processor includes: Control the first processor to update the initial signal of the third communication interface to the target signal; The second processor is woken up based on the target signal.

9. The speech recognition method according to claim 5, characterized in that, The method further includes: If the speech detection function fails to detect key information from the speech to be recognized, the first processor is controlled to poll the signal level value of the second communication interface. If the first chip receives the second preset scenario trigger signal, it updates the signal level value of the second communication interface to the preset level value. The second preset scenario trigger signal is the signal sent by the electronic device to the first chip when the sensor detects the operation that triggers the second preset scenario. If the polled signal level is found to be a preset level value, the first processor is controlled to enter a sleep state.

10. The speech recognition method according to claim 2, characterized in that, The method further includes: If key information is detected from the speech to be recognized based on the speech detection function, the first processor is controlled to enter a sleep state.

11. The speech recognition method according to claim 5, characterized in that, The method further includes: In response to the user's operation to turn off the voice recognition function, the second processor sends a shutdown command to the first processor and the first chip respectively; Based on the shutdown command, the first processor is controlled to pause receiving the voice to be recognized, and the first chip is controlled to shut down the first communication interface.

12. The speech recognition method according to claim 1, characterized in that, The method further includes: Execute the target operation corresponding to the speech recognition result based on the speech recognition result.

13. The method according to claim 3, characterized in that, The first processor is an audio processor, and the second processor is an application processor.

14. The method according to any one of claims 1-13, characterized in that, The first chip is a microcontroller unit (MCU), and the second chip is a system-on-a-chip (SOC).

15. A voice recognition device, characterized in that, Applied to an electronic device, the electronic device includes a first chip and a second chip, the first chip being used to run a first operating system, and the second chip being used to run a second operating system, the device comprising: The voice detection function wake-up module is used to wake up the voice detection function in the second chip when the voice detection function and voice recognition function in the second chip are in a sleep state, if the first chip receives a first preset scenario trigger signal. The voice detection function is a process of performing voice recognition to match key information. The first preset scenario trigger signal is sent to the first chip by the electronic device when the sensor detects an operation that triggers the first preset scenario. When the sensor detects an operation that triggers the first preset scenario, the signal received by the sensor changes. The first preset scenario includes one or more of the following: wrist raise scenario, screen light-up scenario, connection to external device scenario, user in sleep state scenario, and user in preset positioning area scenario. The voice recognition function wake-up module is used to wake up the voice recognition function in the second chip if key information is detected from the voice to be recognized based on the voice detection function. The voice recognition function is the process of performing voice recognition on the voice to obtain semantic information. The speech recognition module is used to perform speech recognition on the speech to be recognized based on the speech recognition function in the second chip to obtain a speech recognition result.

16. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, the processor performs the steps of the speech recognition method as described in any one of claims 1 to 14.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the speech recognition method as described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Low-power-consumption mobile terminal awakening method and device

    CN103686962A

  • Switching control method and device for voice wake-up recognition

    CN110223691A