Control device and audio processing method
By setting up multiple audio acquisition units in the control device and combining them into multi-channel signals, the time-consuming problem under the multi-application mutually exclusive communication mechanism is solved, which improves the user's voice interaction experience and reduces power consumption.
Patent Information
- Application Number
- CN202210106949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-28
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-01-28
AI Technical Summary
When multiple applications adopt mutually exclusive communication mechanisms, the implementation of voice interaction functions takes a long time, affecting the user experience.
A plurality of audio acquisition units are arranged in the control device, and the mono audio signals collected by the plurality of audio acquisition units are combined into a multi-channel audio signal through the first controller, and the second controller detects the wake-up content to wake up the first controller in a sleep state, realizing simultaneous voice pickup of multiple applications.
The simultaneous voice pickup of multiple applications is realized, reducing the time-consuming of the mutually exclusive communication mechanism, improving the user's voice interaction function experience, and reducing power consumption through the sleep state.
Smart Images

Figure CN114596853B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and more specifically, to a control device and an audio processing method. Background Art
[0002] With the development of artificial intelligence technology, the interaction mode is also constantly changing, from simple interaction to the joint development of multiple interaction modes. Among them, the role of voice interaction function is becoming more and more obvious. For example, in the display device scenario, the user can input voice commands into the remote control (such as "play TV series aaa"), the remote control transmits the voice commands to the display device, and the display device responds to the voice commands input by the user, and then plays the TV series the user wants to watch.
[0003] Electronic devices usually have multiple applications installed, and most of these applications support voice interaction functions. Currently, when multiple applications in electronic devices use the voice interaction function, a mutually exclusive communication mechanism is usually established. When a new user needs to use the function, the current user needs to be notified to turn off the voice pickup function and stop voice interaction. After the current user has completed the shutdown, the new user is notified again that the voice pickup function can be used to start voice interaction, and the new user then turns on the voice pickup function.
[0004] In the above manner, when multiple users adopt a mutually exclusive communication mechanism, the implementation of a set of mutually exclusive mechanisms takes a long time, which seriously affects the user experience of using the voice interaction function. Summary of the Invention
[0005] In order to solve the above technical problems, the present application provides a control device and an audio processing method.
[0006] In a first aspect, the present application provides a control device, comprising: a plurality of audio acquisition units and a first controller, wherein the plurality of audio acquisition units are electrically connected to the first controller;
[0007] The first controller is configured to: when the control device is in an awake state, send a first clock signal to each of the multiple audio acquisition units, and obtain the mono audio signals respectively acquired by the multiple audio acquisition units according to the first clock signal;
[0008] The mono-channel audio signals collected by the multiple audio collection units are converted into multi-channel audio signals, so that the target application installed in the control device obtains the audio signals of the corresponding channels from the multi-channel audio signals.
[0009] As a possible implementation manner, the first controller is specifically configured as follows:
[0010] According to the correspondence between the multiple audio collection units and the channel positions, the mono audio signals respectively collected by the multiple audio collection units and the echo signals respectively corresponding to the multiple audio collection units are merged to obtain the multi-channel audio signal.
[0011] As a possible implementation manner, the control device further includes a second controller, wherein the second controller is electrically connected to the multiple audio acquisition units;
[0012] The second controller is configured to: when the control device is in a dormant state, send a second clock signal to each of the plurality of audio acquisition units, and control the plurality of audio acquisition units to respectively acquire a mono audio signal according to the second clock signal;
[0013] And detect whether the mono audio signals respectively collected by the multiple audio collection units include preset voice wake-up content.
[0014] As a possible implementation manner, the second controller is further configured to:
[0015] When the preset voice wake-up content is detected according to the mono audio signals respectively collected by the multiple audio collection units, a wake-up instruction is sent to the first controller to wake up the first controller.
[0016] As a possible implementation manner, the first controller is specifically configured as follows:
[0017] According to the wake-up instruction, controlling the control device to switch from a sleep state to a wake-up state;
[0018] After the control device switches to the awake state, the power supply of the second controller is cut off, and the first clock signal is sent to the multiple audio collection units respectively.
[0019] As a possible implementation, if the target application is a far-field voice service, the target application cyclically obtains the signal corresponding to each audio acquisition unit in the multi-channel audio signal to perform wake-up word detection;
[0020] If the target application is a near-field voice service, the target application performs channel separation on the multi-channel audio signal, extracts the audio signal of the target channel, obtains a voice command based on the audio signal of the target channel, and responds to the voice command.
[0021] As a possible implementation, the multi-channel audio signal is an audio signal in a pulse code modulation (PCM) format.
[0022] In a second aspect, the present application provides an audio processing method, applied to the control device according to any one of the first aspects, wherein the control device includes multiple audio acquisition units; the method includes:
[0023] Sending a first clock signal to each of the plurality of audio acquisition units;
[0024] Acquire mono audio signals collected by the multiple audio collection units respectively according to the first clock signal;
[0025] The mono audio signals respectively collected by the multiple audio collection units are converted into multi-channel audio signals, so that the target application installed on the control device obtains the audio signals of the corresponding channels from the multi-channel audio signals.
[0026] As a possible implementation manner, converting the mono audio signals respectively collected by the multiple audio collection units into multi-channel audio signals includes:
[0027] According to the correspondence between the multiple audio collection units and the channel positions, the mono audio signals respectively collected by the multiple audio collection units and the echo signals respectively corresponding to the multiple audio collection units are merged to obtain the multi-channel audio signal.
[0028] In a third aspect, the present application provides a readable storage medium, comprising: computer program instructions;
[0029] When at least one processor of an electronic device executes the computer program instructions, the electronic device implements the human eye protection control method as shown in the second aspect.
[0030] In a fourth aspect, the present application provides a computer program product. When the computer program product is executed by an electronic device, the electronic device implements the audio processing method shown in the second aspect.
[0031] In a fifth aspect, the present application provides an electronic device, comprising: a memory and a processor;
[0032] The memory is configured to store computer program instructions;
[0033] The processor is configured to execute the computer program instructions so that the electronic device implements the audio processing method as shown in the second aspect.
[0034] In a sixth aspect, the present application also provides a chip system, wherein the chip system includes: a processor; when the processor executes computer program instructions stored in the memory, the electronic device executes the audio processing method shown in the second aspect.
[0035] An embodiment of the present application provides a control device and an audio processing method, wherein multiple audio acquisition units are provided in the control device, and the audio signals respectively collected by the multiple audio acquisition units are merged into a multi-channel audio signal, and the multi-channel audio signal is provided as a shared audio signal to applications with voice pickup requirements in the control device, so that each application with voice requirements extracts the audio signal of the required channel from the shared audio signal, thereby solving the problem of long time consumption and poor user experience when using a mutually exclusive communication mechanism for voice interaction services. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0038] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to one or more embodiments of the present application;
[0039] Figure 2 is a hardware configuration block diagram of a control device 100 according to one or more embodiments of the present application;
[0040] Figure 3 is a software configuration block diagram of the control device 100 according to one or more embodiments of the present application;
[0041] Figure 4 A system framework diagram for performing audio processing according to one or more embodiments of the present application;
[0042] Figure 5 For Figure 4 A schematic diagram of a flow chart of audio processing based on the illustrated embodiment;
[0043] Figure 6a and Figure 6b Schematic diagrams of data structures of multi-channel audio signals obtained by audio processing according to one or more embodiments of the present application;
[0044] Figure 7 A system framework diagram for performing audio processing according to one or more embodiments of the present application;
[0045] Figure 8 For Figure 7A schematic diagram of a flow chart of audio processing based on the illustrated embodiment. DETAILED DESCRIPTION
[0046] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0047] Based on the exemplary embodiments described in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative work shall fall within the scope of protection of the claims attached to this application. In addition, although the disclosure in this application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete embodiment separately. It should be noted that the brief description of the terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0048] At present, when multiple applications use a mutually exclusive communication mechanism to implement voice interaction, the implementation of the mutual exclusion mechanism takes a long time, which seriously affects the user experience of using the voice interaction function. In order to solve this problem, the present application provides an audio data sharing solution that supports multiple applications to collect audio data at the same time according to their respective voice pickup requirements, and will not interfere with each other, thereby solving the time-consuming problem caused by the use of a mutually exclusive communication mechanism, and thus improving the user experience of using the voice interaction function.
[0049] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to one or more embodiments of the present application, such as Figure 1 As shown, the user can operate the display device 200 through the mobile terminal 300 and the control device 100. The control device 100 can be a remote control, and the communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. The user can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc., or the user can input a wake-up command, wake up the voice interaction function, and then input user commands through voice input to control the display device 200. In some embodiments, a mobile terminal, tablet computer, computer, laptop computer, and other smart devices can also be used to control the display device 200. The user can input user commands to the smart device through voice input or touching the display screen of the smart device to control the display device 200.
[0050] In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function. The display device 200 also communicates data with the server 500 through a variety of communication methods. The display device 200 can be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN) and other networks. The server 500 can provide various content and interactions to the display device 200. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing a broadcast receiving television function, the display device 200 can also provide an intelligent network TV function that provides computer support functions.
[0051] Figure 2 Schematically shows a block diagram of the configuration of the control device 100 according to an exemplary embodiment. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 150, a memory, and a power supply. The control device 100 can receive user input commands and convert them into commands that the display device 200 can recognize and respond to, acting as an intermediary for interaction between the user and the display device 200. As mentioned above, the user input commands can be input via voice. The communication interface 130 is used for external communication and includes at least one of a Wi-Fi chip, a Bluetooth module, NFC, or an alternative module.
[0052] The user input / output interface 150 includes at least one of a microphone, a touchpad, a sensor, a button, or an alternative module. The control device 100 provided herein includes multiple microphones, i.e., the control device 100 may include two or more microphones, each of which is used to collect ambient sound and multi-channel re-collected signals. The controller 110 may combine the ambient sound and multi-channel re-collected signals collected by the multiple microphones into multi-channel audio data. Multiple applications installed in the control device may then simultaneously extract desired audio data from the multi-channel audio data.
[0053] Figure 3 FIG. 1 is a block diagram of the software configuration of the control device 100 according to one or more embodiments of the present application. Figure 3 As shown in the figure, the system is divided into four layers, from top to bottom: the application layer (referred to as the "application layer"), the application framework layer (referred to as the "framework layer"), the Android runtime (Android runtime) and system library layer (referred to as the "system runtime library layer"), and the kernel layer.
[0054] The kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0055] The application layer includes at least one of the following applications: far-field voice service, near-field voice service, and remote control voice service. Far-field voice service is a voice service that enables voice interaction through voice wake-up without the user pressing a button on the control device; near-field voice service is a voice service that enables voice interaction by pressing a button on the control device; and remote control voice service is a function in which the control device transmits the user's voice input to the display device, and the display device performs audio processing on the voice content.
[0056] Reference Figure 3 The software framework of the control device shown in the figure, if the operating system of the control device is an Android system, the control device can adopt the ALSA (Advanced Linux Sound Architecture) standard audio architecture in the Android system to provide an audio acquisition solution for each application with voice pickup requirements included in the application layer of the control device, so there is no need to additionally develop application software to implement the audio acquisition solution.
[0057] It should be noted that the audio processing method provided in this application is not limited to application Figure 1 and Figure 3 The scenario shown in the embodiment, for example, can also be applied to other types of devices that can provide voice services, such as display devices with voice functions, smart phones, IPADs, laptops, etc.
[0058] In the following embodiments, the control device is taken as an example, and combined with Figures 1 to 3 The illustrated embodiment introduces the audio processing method provided by the present application by way of example.
[0059] Figure 4 This is a system framework diagram for performing audio processing according to one or more embodiments of the present application. In this embodiment, an example is provided in which a control device includes two microphones, namely microphone 1 and microphone 2, and the control device includes a first controller for controlling microphones 1 and 2 to respectively collect ambient sound and to combine the ambient sound transmitted by microphones 1 and 2.
[0060] Reference Figure 4As shown, microphone 1 and microphone 2 can be electrically connected to the first controller through the PDM interface. The first controller can transmit a clock signal to microphone 1 and microphone 2 through the PDM interface. Microphone 1 and microphone 2 can respectively collect ambient sound according to the rising edge and falling edge of the clock signal sent by the first controller. The ambient sounds collected by microphone 1 and microphone 2 are both mono audio signals.
[0061] The first controller can read the ambient sound collected by microphone 1 and microphone 2 and obtain two-channel return signals, and merge the two mono-channel ambient sounds and the two-channel return signals into a four-channel audio signal.
[0062] Among them, the two echo signals correspond one-to-one to the ambient sounds collected by microphone 1 and microphone 2 respectively. The echo signals can be used for echo cancellation processing to ensure that the target application can obtain noise-eliminated audio data, thereby ensuring the accuracy of subsequent voice command recognition.
[0063] When the first controller merges two monophonic ambient sounds and two channels of collected signals, it can pre-set the correspondence between each channel of signal and the channel position, and merge each channel of signal according to the pre-set correspondence. For example, the first two channels correspond to the two collected signals, and the last two channels correspond to the ambient sounds collected by microphone 1 and microphone 2 respectively; or, the first two channels correspond to the ambient sounds collected by microphone 1 and microphone 2 respectively, and the last two channels correspond to the two collected signals; or, the first two channels correspond to the ambient sound collected by microphone 1 and the collected signal corresponding to microphone 1, and the last two channels correspond to the ambient sound collected by microphone 2 and the collected signal corresponding to microphone 2; it should be noted that the correspondence between each channel of signal and the channel position is not limited to this example.
[0064] The first controller combines the multiple mono audio signals and the echo signal to obtain a multi-channel audio signal, that is, a 4-channel audio signal, which can be transmitted to a designated service node included in the user space.
[0065] As a possible implementation method, an audio sharing service node can be deployed as an intermediate node in the application layer of the control device. The audio sharing service node is the designated service node in the user space mentioned above. The audio sharing service node can carry multi-channel audio signals. Each application in the user space can access the audio sharing service node and obtain the required audio signal from the audio sharing service node. In addition, the audio sharing service node can also perform channel separation, echo cancellation, gain control and other filtering and noise reduction processing on the multi-channel audio signal based on the requirements of the target application with voice pickup needs for the audio signal, to ensure that the target application can obtain an audio signal with higher audio quality.
[0066] It should be noted that it is implemented in the form of an audio sharing service node. The audio sharing service node can be understood as a server, and the target application can be understood as a client. The transmission of audio signals between the target application and the audio sharing service node can be achieved through inter-process communication. When the audio sharing service node is started, the voice pickup requirement of the application can be detected. When it is detected that the application has a voice pickup requirement, the audio acquisition unit is started through the first controller to perform voice pickup, and the relevant information of the application with voice pickup requirement is added to the activated application list, and a data buffer is allocated for the application. The data buffer is used to cache the collected audio signal for the application to return it to the application, so that the application can perform wake-up content detection, voice command recognition, voice command response, etc. according to the audio signal.
[0067] It should be noted that the relevant information of the application with recording requirements mentioned here may include: application name, sampling rate, channels, bit depth, etc. The relevant information of the application with voice pickup requirements can be determined in combination with the requirements of the Android recording implementation (AudioRecord) interface parameters. Among them, bit depth can be understood as the number of bits or bytes occupied by the audio signal collected by the audio acquisition unit. The sampling rate can be understood as the frequency at which the audio acquisition unit collects the audio signal.
[0068] Please continue to refer to Figure 4 As shown, it is assumed that the services with voice requirements in the application layer of the control device include far-field voice services and near-field voice services. The application layer may also include other voice services, which are not limited in this disclosure; when the far-field voice service has a voice pickup requirement, a data buffer a is allocated for the far-field voice service. The far-field voice service can read the ambient sounds collected by microphone 1 and microphone 2 respectively from the audio signal obtained from the audio sharing service node in a cyclic manner, and copy the read data to data buffer a, and then based on the data in data buffer a, perform wake-up content detection on the ambient sound collected by microphone 1 and the ambient sound collected by microphone 2 respectively, that is, detect whether the ambient sound collected by microphone 1 and the ambient sound collected by microphone 2 include preset wake-up content. When the wake-up content is detected, the far-field voice service can further obtain the user's voice command and respond to the user's voice command.
[0069] When the near-field voice service has a voice pickup requirement, a data buffer b is allocated for the near-field voice service. The near-field voice service can read the audio data corresponding to the required microphone from the audio signal obtained by the audio sharing service node, and copy the read data to the data buffer b. The near-field voice service performs voice interaction based on the audio data in the data buffer b. For example, the near-field voice service needs to obtain the ambient sound collected by microphone 1, then the near-field voice service can obtain the audio signals of the second channel and the fourth channel from the data buffer b, and perform filtering and noise reduction processing such as echo cancellation and gain control. After that, the near-field voice service can recognize voice commands and respond to voice commands. For example, the near-field voice service can transmit voice commands to the display device through the control device, and the display device responds to the voice commands.
[0070] If other applications in the user space of the control device have voice pickup requirements, they can be processed in a similar manner.
[0071] In addition, it should be noted that if an exception occurs in the target application during voice pickup through the ALSA standard voice architecture, such as an abnormal exit, the shared audio service node can delete the target application from the activated application list based on the exception information.
[0072] Afterwards, the control device can use other audio acquisition solutions to achieve voice interaction.
[0073] Figure 5 For Figure 4 Schematic diagram of the flow of audio processing based on the embodiment shown. Figure 5 As shown, the method shown in this embodiment includes:
[0074] S501: In the awake state, the first controller sends a first clock signal to each of the plurality of audio acquisition units.
[0075] Combine Figure 4 As shown, the multiple audio acquisition units can be electrically connected to the first controller through the PDM interface, and the first controller can send the first clock signal to the multiple audio acquisition units respectively through the PDM interface.
[0076] Correspondingly, the multiple audio acquisition units respectively receive the first clock signal sent by the first controller.
[0077] The awakening state mentioned here refers to a state in which the control device is in normal working state. When the control device is in normal working state, the first controller and each audio acquisition unit included in the control device are in normal working state.
[0078] The first clock signal is a control signal configured by the first controller to control the multiple audio acquisition units to collect ambient sound. This application does not limit parameters such as the period, duty cycle, and signal level of the first clock signal. For example, the first clock signal can be a square wave signal with a 50% duty cycle, and the audio acquisition units can collect ambient sound based on the rising and falling edges of the square wave signal.
[0079] Assuming that the control device includes two audio collection units (microphones), microphone 1 can collect ambient sound according to the rising edge of the first clock signal, and microphone 2 can collect ambient sound according to the falling edge of the first clock signal. Of course, microphone 1 can also collect ambient sound according to the falling edge of the first clock signal, and microphone 2 can collect ambient sound according to the rising edge of the first clock signal.
[0080] If the control device includes a larger number of audio acquisition units, the multiple audio acquisition units can be divided into two groups, and the first controller sends a first clock signal to the multiple audio acquisition units respectively, and the multiple audio acquisition units determine whether to collect ambient sound according to the rising edge of the first clock signal or the falling edge according to the audio acquisition unit group to which they belong; for example, the control device includes 4 audio acquisition units, among which audio acquisition unit 1 and audio acquisition unit 2 are a group, and audio acquisition unit 3 and audio acquisition unit 4 are a group, audio acquisition unit 1 and audio acquisition unit 2 collect ambient sound according to the rising edge of the first clock signal, and audio acquisition unit 3 and audio acquisition unit 4 collect ambient sound according to the falling edge of the first clock signal.
[0081] S502: The plurality of audio collection units respectively send the mono audio signals collected according to the first clock signal to the first controller.
[0082] Correspondingly, the first controller receives the mono audio signal collected by each audio collection unit, where the mono audio signal is the ambient sound.
[0083] S503: The first controller converts the mono-channel audio signals collected by the multiple audio collection units into multi-channel audio signals.
[0084] The first controller can combine the multiple mono audio signals and the corresponding echo signal of each audio collection unit according to the correspondence between the multiple audio collection units and the channel positions to obtain a multi-channel audio signal. This application does not limit the implementation method of the combination.
[0085] Assume that the sampling bit depth is 16 (i.e. 16 bits represent one sampling point), i.e. one sampling point occupies 2 bytes. When merging channels, the corresponding audio signals can be written to the sampling points in a pre-set order. Figure 4In the embodiment shown, when the control device includes two microphones, the mono audio signal 1 collected by microphone 1, the mono audio signal 2 collected by microphone 2, the collected signal 1 corresponding to microphone 1, and the collected signal 2 corresponding to microphone 2 can be merged in such a manner that the first two channels are collected signals and the last two channels are ambient sounds to obtain a four-channel audio signal.
[0086] For example, referring to Figure 6a As shown, the data format of the merged 4-channel audio signal can be: AABBAABB, where AA represents a sampling point corresponding to microphone 1, which is two bytes; BB represents a sampling point corresponding to microphone 2, which also occupies two bytes; the first sampling point is used to write the corresponding echo signal of microphone 1, the second sampling point is used to write the corresponding echo signal of microphone 2, the third sampling point is used to write the corresponding ambient sound of microphone 1, and the fourth sampling point is used to write the corresponding ambient sound of microphone 2; when the corresponding ambient sound and echo signal at one moment are written, continue to write the sampling points corresponding to microphone 1 and microphone 2 at the next moment until the voice pickup is completed.
[0087] exist Figure 6a Based on the shown embodiment, the first sampling point can also be used to write the ambient sound corresponding to microphone 1, the second sampling point is used to write the ambient sound corresponding to microphone 2, the third sampling point is used to write the echo signal corresponding to microphone 1, and the fourth sampling point is used to write the echo signal corresponding to microphone 2.
[0088] For example, refer to Figure 6b In the embodiment shown, the data format of the merged 4-channel audio signal can be: AAAABBBB, wherein the first sampling point is used to write the collected signal corresponding to microphone 1, the second sampling point is used to write the ambient sound collected by microphone 1, the third sampling point is used to write the collected signal corresponding to microphone 2, and the fourth sampling point is used to write the ambient sound corresponding to microphone 2; when the ambient sound and the collected signal corresponding to a moment are written, the sampling points corresponding to microphone 1 and microphone 2 at the next moment are continued to be written until the voice pickup is completed.
[0089] exist Figure 6b Based on the shown embodiment, the first sampling point can also be used to write the ambient sound corresponding to microphone 1, the second sampling point is used to write the return signal corresponding to microphone 1, the third sampling point is used to write the ambient sound corresponding to microphone 2, and the fourth sampling point is used to write the return signal corresponding to microphone 2.
[0090] After merging the 4-channel audio signals using any of the above methods, if the application needs to extract the audio signal from microphone 1, it only needs to extract the first two bytes of data for every 4 bytes to obtain the ambient sound and echo signal collected by microphone 1.
[0091] It should be noted that the position of the ambient sound and the echo signal corresponding to each microphone in the combined multi-channel audio signal can also be adjusted, and is not limited to the above. Figure 6a as well as Figure 6b The method shown in the embodiment; in addition, if the control device includes more microphones, they can also be merged in a similar manner, for example, writing the ambient sound and echo signal collected by each microphone in the order of the microphone groups.
[0092] S504: The target application obtains the audio signal of the corresponding channel from the multi-channel audio signal.
[0093] The target application may include one or more applications in the user space of the control device, such as Figure 4 In the illustrated embodiment, the target application may include a far-field voice service and / or a near-field voice service. Of course, the target application may also include other applications, such as a remote control voice function, where the remote control voice function refers to the function of controlling the device to transmit the collected audio data to a connected electronic device (such as a display device).
[0094] When the target application includes multiple applications, each application can read the audio signal corresponding to the required channel from the multi-channel audio signal according to its own needs, such as combining Figure 4 In the illustrated embodiment, the target application includes a far-field voice service and a near-field voice service. The far-field voice service can cyclically extract the audio signal corresponding to microphone 1 and the audio signal corresponding to microphone 2 from the 4-channel audio signal to perform wake-up content detection. The near-field service can extract the audio signal corresponding to microphone 1 from the 4-channel audio signal and recognize the voice commands in the audio signal.
[0095] It should be noted that when the target application includes multiple applications, each application may extract the required audio signal from the multi-channel audio signal simultaneously or in the order in which the applications are triggered.
[0096] As a possible implementation, the control device may store the multi-channel audio signal in an audio sharing service node, and the target application accesses the audio sharing service node through inter-process communication to obtain the audio signal corresponding to the required channel.
[0097] The method provided in this embodiment provides multiple audio acquisition units in a control device, and merges the audio signals respectively collected by the multiple audio acquisition units into a multi-channel audio signal, which is provided as a shared audio signal to applications with recording requirements in the control device. This allows each application with recording requirements to extract the audio signal of the required channel from the shared audio signal, thereby solving the problem of long time consumption and poor user experience when using a mutually exclusive communication mechanism for voice interaction services.
[0098] Since the voice interaction function is implemented by the application in the user space of the control device, in order to realize the voice interaction function, the control device needs to be in an awake state at all times. In this state, the power consumption of the control device is relatively large. When the control device is powered by a battery, the working time of the control device is greatly reduced. If the user space is stopped, the voice interaction function will not be realized.
[0099] To solve this problem, the present application puts the user space and the kernel space of the control device into hibernation, sets a chip mounted on the first controller to detect the wake-up content, and when the wake-up content is detected, sends a signal to the first controller to switch the control device from the sleep state to the wake-up state, and realizes the voice interaction function through the first controller.
[0100] Below through Figure 7 as well as Figure 8 The illustrated embodiment details the system framework of the control device and the audio processing flow when setting the mounted chip.
[0101] in, Figure 7 A system framework diagram for performing audio processing according to one or more embodiments of the present application. Figure 7 The embodiment shown in Figure 4 Based on the embodiment shown, the present invention further includes: a second controller. The present application does not limit the type of the second controller. For example, the second controller may be a digital signal processing (DSP) chip.
[0102] The second controller is connected to microphone 1 and microphone 2 via a second PDM interface. When the control device is in a dormant state, the second controller can send a clock signal to microphone 1 and microphone 2 via the second PDM interface, so that microphone 1 and microphone 2 collect ambient sound based on the clock signal sent by the second controller. In addition, the second controller is also used to detect wake-up content based on the ambient sound collected by microphone 1 and microphone 2. The second controller can detect wake-up content based on the audio signals collected by microphone 1 and microphone 2 in a cyclic manner.
[0103] When the second controller detects that the collected audio signal contains preset wake-up content, the second controller can send a signal to the first controller to trigger the first controller and the voice interaction service in the user space to work normally.
[0104] Afterwards, the first controller implements the voice interaction function in the same way as Figure 4 、 Figure 5 、 Figure 6a as well as Figure 6b Similar to the embodiment shown, please refer to the aforementioned Figure 4 、 Figure 5 、 Figure 6a as well as Figure 6b For the sake of brevity, the description of the illustrated embodiment will not be repeated here.
[0105] Figure 8 For Figure 7 Schematic diagram of the flow of audio processing based on the embodiment shown. Figure 8 As shown, the method provided in this embodiment includes:
[0106] S801: A second controller sends a second clock signal to multiple audio acquisition units.
[0107] The control device powers on the second controller, and the first controller loads the firmware program (firmwave) corresponding to the second controller to the second controller, and configures multiple audio acquisition units to be controlled by the second controller. After that, the first controller can enter the sleep state, that is, the control device enters the sleep state, and the second controller enters the wake-up content detection mode.
[0108] The second controller may send a second clock signal to the multiple audio acquisition units. The present disclosure does not limit the second clock signal. For example, the second clock signal may be a square wave signal similar to the first clock signal.
[0109] S802: Multiple audio collection units respectively collect mono audio signals according to the second clock signal.
[0110] S803: The multiple audio collection units send the collected mono audio signals to the second controller.
[0111] That is, the audio signals collected by the multiple audio collection units are sent to the second controller.
[0112] S804: The second controller performs wake-up content detection according to the mono audio signal sent by each audio acquisition unit.
[0113] When the preset wake-up content is detected, step S805 is executed.
[0114] S805: The second controller sends a wake-up instruction to the first controller.
[0115] Exemplarily, the wake-up instruction may be transmitted to the first controller in the form of an interrupt signal, that is, the second controller may send an interrupt signal to the first controller to transmit the wake-up instruction to the first controller.
[0116] S806: The first controller responds to the wake-up instruction, causing the control device to switch from the sleep state to the wake-up state.
[0117] That is, when the first controller detects an interrupt signal, it responds to the interrupt signal and configures the clock signals of multiple audio acquisition units to the first controller, thereby shutting down the function of inputting the audio signals collected by the multiple audio acquisition units to the second controller, and inputting the audio signals collected by the multiple audio acquisition units to the first controller.
[0118] S807: The first controller sends a first clock signal to each of the multiple audio acquisition units.
[0119] S808 . The multiple audio collection units respectively send the mono audio signals collected according to the first clock signal to the first controller.
[0120] S809: The first controller converts the mono-channel audio signals collected by the multiple audio collection units into multi-channel audio signals.
[0121] S810: The target application program obtains an audio signal of a corresponding channel from a multi-channel audio signal.
[0122] Among them, steps S807 to S810 of this embodiment are respectively Figure 5 In the embodiment shown, steps S501 to S504 are similar, and can be referred to Figure 5 For the sake of brevity, the detailed description of the illustrated embodiments will not be repeated here.
[0123] In this embodiment, a mounting chip (i.e., a second controller) is provided to perform wake-up content detection. During the process of the mounting chip performing wake-up content detection, the control device is in a sleep state, the overall power consumption is relatively low, and the use time of the control device can be increased.
[0124] In addition, after the control device switches to the awake state, if the second controller and the first controller work simultaneously, the power consumption is relatively large. Therefore, the second controller may be powered off to further reduce the overall power consumption.
[0125] In addition, in this solution, the control device can realize audio acquisition through the ALSA standard audio architecture, and the software application for audio acquisition does not need to be customized, can be universal, and has high universality.
[0126] An embodiment of the present disclosure further provides a readable storage medium, comprising: computer program instructions; when the computer program instructions are executed by a processor of an electronic device, the audio processing method shown in any of the above method embodiments is implemented.
[0127] The embodiments of the present disclosure further provide a computer program product. When the computer program product is executed by a computer, the computer is enabled to implement the audio processing method shown in any of the above method embodiments.
[0128] For ease of explanation, the above description has been made in conjunction with specific embodiments. However, the above discussion of some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are intended to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.
Claims
1. A control device, characterized in that: include: A plurality of audio acquisition units and a first controller, wherein the plurality of audio acquisition units are electrically connected to the first controller; The first controller is configured to: when the control device is in an awake state, send a first clock signal to each of the multiple audio acquisition units, and obtain the mono audio signals respectively acquired by the multiple audio acquisition units according to the first clock signal; converting the mono-channel audio signals respectively collected by the multiple audio collection units into multi-channel audio signals, so that the target application installed in the control device obtains audio signals of corresponding channels from the multi-channel audio signals; The first controller is specifically configured to: combine the mono audio signals respectively collected by the multiple audio collection units and the echo signals respectively corresponding to the multiple audio collection units according to the correspondence between the multiple audio collection units and the channel positions, to obtain the multi-channel audio signal; When the target application includes multiple applications, each application can obtain the audio signal of the corresponding channel from the multi-channel audio signal according to its own needs; If the target application is a far-field voice service, the target application cyclically obtains a signal corresponding to each of the audio acquisition units in the multi-channel audio signal to perform wake-up word detection; If the target application is a near-field voice service, the target application performs channel separation on the multi-channel audio signal, extracts the audio signal of the target channel, obtains a voice command based on the audio signal of the target channel, and responds to the voice command.
2. The control device according to claim 1, characterized in that The control device further includes a second controller electrically connected to the plurality of audio acquisition units; The second controller is configured to: when the control device is in a dormant state, send a second clock signal to each of the plurality of audio acquisition units, and control the plurality of audio acquisition units to respectively acquire a mono audio signal according to the second clock signal; And detect whether the mono audio signals respectively collected by the multiple audio collection units include preset voice wake-up content.
3. The control device according to claim 2, characterized in that The second controller is further configured to: when the preset voice wake-up content is detected according to the mono audio signals respectively collected by the multiple audio collection units, send a wake-up instruction to the first controller to wake up the control device.
4. The control device according to claim 3, characterized in that The first controller is further configured to: According to the wake-up instruction, controlling the control device to switch from a sleep state to a wake-up state; After the control device switches to the awake state, the power supply of the second controller is cut off, and the first clock signal is sent to the multiple audio collection units respectively.
5. The control device according to claim 1 or 2, characterized in that: The multi-channel audio signal is an audio signal in a pulse code modulation (PCM) format.
6. An audio processing method, characterized in that: The control device according to any one of claims 1 to 5, wherein the control device comprises a plurality of audio acquisition units; and the method comprises: Sending a first clock signal to each of the plurality of audio acquisition units; Acquire mono audio signals collected by the multiple audio collection units respectively according to the first clock signal; Converting the mono audio signals respectively collected by the plurality of audio collection units into multi-channel audio signals, so that a target application installed on the control device obtains audio signals of corresponding channels from the multi-channel audio signals; The converting of the mono audio signals respectively collected by the plurality of audio collection units into a multi-channel audio signal comprises: According to the correspondence between the multiple audio collection units and the channel positions, the mono audio signals respectively collected by the multiple audio collection units and the collected signals respectively corresponding to the multiple audio collection units are merged to obtain the multi-channel audio signal; When the target application includes multiple applications, each application can obtain the audio signal of the corresponding channel from the multi-channel audio signal according to its own needs; If the target application is a far-field voice service, the target application cyclically obtains a signal corresponding to each of the audio acquisition units in the multi-channel audio signal to perform wake-up word detection; If the target application is a near-field voice service, the target application performs channel separation on the multi-channel audio signal, extracts the audio signal of the target channel, obtains a voice command based on the audio signal of the target channel, and responds to the voice command.
7. The method according to claim 6, characterized in that The control device further includes a second controller electrically connected to the plurality of audio acquisition units; The method further comprises: When the control device is in a dormant state, sending a second clock signal to each of the plurality of audio acquisition units through the second controller, and controlling the plurality of audio acquisition units to respectively acquire a mono audio signal according to the second clock signal; and detecting whether the mono audio signals respectively collected by the multiple audio collection units include preset voice wake-up content; When it is detected that the mono audio signals respectively collected by the multiple audio collection units include preset voice wake-up content, a wake-up instruction is sent to the first controller to wake up the control device.
Citation Information
Patent Citations
Multi-sound-channel mixed audio signal receiving and processing method
CN103985246A
Implementation method, device and system based on voice control network, and storage medium
CN107767867A