An audio acquisition method and apparatus

By setting priorities for voice service types in intelligent playback devices and rationally scheduling recording channels, the problems of resource waste and conflicts among multiple recording channels are solved, enabling multi-channel concurrent recording and efficient resource utilization.

CN115966203BActive Publication Date: 2025-12-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111169887.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-08
Publication Date
2025-12-19
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

In existing smart playback devices, the resources of multiple recording channels are not being used properly, resulting in resource waste and conflicts. In particular, in multi-recording channel environments, high-priority voice services cannot preempt occupied channels, causing a mismatch in recording resources.

Method used

By setting priorities for different voice service types and rationally scheduling multiple recording channels, recording requests for high-priority voice service types are allowed, preventing low-priority voice services from interfering with high-priority services, thus enabling multi-channel concurrent recording.

Benefits of technology

It improves the utilization rate of recording channel resources, ensures the recording accuracy of high-priority voice services, prevents low-priority services from interfering with high-priority services, and enables simultaneous recording of multiple voice services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966203B_ABST
    Figure CN115966203B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice processing, and provides an audio acquisition method and device. When a first recording request is received, a first voice service type corresponding to the first recording request is determined. When there is no authorized target voice service type or the priority of the target voice service type is lower than the priority of the first voice service type, recording is allowed, and voice data is acquired based on a recording channel corresponding to the first voice service type. Whether to acquire voice data of the first voice service type is determined by comparing the priority of the first voice service type and the priority of the target voice service type. In this way, the interference of a low-priority voice service type on a high-priority voice service type can be effectively prevented, and the accuracy of acquired voice data is improved. Furthermore, a corresponding recording channel is called based on the first voice service type, so that the rational use of channel resources is realized, and the utilization rate of the channel resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and discloses an audio acquisition method and device. BACKGROUND

[0002] With the development of artificial intelligence technology, most intelligent playing devices (for example, mobile phones, sound boxes, etc.) have voice recognition function and audio / video call function.

[0003] At present, most intelligent playing devices acquire audio data based on a single recording channel, that is, different voice service types share one recording channel. Obviously, this method will cause channel resource conflict.

[0004] In order to solve the above resource conflict problem, multiple recording channels are set in some new intelligent playing devices (for example, smart televisions). However, there is no audio data acquisition method that can reasonably use multiple recording channels at present, causing waste of channel resources. SUMMARY

[0005] The embodiments of the present application provide an audio acquisition method and device to improve the utilization rate of channel resources.

[0006] In a first aspect, the embodiments of the present application provide an audio acquisition method, which comprises:

[0007] In response to a first recording request, determining a first voice service type corresponding to the first recording request;

[0008] When it is determined that there is no authorized target voice service type at present or the priority of the first voice service type is higher than the priority of the currently authorized target voice service type, allowing the first recording request, and acquiring voice data of the first voice service type based on a recording channel corresponding to the first voice service type.

[0009] In a second aspect, the embodiments of the present application provide an intelligent playing device, which comprises:

[0010] A response module, configured to determine a first voice service type corresponding to a first recording request in response to the first recording request;

[0011] An acquisition module, configured to allow the first recording request when it is determined that there is no authorized target voice service type at present or the priority of the first voice service type is higher than the priority of the currently authorized target voice service type, and acquire voice data of the first voice service type based on a recording channel corresponding to the first voice service type.

[0012] Optionally, the intelligent playing device is installed with a first application, the first application comprises an audio and video call module, and the audio and video call module is used for cross-process communication with a second application; and the collection module is used for:

[0013] The second application is used to collect voice data of the audio and video call type based on a recording channel corresponding to the audio and video call type.

[0014] The audio and video call module is used to call a reading interface of the second application to obtain voice data of the audio and video call type collected by the second application.

[0015] Optionally, the audio and video call module comprises a first communication unit, and the second application comprises a second communication unit, and the collection module is specifically used for:

[0016] The first communication unit is used to send a data acquisition request to the second communication unit, so that the second communication unit initializes recording parameters based on the received data acquisition request, and collects voice data from a recording channel corresponding to the audio and video call type based on the initialized recording parameters.

[0017] Optionally, the first recording request is sent after the audio and video call module detects an audio and video call instruction.

[0018] Optionally, the intelligent playing device further comprises a sending module, which is used for:

[0019] The sending module is used to send a first control instruction to at least one first recording module, so that the at least one first recording module does not collect voice data by using a corresponding recording channel; wherein a priority of a voice service type corresponding to the at least one first recording module is lower than a priority of the first voice service type.

[0020] Optionally, the intelligent playing device further comprises a sending module, which is used for:

[0021] The sending module is used to receive a recording completion instruction corresponding to the first voice service type, and send a second control instruction to the at least one first recording module, so that the at least one first recording module restores a use state of the corresponding recording channel.

[0022] Optionally, the first application further comprises a near-field voice module, and the response module is specifically used for:

[0023] When it is determined that the first recording request is sent after the near-field voice module detects a voice key event, the response module determines that the first voice service type is a near-field voice service type.

[0024] Optionally, the first application further comprises a far-field voice module, the far-field voice module enters a listening state when starting up, and the response module is specifically configured to:

[0025] When it is determined that the first voice recording request is sent by the far-field voice module after the first wake-up word is listened to through a corresponding voice recording channel, it is determined that the first voice service type is a far-field voice service type.

[0026] Optionally, the voice service type comprises a near-field voice service type, a far-field voice service type and an audio-video call type, the priority of the near-field voice service type is higher than the priority of the far-field voice service type, and the priority of the far-field voice service type is higher than the priority of the audio-video call type.

[0027] Optionally, when the first voice service type is the audio-video call type and the voice data comprises a second wake-up word, the sending module is further configured to:

[0028] send a third control instruction to an audio-video call module corresponding to the first voice service type, so that the audio-video call module stops collecting voice data through a voice recording channel corresponding to the audio-video call type; and

[0029] send an audio recording permission instruction to a far-field voice module woken up based on the second wake-up word, so that the far-field voice module collects voice data through a voice recording channel corresponding to the far-field voice service type.

[0030] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement an audio collection method.

[0031] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and when the computer instructions are executed on a computer, the computer executes an audio collection method.

[0032] In the above embodiments of the present application, after receiving the first recording request, the voice service type corresponding to the first recording request is determined. When there is no authorized target voice service type at present, or there is an authorized target voice service type at present, but the priority of the first voice service type is higher than the priority of the target voice service type, recording is allowed, and voice data is collected based on the recording channel corresponding to the first voice service type. By comparing the priority of the first voice service type with the priority of the target voice service type, it is determined whether to collect voice data of the first voice service type. In this way, the interference of a voice service type with low priority on a voice service type with high priority can be effectively prevented, and the accuracy of the collected voice data is improved. Moreover, the corresponding recording channel is called based on the first voice service type, so that the channel resources are reasonably used, and the utilization rate of the channel resources is improved. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 A hardware structure diagram of the intelligent playing device provided by the embodiments of the present application is provided.

[0034] Figure 2 A software structure diagram of the intelligent playing device provided by the embodiments of the present application is provided.

[0035] Figure 3 A priority timing diagram of the voice service type provided by the embodiments of the present application is provided.

[0036] Figure 4 A flowchart of the multi-path concurrent recording method provided by the embodiments of the present application is provided.

[0037] Figure 5 A recording switching process schematic diagram provided by the embodiments of the present application is provided.

[0038] Figure 6 A flowchart of the scheduling method of each functional module in the intelligent playing device provided by the embodiments of the present application is provided.

[0039] Figure 7 A functional structure diagram of the intelligent playing device provided by the embodiments of the present application is provided.

[0040] Figure 8 A structure diagram of the electronic device provided by the embodiments of the present application is provided.

[0041] Figure 9 A structure diagram of the terminal device provided by the embodiments of the present application is provided. DETAILED DESCRIPTION

[0042] In order to better understand the technical solutions provided by the embodiments of the present application, the following will be described in detail in combination with the drawings of the specification and specific implementation manners.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] In order for those skilled in the art to better understand the technical solutions of the present application, the basic concepts involved in the present application are described below.

[0045] Near-field voice: refers to the voice input by the user through an external device (for example, a remote control, a wireless terminal). It needs to be triggered by a physical key. The key is pressed, the near-field voice starts recording, the key is released, and the near-field voice recording ends. Usually, the distance between the external device and the smart player is not more than 50 centimeters (cm), so the voice input in this way is called near-field voice.

[0046] Far-field voice: refers to the voice input by the user through the microphone array provided by the smart player. It needs to be triggered by a specific wake-up word. When the wake-up word is detected, the far-field voice starts recording. After the user finishes speaking and remains still for a set period of time, the far-field voice recording ends. Usually, the distance between the user and the microphone array is about 3 to 5 meters (m), so the voice input in this way is called far-field voice.

[0047] Audio and video call: a new function of a smart player, which needs to use the built-in or external camera of the smart player, and the camera has a sound collection function (for example, the camera has a microphone). After the user initiates a call or is called passively, the camera starts collecting images while the microphone collects voice. When the call ends, the collection of images and voice ends.

[0048] The design idea of the present application is described below.

[0049] At present, for a single recording channel smart player (for example, a mobile phone), a built-in microphone is mostly used to collect audio data. In such a smart player, each voice service type corresponds to an application program (Application, APP), and each APP shares a recording channel to receive audio data of the corresponding type. This way of sharing a recording channel by multiple APPs is called a single-channel recording mechanism.

[0050] The intelligent playing device using single-channel recording mechanism usually receives audio data from the outside based on the principle of "first occupation first use". Specifically, the APP that starts recording function first will occupy the recording channel to receive the corresponding type of audio data, at this time, other APPs cannot use the recording channel, only after the APP that starts recording function first ends recording and releases the recording channel, other APPs can continue to use the recording channel. Obviously, the principle of "first occupation first use" will cause resource conflicts among various APPs.

[0051] In some new intelligent playing devices (for example, smart TV), multiple recording channels are set, each recording channel corresponds to a voice service type, thereby solving the problem of resource conflicts caused by single recording mechanism.

[0052] With the continuous development of intelligent playing devices, the functions of intelligent playing devices are more and more, including but not limited to near-field voice recognition, far-field voice recognition and audio-video call function, the three functions respectively depend on different recording channels.

[0053] From the principle of "first occupation first use", different voice service types do not distinguish priority, and treat various voice services equally. When a high-priority voice service requests recording, since the recording channel has been occupied, the high-priority voice service cannot occupy the recording channel, and can only wait. Obviously, the principle of "first occupation first use" in single-channel recording mechanism is no longer applicable to intelligent playing devices with multiple recording channels. Since the channel resource scheduling strategy corresponding to the single-channel recording mechanism is not applicable to the multi-channel recording mechanism, there will be a problem of mismatch between software mechanism and hardware resources, which causes the hardware resources to be unable to be reasonably utilized, resulting in resource waste.

[0054] For the intelligent playing device with multiple recording channels, the embodiment of the present application provides an audio acquisition method and device to reasonably schedule multiple recording channels to realize multi-channel concurrent acquisition of audio data and improve the utilization rate of channel resources. Specifically, in the embodiment of the present application, the function modules of different voice service types are integrated in the same application, and priorities of different types of voice services are set in advance, wherein the priority of near-field voice service type is higher than that of far-field voice service type, and the priority of far-field voice service type is higher than that of audio-video call service type. When a recording request is received, the priority of the voice service type corresponding to the recording request is compared with the priority of the current voice service type, and when the former is higher than the latter, voice data is acquired based on the recording channel corresponding to the high-priority voice service type, thereby preventing the interference of low-priority voice service type to high-priority voice service type and improving the accuracy of recording. Moreover, the corresponding recording channel is called based on the voice service type corresponding to the recording request, thereby realizing reasonable utilization of channel resources and improving the utilization rate of channel resources.

[0055] By the audio acquisition method of the embodiment of the present application, the coexistence of near-field voice, far-field voice and audio-video call three-way recording can be realized, and when the audio-video call is performed, the voice data is acquired by the embodiment of the present application in a cross-process communication mode, thereby ensuring the simultaneous recording of far-field voice and audio-video call.

[0056] Taking the smart player device as a smart TV for example, Figure 1 The hardware structure diagram of the smart player device provided by the embodiment of the present application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the smart TV 10 is configured with an external remote controller 101, and the remote controller 101 is internally provided with a microphone. When the voice button of the remote controller 101 is pressed, the remote controller microphone starts recording the near-field voice, and when the voice button of the remote controller 101 is released, the near-field voice recording is ended. The remote controller 101 sends the recorded near-field voice to the smart TV 10 through a communication protocol, and the TV assistant on the smart TV 10 acquires the near-field voice data through a system interface and performs a series of operations such as voice recognition, semantic understanding and voice response on the obtained near-field voice data. In the embodiment of the present application, the type of the remote controller 101 is not limited, including but not limited to a Bluetooth remote controller, an infrared remote controller and a wireless remote controller.

[0057] It should be noted that, Figure 1 The voice button in the above embodiment is only an example, and in addition to being arranged on the remote controller for controlling the smart TV, the voice button can also be arranged on the smart TV itself, and the type of the voice button can be a press type or a touch type, and the embodiment of the present application does not make a restrictive requirement.

[0058] As shown in FIG. 2, Figure 1 The smart TV 10 is internally provided with a microphone array 102, and the recording function of the microphone array 102 is started with the start of the smart TV 10, that is, after the smart TV 10 is started, the microphone array 102 enters a listening state, and after the smart TV 10 is turned off, the microphone array 102 ends the listening state. After a specific wake-up word is listened to in the listening state, the microphone array 102 starts recording the far-field voice, and after the user finishes speaking and remains still for a period of time, the far-field voice recording is ended. The TV assistant on the smart TV 10 acquires the far-field voice data through a system interface and performs a series of operations such as voice recognition, semantic understanding and voice response on the obtained far-field voice data.

[0059] As shown in FIG. 3, Figure 1 The smart TV 10 is internally provided with or externally connected with a camera 103, and the camera 103 is provided with a microphone for call. When the audio-video call is connected, the camera 103 simultaneously acquires video images and audio data, and after the call is ended, the audio-video data recording is completed.

[0060] It should be noted that, Figure 1The illustrated intelligent playing device is only an example, and the intelligent playing device in the embodiments of the present application includes, but is not limited to, a vehicle-mounted terminal, a notebook computer, a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, and the like, which are terminals with audio and video playing functions.

[0061] Based on Figure 1 The illustrated intelligent playing device, Figure 2 An exemplary software structure diagram of the intelligent playing device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, Figure 2 As shown, the playing assistant of the intelligent playing device includes a near-field voice module, a far-field voice module, an audio and video call module, and a recording scheduling management module.

[0062] The near-field voice module corresponds to a “VOIC_RECOGNITION” recording channel, and is configured to provide a near-field voice service, including listening to a voice key, acquiring near-field voice data, and identifying the near-field voice data. Specifically, the near-field voice module detects a voice key event of a remote controller, and after detecting the voice key event, receives near-field voice data recorded by a microphone of the remote controller through the recording channel “VOIC_RECOGNITION”. The priority corresponding to the near-field voice service is set to high.

[0063] The far-field voice module corresponds to a “HOTWORD” recording channel, and is configured to provide a far-field voice service, including listening to a wake-up word, acquiring far-field voice data, and identifying the far-field voice data. The far-field voice module enters a listening state when the intelligent playing device is powered on, and continuously listens to the wake-up word for waking up the far-field voice service. After listening to the wake-up word, the far-field voice module receives far-field voice data recorded by a built-in microphone array through the recording channel “HOTWORD”. The priority corresponding to the far-field voice service is set to intermediate.

[0064] The audio and video call module corresponds to a “MIC” recording channel, and is configured to provide an audio and video call service, including acquiring voice data and image data. The priority corresponding to the audio and video call service is set to low.

[0065] The recording scheduling management module is configured to send corresponding scheduling instructions to the near-field voice module, the far-field voice module, and the audio and video call module, so as to reasonably schedule multiple recording channels to realize multi-path concurrent recording.

[0066] The configuration information of the function modules corresponding to various voice service types in the embodiments of the present application is shown in Table 1.

[0067] Table 1

[0068] Module Voice service type Priority Recording channel Near-field voice module Near-field voice service High VOIC_RECOGNITION Far-field voice module Far-field voice service Medium HOTWORD Audio-video call module Audio-video call service Low MIC

[0069] Figure 2 The running environment of the intelligent playing device shown is an Android operating system. Limited by the Android system, the playing assistant APP (denoted as a first application) only starts one recording process, i.e., only one recording channel can record at a time. Thus, the continuous monitoring of the far-field voice module and the audio-video call cannot coexist. In order to solve the problem of simultaneous recording, the application sets an independent call APP (denoted as a second application) for the audio-video call.

[0070] As shown in Figure 2 As shown in the audio-video call module and the call APP, an AKAudioRecord unit is designed to realize the cross-process communication between the audio-video call module and the call APP, and to receive the audio-video data recorded by the camera, so that the audio-video call module and the far-field voice module can simultaneously record through two different processes.

[0071] The AKAudioRecord unit is divided into a server (Server) end and a client (Client) end. The Server end (denoted as a second communication unit) of the AKAudioRecord unit integrated in the call APP is used to initialize the recording parameters (including the number of recording channels, the audio-video sampling frequency, the audio-video sampling depth, etc.), control the camera to start or stop recording audio-video data, and receive the audio-video data recorded by the camera through the "MIC" recording channel. The Client end (denoted as a first communication unit) of the AKAudioRecord unit integrated in the audio-video call module in the playing assistant APP is used to obtain the audio-video data recorded by the camera microphone through the AudioRecord interface, thereby realizing the audio-video call function.

[0072] It should be noted that Figure 2 The cross-process implementation manner in the above is only an example. In addition to using the AudioRecord recording interface in the Android system, the Alsa recording interface, the Binder AIDL interface, the pipe, the shared memory, etc. can also be used.

[0073] The application sets different priorities for various types of voice services. Taking the voice services including the near-field voice, the far-field voice, and the audio-video call as examples, the priorities of the three types of voice services are as follows: the priority of the near-field voice service type is higher than that of the far-field voice service type, and the priority of the far-field voice service type is higher than that of the audio-video call type. Based on the set priorities, the three channel resources are reasonably scheduled to realize multi-channel concurrent recording.

[0074] Figure 3 The embodiment of the application is a schematic diagram of the embodiment of the application for the priorities of the three types of voice services on the time axis; as shown inFigure 3 As shown, the far-field voice service enters the listening state at startup, the near-field voice service and the audio-video call service enter the recording state after reaching the trigger condition respectively. The priority of the near-field voice is higher than that of the far-field voice, and the priority of the far-field voice is higher than that of the audio-video call, which ensures that the near-field voice recording of the remote controller can interrupt the recording of the far-field voice, and the recording of the near-field voice or the far-field voice can control the audio-video call function.

[0075] As shown in (a) in FIG. 1, Figure 3 As shown, when the voice key event is detected, the near-field voice starts to be recorded, at this time, the far-field voice ends the listening state, and the camera pauses recording the audio-video data in the audio-video call to avoid the far-field voice being awakened or the voice in the audio-video call being recorded into the near-field voice, which interferes with the real intention of the near-field voice recorded by the user. When the voice key event ends, the recording of the near-field voice ends, the far-field voice resumes the listening state, and the camera continues to record the audio-video data in the audio-video call.

[0076] As shown in (b) in FIG. 1, Figure 3 As shown, when the microphone array receives a specific wake-up word, the far-field voice starts to be recorded, at this time, the camera pauses recording the audio data in the audio-video call to avoid the voice in the audio-video call being recorded into the far-field voice, which interferes with the real intention of the far-field voice. After the user stops speaking and remains still for a period of time, the recording of the far-field voice ends, and the camera continues to record the audio-video data in the audio-video call.

[0077] As shown in (c) in FIG. 1, Figure 3 As shown, when the audio-video call starts, the far-field voice can maintain the listening state, so that the audio-video call and the far-field voice use the two channels at the same time, realizing the coexistence of the far-field voice and the audio-video call.

[0078] Based on the device structure shown in FIG. 1, Figures 1-2 The multi-path concurrent recording method involved in the embodiments of the present application is described in detail. Referring to FIG. 2, Figure 4 The method is executed by the intelligent playing device and mainly includes the following steps:

[0079] S401: In response to a first recording request, determining a first voice service type corresponding to the first recording request.

[0080] Taking the intelligent playing device that can provide three types of voice services, i.e., near-field voice service, far-field voice service and audio-video call service, as an example, the first recording request can be a request sent by a near-field voice module providing the near-field voice service, a request sent by a far-field voice module providing the far-field voice service, or a request sent by an audio-video call module providing the audio-video call service. The three modules are integrated in the first application at the same time.

[0081] Taking the first recording request sent by the near-field voice module as an example, when S401 is executed, the first recording request is sent after the near-field voice module detects a voice key event. The first recording request carries an identifier of a near-field voice service corresponding to the near-field voice module. Based on the identifier, it is determined that the first voice service type corresponding to the first recording request is a near-field voice service type.

[0082] The voice key event is not limited in the embodiments of the present application. It can be triggered by an external device (for example, a remote controller) that controls the intelligent playback room device, or by a key of the intelligent playback device itself. The triggering mode can be a key mode or a touch mode.

[0083] Taking the first recording request sent by the far-field voice module as an example, when S401 is executed, the first recording request is sent after the far-field voice module listens to a first wake-up word through a "HOTWORD" recording channel. The first recording request carries an identifier of a far-field voice service corresponding to the far-field voice module. Based on the identifier, it is determined that the first voice service type corresponding to the first recording request is a far-field voice service type.

[0084] The first wake-up word can be set according to actual needs. It can be one or multiple. The listening algorithm of the wake-up word is not limited in the embodiments of the present application. It includes but is not limited to a voice activity detection (VAD) algorithm and a key word search (KWS) algorithm.

[0085] Taking the first recording request sent by the audio-video call module as an example, when S401 is executed, the first recording request is sent after the audio-video call module detects an audio-video call instruction. The first recording request carries an identifier of an audio-video call service corresponding to the audio-video call module. Based on the identifier, it is determined that the first voice service type corresponding to the first recording request is an audio-video call service type.

[0086] The triggering mode of the audio-video call instruction is not limited in the embodiments of the present application. It can be triggered when an audio-video call is actively dialed, or it can be triggered when an audio-video call is passively received.

[0087] S402: Determine whether there is an authorized target voice service type at present. If there is, execute S403. If there is not, execute S404.

[0088] Since the intelligent playback device can provide multiple types of voice services in the embodiments of the present application, after receiving the first recording request, it is necessary to determine whether there is an authorized target voice service type at present. If there is, the priority of the target voice service type is obtained. If there is not, recording is directly performed.

[0089] S403: Determine whether the priority of the first voice service type is higher than the priority of the target voice service type. If yes, execute S404, otherwise, execute S405.

[0090] In the execution of S403, the priority of the first voice service type is obtained, and the priority of the first voice service type is compared with the priority of the target voice service type. If the former is higher than the latter, it indicates that the importance level of the first voice service type is higher, and the target voice service should be interrupted. If the former is lower than the latter, it indicates that the importance level of the target voice service is higher, and the first recording request should be rejected to prevent the first voice service from interfering with the target voice service.

[0091] For example, taking the first voice service type as a near-field voice service type and the target voice service type as a far-field voice service type as an example, since the priority of the near-field voice service type is higher than the priority of the far-field voice service type, S404 is executed.

[0092] For another example, taking the first voice service type as an audio-video call type and the target voice service type as a far-field voice service type as an example, since the priority of the audio-video call type is lower than the priority of the far-field voice service type, the first recording request is rejected.

[0093] S404: Allow the first recording request, and collect voice data of the first voice service type based on a recording channel corresponding to the first voice service type.

[0094] When the first voice service type is a near-field voice service type, S404 is executed, and the recording scheduling management module in the intelligent playing device sends a first recording instruction to the near-field voice module to inform the near-field voice module to allow the entry of near-field voice. After receiving the first recording instruction, the near-field voice module collects near-field voice data through a “VOIC_RECOGNITION” recording channel corresponding to the near-field voice service type.

[0095] When the first voice service type is a far-field voice service type, S404 is executed, and the recording scheduling management module in the intelligent playing device sends a first recording instruction to the far-field voice module to inform the far-field voice module to allow the entry of far-field voice. After receiving the first recording instruction, the far-field voice module collects far-field voice data through a “HOTWORD” recording channel corresponding to the far-field voice service type.

[0096] When the first voice service type is an audio-video call type, since the audio-video call module includes a first communication unit and the second unit includes a second communication unit, as Figure 2Therefore, even if the far-field voice module enters the listening state to continuously receive sound after the smart playing device is powered on, the audio and video call module and the second application can perform cross-process communication through the first communication unit and the second communication unit to realize coexistence of the far-field voice and the audio and video call.

[0097] When S404 is performed, specifically, the first application and the second application in the smart playing device perform communication, through the second application, voice data of the audio and video call type is collected based on the audio and video call type corresponding recording channel, and a read interface (for example, AudioRecord interface) of the second application is called through the audio and video call module to obtain the voice data of the audio and video call type collected by the second application.

[0098] Specifically, the first application sends a data acquisition request to the second communication unit in the second application through the first communication unit in the audio and video call module, after the second communication unit receives the data acquisition request, initializes recording parameters, including the number of recording channels, audio and video sampling frequency, audio and video sampling depth, and collects voice data from the "MIC" recording channel corresponding to the audio and video call type based on the initialized recording parameters.

[0099] It should be noted that while collecting voice data from the "MIC" recording channel corresponding to the audio and video call type, image data collected by the camera can also be received through the "MIC" recording channel.

[0100] S405: reject the first recording request.

[0101] When S405 is performed, since the priority of the first voice service type is lower than the priority of the target service type, in order to prevent the first voice service from interfering with the target voice service, the first recording request should be rejected.

[0102] The embodiments of the present application set different priorities for different types of voice services, therefore, in an optional implementation, after receiving the recording request, a control instruction can also be sent to the module corresponding to the voice service type with low priority according to the priority order of the voice service type, to ensure that during the recording process, the interference of the voice service with low priority on the voice service with high priority is reduced.

[0103] Specifically, when there is at least one voice service type with a priority lower than that of the first voice service type, a first control instruction is sent to the first recording module corresponding to the at least one voice service type with low priority, so that the at least one first recording module does not use the corresponding recording channel to collect voice data.

[0104] Taking the first voice service type as the near-field voice service type as an example, the voice service types with a priority lower than that of the near-field voice service type include the far-field voice service type and the audio-video call type. After receiving the first recording request sent by the near-field voice module, the recording scheduling management module sends a first control instruction to the far-field voice module and the audio-video call module respectively, so that the far-field voice module does not collect far-field voice data through the "HOTWORD" recording channel, and the audio-video call module does not collect audio-video voice data through the "MIC" recording channel.

[0105] Taking the first voice service type as the far-field voice service type as an example, the voice service types with a priority lower than that of the far-field voice service type include the audio-video call type. After receiving the first recording request sent by the far-field voice module, the recording scheduling management module sends a first control instruction to the audio-video call module, so that the audio-video call module does not collect audio-video voice data through the "MIC" recording channel.

[0106] It should be noted that when the first voice service type is the audio-video call type, there is no voice service type with a priority lower than that of the first voice service type, and therefore the recording scheduling management module in the intelligent playing device does not send a first control instruction.

[0107] In the embodiments of the present application, after the voice data of the first voice service type is collected, the following operations can also be performed:

[0108] A recording completion instruction corresponding to the first voice service type is received, and a second control instruction is sent to at least one first recording module to make the at least one first recording module resume the use state of the corresponding recording channel.

[0109] Taking the first voice service type as the near-field voice service type as an example, when the user releases the voice key of the remote controller, the near-field voice data collection is completed, triggering the near-field voice module to send a recording completion instruction to the recording scheduling management module in the intelligent playing device. After receiving the recording completion instruction, the recording scheduling management module sends a second control instruction to the far-field voice module and the audio-video module respectively. After receiving the second control instruction, the far-field voice module continues to listen to the first wake-up word through the "HOTWORD" recording channel, i.e., resumes to the listening state. After receiving the second control instruction, if the audio-video call is still continuing, the audio-video call module resumes the use state of the "MIC" recording channel, collects audio-video voice data through the "MIC" recording channel, if the audio-video call has ended, the "MIC" recording channel resumes to the idle state and no longer collects audio-video voice data through the "MIC" recording channel.

[0110] Taking the first voice service type as the far-field voice service type as an example, when the user stops speaking for a period of time, the far-field voice data acquisition ends, triggering the far-field voice module to send a recording completion instruction to the recording scheduling management module in the intelligent playback device. After receiving the recording completion instruction, the recording scheduling management module sends a second control instruction to the audio and video call module. If the audio and video call continues after the far-field voice ends, the audio and video call module continues to acquire audio, video and voice data through the "MIC" recording channel. If the audio and video call has ended after the far-field voice ends, the "MIC" recording channel returns to the idle state.

[0111] It should be noted that when the first voice service type is an audio / video call, since there is no voice service type with lower priority, the recording scheduling management module in the intelligent playback device does not send the second control command.

[0112] In some embodiments of this application, cross-process communication enables audio / video calls and far-field voice to coexist. During an audio / video call, the far-field voice module still listens for the wake word through the "HOTWORD" recording channel. Thus, when the first voice service type is an audio / video call and the collected voice data contains the second wake word, the far-field voice module is woken up because the far-field voice service type has a higher priority than the audio / video call type. The recording scheduling and management module in the smart device needs to switch the audio / video call to far-field voice.

[0113] In specific implementation, the recording scheduling and management module in the intelligent playback device sends a third control command to the audio and video call module corresponding to the first voice service type. After receiving the third control command, the audio and video call module stops using the recording channel corresponding to the audio and video call type to collect voice data, and sends a recording permission command to the far-field voice module that is woken up based on the second wake-up word. After receiving the recording permission command, the far-field voice module collects voice data through the recording channel corresponding to the far-field voice service type.

[0114] For example, such as Figure 5 As shown, during the audio / video call, a second wake-up word appears in the voice data, which can wake up the far-field voice module. The recording management and scheduling module in the smart playback device sends a third control command to the audio / video call application. After receiving the third control command, the audio / video call application stops receiving audio / video data through the recording channel "MIC". Furthermore, the recording scheduling and management module in the smart playback device sends a recording permission command to the far-field voice module. After receiving the recording permission command, the far-field voice module collects far-field voice data through the "HOTWORD" recording channel until the far-field voice ends. The audio / video module then continues to collect voice data from the audio / video through "MIC".

[0115] In the above embodiments of the present application, the intelligent playing device is configured with a remote controller, a microphone array and a camera, and the remote controller and the camera are provided with microphones. Based on these hardware, three types of voice services, i.e., near-field voice, far-field voice and audio-video call, are provided through a playing assistant APP. The priority of the near-field voice service type is higher than that of the far-field voice service type, and the priority of the far-field voice service type is higher than that of the audio-video call type. After the intelligent playing device receives a recording request, the recording dispatch management module sends instructions to the media applications corresponding to each voice service type in the order of the priority of each voice service type, and through reasonable calling of the three recording channels, the three types of voice services are realized to coexist simultaneously, and the interference of the voice service of low priority on the voice service of high priority is prevented, and the recording channel called at that time is ensured to be consistent with the voice input intention of the user at that time. Figure 1

[0116] Meanwhile, a call APP is designed for the audio-video call to send the audio-video data recorded by the microphone of the camera to the playing assistant APP in the intelligent playing device through cross-process communication, thereby solving the limitation of the Android system and realizing simultaneous audio-video call and far-field voice.

[0117] To clearly describe the multi-channel audio data acquisition method provided by the embodiments of the present application, Figure 6 a schematic diagram of the dispatch of the recording dispatch management module to each voice service module is shown. As shown in the figure, Figure 6 S601-S603 are the process of starting up the intelligent playing device, S604-S608 are the process of acquiring near-field voice data, S609-S613 are the process of acquiring far-field voice data, and S614-S616 are the process of acquiring audio-video data in the audio-video call. The specific contents are as follows:

[0118] S601: After the intelligent playing device is started up, the recording dispatch management module, the near-field voice module, the far-field voice module and the audio-video call module receive the start-up instruction respectively.

[0119] In an optional embodiment, the user starts up the intelligent playing device through the "on-off" key of the remote controller or the "on-off" key built in the intelligent playing device.

[0120] S602: After receiving the start-up instruction, the far-field voice module enters a listening state and continuously listens for the wake-up word in the user voice.

[0121] In S602, after receiving the start-up instruction, the far-field voice module enters a listening state and continuously collects audio through the "HOTWORD" recording channel to listen for the wake-up word

[0122] ​S603: The near-field voice module, the far-field voice module, and the audio / video call module respectively send a scheduling strategy to the recording scheduling management module, and the scheduling strategy contains a voice service type provided by the corresponding module.

[0123] After the intelligent playing device is started, the near-field voice module, the far-field voice module, and the audio / video call module respectively send a scheduling strategy to the recording scheduling management module, and the scheduling strategy contains a voice service type corresponding to each voice service module. Each voice service module provides one type of voice service, each voice service module corresponds to one recording channel, and different priorities are respectively set for each voice service type. For details, refer to the foregoing embodiments, which are not repeated here.

[0124] S604: The near-field voice module sends a recording request to the recording scheduling management module.

[0125] In an optional implementation, when S604 is performed, the user presses the voice key of the remote controller to trigger the near-field voice module to send a recording request to the recording scheduling management module.

[0126] S605: The recording scheduling management module sends a first control instruction to the far-field voice module and the audio / video call module based on the received recording request, so that the far-field voice module and the audio / video call module stop recording.

[0127] In S605, since the priority of the near-field voice service type is higher than the priorities of the far-field voice service type and the audio / video call type, after receiving the recording request of the near-field voice module, the first control instruction is sent to the far-field voice module and the audio / video call module, so that the far-field voice module stops collecting far-field voice data using the “HOTWORD” recording channel, and the audio / video call module stops collecting audio / video data using the “MIC” recording channel, thereby preventing the far-field voice and the audio / video call from interfering with the near-field voice, and improving the accuracy of the near-field voice input.

[0128] S606: The recording scheduling management module sends a recording instruction to the near-field voice module based on the received recording request, so that the near-field voice module collects near-field voice data through the corresponding recording channel.

[0129] In S606, after receiving the recording instruction sent by the recording scheduling management module, the near-field voice module collects near-field voice data through the “VOIC_RECOGNITION” recording channel.

[0130] S607: The near-field voice module sends a recording completion instruction to the recording scheduling management module.

[0131] In some embodiments, the execution of the recording completion instruction in S607 can be triggered by the user releasing the voice button of the remote controller, or triggered after a preset period of time of inactivity.

[0132] S608: The recording scheduling management module sends a second control instruction to the far-field voice module and the audio / video call module respectively based on the received recording completion instruction, so that the far-field voice module and the audio / video call module return to the original state.

[0133] In S608, after the near-field voice recording is completed, the second control instruction is sent to the far-field voice module and the audio / video call module respectively, the far-field voice module resumes the listening state of the "HOTWORD" recording channel based on the second control instruction, and the audio / video call module resumes the use state of the "MIC" recording channel based on the second control instruction.

[0134] S609: The far-field voice module sends a recording request to the recording scheduling management module.

[0135] In an optional implementation, when S609 is executed, the user triggers the far-field voice module to send a recording request to the recording scheduling management module through the wake-up word.

[0136] S610: The recording scheduling management module sends a first control instruction to the audio / video call module based on the received recording request, so that the audio / video call module stops recording.

[0137] In S610, since the priority of the far-field voice service type is higher than the priority of the audio / video call type, after receiving the recording request sent by the far-field voice module, the first control instruction is sent to the audio / video call module, so that the audio / video call module stops using the "MIC" recording channel to collect audio / video data, thereby preventing the audio / video call from interfering with the far-field voice, and improving the accuracy of far-field voice input.

[0138] S611: The recording scheduling management module sends a recording instruction to the far-field voice module based on the received recording request, so that the far-field voice module collects far-field voice data through the corresponding recording channel.

[0139] In S611, after the far-field voice module receives the recording instruction sent by the recording scheduling management module, it collects far-field voice data through the "HOTWORD" channel.

[0140] S612: The far-field voice module sends a recording completion instruction to the recording scheduling management module.

[0141] In an optional implementation, when S612 is executed, the user stops speaking and inactivates for a preset period of time, and then the far-field voice ends, triggering the far-field voice module to send a recording completion instruction to the recording scheduling management module.

[0142] S613: The recording scheduling management module sends a second control instruction to the audio-video call module based on the recording completion instruction, so that the audio-video call module restores the original state.

[0143] In S613, after the far-field voice recording is completed, a second control instruction is sent to the audio-video call module, and the audio-video call module restores the use state of the "MIC" channel based on the second control instruction.

[0144] S614: The audio-video call module sends a recording request to the recording scheduling management module.

[0145] In an optional embodiment, when S614 is executed, after the user makes or receives an audio-video call, the audio-video call module sends a recording request to the recording scheduling management module.

[0146] S615: The recording scheduling management module sends a recording instruction to the audio-video call module based on the received recording scheduling request, so that the audio-video call module collects audio-video data through the corresponding recording channel.

[0147] In an optional embodiment, when S615 is executed, after the audio-video call module receives the recording instruction, the AKAudioRecord interface is called to communicate with the call APP across processes, and the audio-video data collected by the call APP through the "MIC" recording channel is received. Since the cross-process communication mode is adopted between the audio-video call module and the call APP, the listening process of the far-field voice module is independent, so that the audio-video call and the far-field voice coexist.

[0148] S616: The audio-video call module sends a recording completion instruction to the recording scheduling management module.

[0149] In some embodiments, when S616 is executed, after the user hangs up the audio-video call, the audio-video call module sends a recording completion instruction to the recording scheduling management module.

[0150] It should be noted that, Figure 6 The scheduling process in the above is not a strict execution sequence, and in an optional manner, only one scheduling mode or any combination can be performed, for example, only S601-S608 are executed.

[0151] Based on the same technical concept, an intelligent playing device is provided. Referring to Figure 7 , the intelligent playing device comprises:

[0152] The response module 701 is configured to determine a first voice service type corresponding to the first recording request in response to the first recording request.

[0153] The collection module 702 is configured to determine that there is no authorized target voice service type currently, or determine that the priority of the first voice service type is higher than the priority of the currently authorized target voice service type, allow the first recording request, and collect voice data of the first voice service type based on a recording channel corresponding to the first voice service type.

[0154] Optionally, the intelligent playing device is installed with a first application, and the first application includes an audio and video call module configured to perform cross-process communication with a second application.

[0155] The second application is configured to collect voice data of the audio and video call type based on a recording channel corresponding to the audio and video call type.

[0156] The audio and video call module is configured to call a reading interface of the second application to obtain the voice data of the audio and video call type collected by the second application.

[0157] Optionally, the audio and video call module includes a first communication unit, and the second application includes a second communication unit.

[0158] The first communication unit is configured to send a data acquisition request to the second communication unit, so that the second communication unit initializes recording parameters based on the received data acquisition request, and collects voice data from a recording channel corresponding to the audio and video call type based on the initialized recording parameters.

[0159] Optionally, the first recording request is sent by the audio and video call module after detecting an audio and video call instruction.

[0160] Optionally, the intelligent playing device further includes a sending module 703 configured to:

[0161] The sending module 703 is configured to send a first control instruction to at least one first recording module, so that the at least one first recording module does not collect voice data using a corresponding recording channel.

[0162] Optionally, the intelligent playing device further includes a receiving module 704 configured to:

[0163] The receiving module 704 is configured to receive a recording completion instruction corresponding to the first voice service type, and send a second control instruction to the at least one first recording module, so that the at least one first recording module restores a use state of the corresponding recording channel.

[0164] Optionally, the first application further comprises a near-field voice module, and the response module 701 is specifically configured to:

[0165] When it is determined that the first voice recording request is sent after the near-field voice module detects a voice key event, it is determined that the first voice service type is a near-field voice service type.

[0166] Optionally, the first application further comprises a far-field voice module, and the far-field voice module enters a listening state when starting up, and the response module 701 is specifically configured to:

[0167] When it is determined that the first voice recording request is sent after the far-field voice module listens to a first wake-up word through a corresponding voice recording channel, it is determined that the first voice service type is a far-field voice service type.

[0168] Optionally, the voice service type comprises a near-field voice service type, a far-field voice service type, and an audio-video call type, the priority of the near-field voice service type is higher than the priority of the far-field voice service type, and the priority of the far-field voice service type is higher than the priority of the audio-video call type.

[0169] Optionally, when the first voice service type is an audio-video call type and the voice data comprises a second wake-up word, the sending module 703 is further configured to:

[0170] send a third control instruction to an audio-video call module corresponding to the first voice service type, so that the audio-video call module stops collecting voice data using a voice recording channel corresponding to the audio-video call type; and

[0171] send an audio recording permission instruction to a far-field voice module woken up based on the second wake-up word, so that the far-field voice module collects voice data through a voice recording channel corresponding to the far-field voice service type.

[0172] As an embodiment, Figure 7 The modules in the device can be used for the audio collection method provided by the intelligent playing device of the embodiments of the present application, and can achieve the same technical effects, which will not be described here.

[0173] The device as an example of a hardware entity is an electronic device as shown in Figure 8 The electronic device comprises a processor 801, a storage medium 802, and at least one external communication interface 803; the processor 801, the storage medium 802, and the external communication interface 803 are all connected through a bus 804.

[0174] The storage medium 802 stores a computer program.

[0175] The processor 801 implements the audio acquisition method discussed above when executing the computer program.

[0176] Figure 8 The processor 801 is taken as an example, but the number of processors 801 is not limited in practice.

[0177] The storage medium 802 can be a volatile memory, such as a random-access memory (RAM), or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The storage medium 802 can be a combination of the above storage media.

[0178] Based on the same inventive concept, the present application provides a terminal device, which is described below.

[0179] Please refer to Figure 9 The terminal device includes a display unit 940, a processor 980, and a storage 920. The display unit 940 includes a display panel 941, and is configured to receive user input and provide various operation interfaces and display pages.

[0180] Optionally, the display panel 941 can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).

[0181] The processor 980 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 980 reads a media application and runs the media application on a playback assistant APP of the terminal device, and displays the interface of the application on the display unit 940. The processor 980 can include one or more general-purpose processors, and can also include one or more digital signal processors (DSPs) for performing related operations to implement the technical solutions provided in the present application.

[0182] The memory 920 generally includes an internal memory and an external memory. The internal memory can be a random access memory (RAM), a read only memory (ROM), a cache memory (CACHE), etc. The external memory can be a hard disk, an optical disk, a USB disk, a floppy disk, a magnetic tape, etc. The memory 920 is configured to store computer programs and other data. The computer programs include application programs corresponding to the client, etc. The other data can include data generated after the operating system or the application programs are run, including system data (for example, configuration parameters of the operating system) and user data. In the embodiments of the present application, program instructions are stored in the memory 920, and the processor 980 executes the program instructions in the memory 920 to implement any one of the audio acquisition methods discussed in the foregoing figures.

[0183] In addition, the terminal device can further include a display unit 940 configured to receive inputted digital information, word information or contact touch operation or non-contact gesture, and generate signal input related to user settings and function control of the terminal device, etc. Specifically, in the embodiments of the present application, the display unit 940 can include a display panel 941. The display panel 941, for example, a touch screen, can collect touch operations (such as a user using a finger, a stylus or any suitable object or accessory on or near the display panel 941) of the user on or near the display panel 941, and drive the corresponding connection device according to the pre-set program. Optionally, the display panel 941 can include two parts of a touch detection device and a touch controller. The touch detection device detects the touch position of the player and detects the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch coordinates, and sends it to the processor 980, and can also receive commands from the processor 980 and execute them.

[0184] The display panel 941 can be implemented in various types such as resistive, capacitive, infrared and surface acoustic wave, etc. In addition to the display unit 940, the terminal device can further include an input unit 930. The input unit 930 can include but is not limited to an image input device 931 and other input devices 932. The other input devices 932 can include but are not limited to one or more of a physical keyboard, a function key (such as a volume control button, an on-off button, etc.), a trackball, a mouse, a joystick, etc.

[0185] In addition to the above, the terminal device can further include a power supply 990 for supplying power to other modules, an audio circuit 960, a near field communication module 970 and an RF circuit 910. The terminal device can further include one or more sensors 950, such as an acceleration sensor, a light sensor, a pressure sensor, a camera, etc. The audio circuit 960 specifically includes a speaker 961 and a microphone 962, etc. For example, the terminal device can collect the user's voice through the microphone 962, and perform corresponding operations, etc.

[0186] As an embodiment, the number of the processor 980 can be one or more, and the processor 980 and the memory 920 can be coupled or relatively independent.

[0187] As an embodiment, Figure 9 The processor 980 in the foregoing embodiment can be used to implement the functions of various modules in the foregoing embodiment. Figure 7 The processor 980 in the foregoing embodiment can be used to implement the functions of various modules in the foregoing embodiment.

[0188] As an embodiment, Figure 9 The processor 980 in the foregoing embodiment can be used to implement the functions of various modules in the foregoing embodiment.

[0189] Those skilled in the art can understand that all or part of the steps of the foregoing method embodiments can be completed by a program instruction related hardware, and the foregoing computer program can be stored in a computer readable storage medium, and when the computer program is executed, the steps of the foregoing method embodiments are executed; and the foregoing storage medium includes: a mobile storage device, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0190] Alternatively, the integrated unit of the present application can be stored in a computer readable storage medium if it is realized in the form of a software function module and sold or used as an independent product. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the embodiments of the present application. The foregoing storage medium includes: a mobile storage device, a ROM, a RAM, a magnetic disk or an optical disk, and various media that can store program codes.

[0191] Based on the same technical concept, the embodiments of the present application also provide a computer readable storage medium, which stores computer instructions, and when the computer instructions are executed on a computer, the computer executes the audio acquisition method as discussed above.

[0192] Those skilled in the art will appreciate that embodiments of the present application can be devised for a variety of other systems which are currently developed or later developed. Therefore, the present application is intended to cover all such modifications and variations of this application that are within the scope of the appended claims and their equivalents. It is intended that each element of claim 1 and 2 is independent of one another. No element of claim 1 and 2, or any other claim, is implied to depend on any other element or limitation of claim 1 and 2 or any other claim except where expressly recited in that claim.

[0193] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to this application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0194] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0195] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.

[0196] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. An audio acquisition method, characterized in that, The method includes: In response to the first recording request, determine the first voice service type corresponding to the first recording request; When it is determined that there is no authorized target voice service type, or when it is determined that the priority of the first voice service type is higher than the priority of the currently authorized target voice service type, the first recording request is allowed, and voice data of the first voice service type is collected based on the recording channel corresponding to the first voice service type. When there is at least one voice service type with a lower priority than the first voice service type, a first control instruction is sent to the first recording module corresponding to the at least one voice service type, and after receiving the recording completion instruction corresponding to the first voice service type, a second control instruction is sent to the at least one first recording module; wherein, the first control instruction is used to instruct the at least one first recording module not to use the corresponding recording channel to collect voice data, and the second control instruction is used to instruct the at least one first recording module to restore the usage status of the corresponding recording channel.

2. The method as described in claim 1, characterized in that, It is applied to a first application, which includes an audio and video call module for cross-process communication with a second application; When the first voice service type corresponding to the first recording request is an audio / video call type, the step of collecting voice data of the first voice service type based on the recording channel corresponding to the first voice service type includes: Through the second application, voice data of the audio and video call type is collected based on the recording channel corresponding to the audio and video call type; The audio and video call module calls the reading interface of the second application to obtain the voice data of the audio and video call type collected by the second application.

3. The method as described in claim 2, characterized in that, The audio and video call module includes a first communication unit, and the second application includes a second communication unit; The step of collecting voice data for the audio / video call type through the second application, based on the recording channel corresponding to the audio / video call type, includes: The first communication unit sends a data acquisition request to the second communication unit, so that the second communication unit initializes the recording parameters based on the received data acquisition request, and acquires voice data from the recording channel corresponding to the audio / video call type based on the initialized recording parameters.

4. The method as described in claim 2 or 3, characterized in that, The first recording request is sent by the audio / video call module after it detects an audio / video call instruction.

5. The method according to any one of claims 1-3, characterized in that, It is applied to a first application, which includes a near-field speech module; Determining the first voice service type corresponding to the first recording request includes: When it is determined that the first recording request was sent after the near-field voice module detected a voice button event, the first voice service type is determined to be a near-field voice service type.

6. The method according to any one of claims 1-3, characterized in that, It is applied to a first application, which includes a far-field voice module, and the far-field voice module enters a listening state when powered on; Determining the first voice service type corresponding to the first recording request includes: When it is determined that the first recording request was sent by the far-field voice module after listening to the first wake-up word through the corresponding recording channel, the first voice service type is determined to be a far-field voice service type.

7. The method as described in claim 1, characterized in that, Voice service types include near-field voice service type, far-field voice service type, and audio / video call type. The near-field voice service type has a higher priority than the far-field voice service type, and the far-field voice service type has a higher priority than the audio / video call type.

8. The method as described in claim 7, characterized in that, When the first voice service type is an audio / video call type, and the voice data contains a second wake-up word, the method further includes: Send a third control command to the audio / video call module corresponding to the first voice service type, so that the audio / video call module stops using the recording channel corresponding to the audio / video call type to collect voice data; and... Send a recording permission command to the far-field voice module that is woken up based on the second wake word, so that the far-field voice module can collect voice data through the recording channel corresponding to the far-field voice service type.

9. A smart playback device, characterized in that, include: The response module is used to respond to the first recording request and determine the first voice service type corresponding to the first recording request; Acquisition module: If there is no currently authorized target voice service type, or the priority of the first voice service type is higher than the priority of the currently authorized target voice service type, then allow the first recording request and acquire voice data of the first voice service type based on the recording channel corresponding to the first voice service type; The sending module is configured to send a first control instruction to the first recording module corresponding to the at least one voice service type when there is at least one voice service type with a lower priority than the first voice service type, and to send a second control instruction to the at least one first recording module after receiving the recording completion instruction corresponding to the first voice service type; wherein, the first control instruction is used to instruct the at least one first recording module not to use the corresponding recording channel to collect voice data, and the second control instruction is used to instruct the at least one first recording module to resume the use of the corresponding recording channel.

Citation Information

Patent Citations

  • Traffic network recording method and relevant device

    CN103841245A