Audio processing method and apparatus, and electronic device, storage medium and program product

WO2025185380A8PCT designated stage Publication Date: 2025-10-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/075912
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-04
Filing Date
2025-02-06
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

When the cascaded audio capture device captures the speaker's voice, it is easily interfered with by other people's voices, resulting in poor sound quality.

Method used

Through the main audio acquisition device and the main audio acquisition device of the main mode, the audio data is filtered and processed, and specific audio data is suppressed or enhanced to improve the sound quality.

Benefits of technology

It effectively reduces the interference of other sounds on the speaker and improves the sound effect collected by cascaded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075912_02102025_PF_FP_ABST
    Figure CN2025075912_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are an audio processing method and apparatus, and an electronic device, a storage medium and a program product. A master audio collection device is connected to at least one slave audio collection device. When executed by a master audio collection device, the method comprises: in response to a presentation mode enable operation triggered with respect to a first device, setting the first device as a presentation device, and setting a presentation identifier of the presentation device; acquiring a plurality of pieces of audio data collected by the master audio collection device and at least one slave audio collection device, and on the basis of the presentation identifier, screening the plurality of pieces of audio data to obtain first audio data collected by the presentation device; processing the plurality of pieces of audio data, so as to obtain audio data to be transmitted; and transmitting to a terminal device the audio data to be transmitted, and playing same.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, device, electronic device, storage medium and program product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 4, 2024, with application number 202410245694.9 and application name “An audio processing method, device, electronic device and storage medium”. Technical Field

[0002] The present application relates to the field of audio processing technology, and in particular to an audio processing method, device, electronic device, storage medium and program product.

[0003] Background of the Invention

[0004] Audio capture devices (e.g., microphones) are intelligent hardware that can be used in various scenarios, such as conferences and lectures. In larger venues (e.g., conference rooms, classrooms, theaters, etc.), multiple audio capture devices are typically cascaded to expand the pickup range. One of these devices is then connected to a terminal device (e.g., a conference terminal).

[0005] In the related art, to facilitate the use of cascaded devices, multiple audio capture devices in the cascade are typically controlled in a unified manner, that is, they are turned on and off uniformly. When the cascaded audio capture devices are turned on, when the speaker of a meeting uses any of the cascaded audio capture devices to speak, while that audio capture device captures the speaker's voice, other audio capture devices may also capture the voices of others. Ultimately, the sound captured by the terminal device contains not only the speaker's voice, but also the voices of others, which may interfere with the speaker's voice and affect the speaker's sound quality.

[0006] Therefore, the sound collected by the existing cascade equipment is relatively poor. Summary of the Invention

[0007] Embodiments of the present application provide an audio processing method, apparatus, electronic device, storage medium, and program product for improving the effect of sound collected by cascaded devices.

[0008] In one aspect, an embodiment of the present application provides an audio processing method, which is performed by a master audio acquisition device, the master audio acquisition device being connected to at least one slave audio acquisition device, the method comprising:

[0009] In response to a speaker mode activation operation triggered on a first device, setting the first device as a speaker device and setting a speaker identifier of the speaker device; wherein the first device is any one of the master audio acquisition device and the at least one slave audio acquisition device;

[0010] Acquire multiple audio data collected by the master audio collection device and the at least one slave audio collection device, and based on the speaker identifier, filter out first audio data collected by the master audio collection device from the multiple audio data;

[0011] Processing the plurality of audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, suppressing other audio data in the plurality of audio data except the first audio data; and

[0012] The audio data to be transmitted is transmitted to the terminal device for playback.

[0013] On the other hand, an embodiment of the present application provides an audio processing method, which is performed by a first slave audio acquisition device among at least one slave audio acquisition device, wherein the at least one slave audio acquisition device is connected to a master audio acquisition device, and the method includes:

[0014] In response to a speaker mode activation operation triggered for the first slave audio acquisition device, sending a speaker activation request to the master audio acquisition device, wherein the speaker activation request includes a device identifier of the first slave audio acquisition device, so that the master audio acquisition device sets the first slave audio acquisition device as the speaker device and sets the device identifier of the first slave audio acquisition device to a speaker identifier; and

[0015] Collect first audio data and send the first audio data to the main audio acquisition device, so that the main audio acquisition device filters out the first audio data from the multiple audio data received based on the main speaker identifier, processes the multiple audio data, and transmits the processed audio data to the terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

[0016] On the other hand, an embodiment of the present application provides an audio processing device, which is connected to at least one slave audio acquisition device in a cascade manner, and includes:

[0017] A setting unit, configured to, in response to a speaker mode activation operation triggered on a first device, set the first device as a speaker device and set a speaker identifier of the speaker device; wherein the first device is any one of the master audio acquisition device and the at least one slave audio acquisition device;

[0018] a screening unit, configured to obtain a plurality of audio data collected by the master audio collection device and the at least one slave audio collection device, and based on the speaker identifier, screen out the first audio data collected by the master audio collection device from the plurality of audio data;

[0019] a processing unit, configured to process the plurality of audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, suppressing other audio data in the plurality of audio data except the first audio data; and

[0020] The first transmission unit is configured to transmit the audio data to be transmitted to a terminal device for playback.

[0021] On the other hand, an embodiment of the present application provides an audio processing device, including at least one slave audio acquisition device including the device, the at least one slave audio acquisition device being connected to a master audio acquisition device, the device including:

[0022] a first sending unit, configured to, in response to a speaker mode start operation triggered on the first slave audio acquisition device, send a speaker start request to the master audio acquisition device, wherein the speaker start request includes a device identifier of the first slave audio acquisition device, so that the master audio acquisition device sets the first slave audio acquisition device as the speaker device and sets the device identifier of the first slave audio acquisition device to a speaker identifier; and

[0023] The acquisition unit is configured to acquire first audio data and transmit the first audio data to the primary audio acquisition device, so that the primary audio acquisition device filters out the first audio data from the received multiple audio data based on the speaker identifier, processes the multiple audio data, and transmits the processed audio data to the terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

[0024] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any one of the above-mentioned audio processing methods.

[0025] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to perform the steps of any one of the above-mentioned audio processing methods.

[0026] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the steps of any one of the above-mentioned audio processing methods.

[0027] BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0029] FIG1 is a connection diagram of a cascade device in one embodiment of the present application;

[0030] FIG2 is a schematic diagram of an application scenario of an audio processing method in an embodiment of the present application;

[0031] FIG3 is a connection diagram of a cascade device in an embodiment of the present application;

[0032] FIG4 is a schematic diagram of a PoE according to an embodiment of the present application;

[0033] FIG5 is a flow chart of an audio processing method according to an embodiment of the present application;

[0034] FIG6 is a schematic diagram of a master audio acquisition device transmitting audio data in a speaker mode according to an embodiment of the present application;

[0035] FIG7 is a logic diagram of an audio processing method in an embodiment of the present application;

[0036] FIG8 is a schematic diagram of a switching method of a speaker mode in an embodiment of the present application;

[0037] FIG9 is a schematic diagram of switching from an open microphone state to a speaker mode in an embodiment of the present application;

[0038] FIG10 is a schematic diagram of switching from a microphone-mute state to a speaker mode in an embodiment of the present application;

[0039] FIG11 is a schematic diagram of a speaker mode switched to an open microphone state in an embodiment of the present application;

[0040] FIG12 is a schematic diagram of a master audio acquisition device transmitting audio data in a normal mode according to an embodiment of the present application;

[0041] FIG13 is a flowchart of another audio processing method in an embodiment of the present application;

[0042] FIG14 is a schematic diagram of the structure of an audio processing device according to an embodiment of the present application;

[0043] FIG15 is a schematic diagram of the structure of another audio processing device in an embodiment of the present application;

[0044] FIG16 is a schematic diagram of the hardware structure of an electronic device to which an embodiment of the present application is applied;

[0045] FIG17 is a schematic diagram of the hardware structure of an electronic device using an embodiment of the present application.

[0046] Implementation Method

[0047] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.

[0048] Figure 1 shows a schematic diagram of a cascade device connection in one embodiment of the present application. Here, audio capture device 110 is connected to terminal device 150, and audio capture device 110, audio capture device 120, audio capture device 130, and audio capture device 140 are sequentially connected to form a cascade device. Thus, the cascade device, as a combination, can serve as an extension of the terminal device's sound pickup capabilities.

[0049] The audio acquisition device 110 connected to the terminal device 150 is called a primary audio acquisition device; the remaining audio acquisition devices 120 , 130 , and 140 are called secondary audio acquisition devices.

[0050] Since cascaded devices can cover a wider area, they can pick up sounds in a larger range, such as in a conference hall, making it easier for conference equipment to capture sounds clearly in larger venues.

[0051] The following is an introduction to some concepts involved in the embodiments of this application.

[0052] Cascading devices refers to connecting multiple audio capture devices together so they can work together and share audio signals. This includes a master audio capture device and multiple slave audio capture devices. Cascading devices can be connected in series or in one-to-many configurations. Cascading devices are often used to improve audio capture quality, increase audio input sources, or provide coverage over a wider area. They can be used in a variety of scenarios, such as conferences, presentations, performances, and recordings.

[0053] Power over Ethernet (PoE) is a technology that allows data and power to be transmitted simultaneously over Ethernet cables (such as Cat5e or Cat6). PoE simplifies network device wiring, reduces costs, and increases flexibility.

[0054] The word “exemplary” is used hereinafter to mean “serving as an example, example, or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0055] The terms "first" and "second" are used for descriptive purposes only and should not be construed as explicitly or implicitly indicating relative importance or the number of the technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0056] The present application relates to the field of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to realize data computing, storage, processing, and sharing.

[0057] Cloud conferencing is an efficient, convenient, and low-cost conferencing format based on cloud computing technology. Users can quickly and efficiently share voice, data, and video with teams and clients around the world through simple online operations. Cloud conferencing service providers handle the complex technical aspects of data transmission and processing.

[0058] At present, cloud conferencing mainly focuses on the Software as a Service (SaaS) model, providing services including telephone, network, video and other service forms. Video conferencing based on cloud computing is called cloud conferencing.

[0059] In cloud conferencing applications, data transmission, processing, and storage are all handled by the computer resources of the video conferencing manufacturer. Users no longer need to purchase expensive hardware or install cumbersome software. They only need to open a browser and log in to the corresponding interface to conduct efficient remote meetings.

[0060] Cloud conferencing systems support dynamic multi-server cluster deployment and offer multiple high-performance servers, significantly improving conference stability, security, and availability. In recent years, video conferencing has gained widespread popularity due to its ability to significantly improve communication efficiency, continuously reduce communication costs, and enhance internal management. It has been widely adopted in various sectors, including government, transportation, finance, carriers, education, and enterprises. The use of cloud computing has made video conferencing even more appealing in terms of convenience, speed, and ease of use, and is poised to usher in a new wave of video conferencing applications.

[0061] The audio processing method described in the embodiment of the present application can be applied to audio collection in cloud conference scenarios. The main audio collection device can obtain audio data and transmit it to the terminal device in the cloud conference scenario.

[0062] The following is a brief overview of the design concepts of the embodiments of the present application.

[0063] In the related art, in larger venues (such as conference rooms, classrooms, theaters, etc.), in order to expand the sound pickup range, multiple audio acquisition devices are usually cascaded to form a cascade device. In order to facilitate the use of the cascade device, the multiple audio acquisition devices in the cascade device are usually controlled in a unified manner, that is, they are turned off and on in a unified manner. After the cascade audio acquisition device is turned on, when the speaker of the meeting uses any one of the audio acquisition devices in the cascade device to speak, while the audio acquisition device is collecting the speaker's voice, other audio acquisition devices may also collect the voices of other people. Ultimately, the sound acquired by the terminal device contains not only the speaker's voice, but may also contain the voices of other people, which interferes with the speaker's voice and affects the speaker's sound effect. Therefore, the sound collected by the existing cascade device is relatively poor.

[0064] In view of this, the embodiments of the present application provide an audio processing method, apparatus, electronic device, storage medium and program product. Any device in the cascade device can trigger the speaker mode. In the speaker mode, the main audio acquisition device in the cascade device can suppress audio data other than the audio data collected by the speaker device (i.e., the device with the speaker mode turned on), or enhance the audio data collected by the speaker device, or both suppress audio data other than the audio data collected by the speaker device and enhance the audio data collected by the speaker device, thereby reducing the interference of other sounds on the voice of the speaker using the speaker device and improving the effect of the sound collected by the cascade device.

[0065] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application can be combined with each other if there is no conflict.

[0066] As shown in Figure 2, it is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario diagram includes multiple audio acquisition devices, a terminal device 230, other terminal devices 240, and a server 250. The multiple audio acquisition devices include a master audio acquisition device 210 and at least one slave audio acquisition device 220.

[0067] In an embodiment of the present application, the audio acquisition device may be a device with an audio acquisition function, such as a device including a microphone, including but not limited to a speaker, a recording device, etc.

[0068] The terminal device 230 and other terminal devices 240 include but are not limited to smart phones, tablet computers, laptops, desktop computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, smart speakers, smart watches and other devices.

[0069] Server 250 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0070] In an optional embodiment, the master audio acquisition device 210 and at least one slave audio acquisition device 220 among the multiple audio acquisition devices can be connected via a wired network or a wireless network. The cascade connection shown in FIG2 is merely exemplary, and connection methods include, but are not limited to, serial connection, one-to-many connection, and other connection methods. Each slave audio acquisition device can be directly or indirectly connected to the master audio acquisition device. The master audio acquisition device 210 can be connected to the terminal device 230 via a wired network or a wireless network.

[0071] In an optional implementation, the terminal device 230, other terminal devices 240 and the server 250 may be directly or indirectly connected via a wired network or a wireless network, which is not limited in this application.

[0072] It should be noted that the audio processing method in each embodiment of the present application can be executed by a master audio acquisition device or a slave audio acquisition device.

[0073] In some embodiments, taking a conference scenario as an example, a terminal device 230 and other terminal devices 240 can access an online conference, and a main audio capture device 210 is connected to the terminal device 230. During the meeting, a participant can trigger the main audio capture device 210 or enable speaker mode from the audio capture device 220. In response to a speaker mode activation operation triggered for any device, the main audio capture device 210 sets the device as the main speaker. In speaker mode, a speaker who needs to speak can speak through the main speaker device. At this point, the main speaker device can collect the speaker's audio data.

[0074] The slave audio collection device 220 may also collect other people's voices or noises, and the slave audio collection device 220 may transmit the collected audio data to the master audio collection device 210 .

[0075] The main audio acquisition device 210 can also collect audio data and filter out the audio data of the main speaker from the acquired audio data; then, it can suppress the audio data collected by non-main speaker devices, or enhance the audio data collected by the main speaker device, or both suppress the audio data collected by non-main speaker devices and enhance the audio data collected by the main speaker device, thereby obtaining the audio data to be transmitted and transmitting the audio data to be transmitted to the terminal device 230.

[0076] The terminal device 230 can send the received audio data to the server 250, and the server 250 sends the audio data to other terminal devices 240 for playback to achieve an online conference. At the same time, the terminal device 230 can also play the audio data.

[0077] In other embodiments, still taking the conference scenario as an example, multiple participants can conduct a local conference through multiple audio capture devices and terminal device 230, and the main audio capture device 210 is connected to the terminal device 230. In the speaker mode, the main audio capture device 210 filters out the audio data of the main speaker from the acquired audio data; then, it can suppress the audio data collected by non-main speaker devices, or enhance the audio data collected by the main speaker device, or both suppress the audio data collected by non-main speaker devices and enhance the audio data collected by the main speaker device, thereby obtaining the audio data to be transmitted, and transmitting the audio data to be transmitted to the terminal device 230, which then plays the received audio data to achieve sound amplification.

[0078] In other embodiments, taking a recording scenario as an example, multiple recorders can record audio through multiple audio capture devices and a terminal device 230, wherein a master audio capture device 210 is connected to the terminal device 230. During the recording process, the user can trigger the master audio capture device 210 or the slave audio capture device 220 to enable speaker mode. In response to the speaker mode activation operation triggered on either device, the master audio capture device 210 sets the device as the master speaker. In speaker mode, the speaker who needs to record the audio can record the audio through the master speaker device. At this time, the main speaker device can collect the audio data of the speaker, and the non-main speaker devices may also collect the voices or noises of other people. The slave audio collection device 220 can transmit the collected audio data to the main audio collection device 210. The main audio collection device 210 can also collect audio data and filter out the audio data of the main speaker device from the acquired audio data. Afterwards, the audio data collected by the non-main speaker devices can be suppressed, or the audio data collected by the main speaker device can be enhanced, or the audio data collected by the non-main speaker devices can be suppressed and the audio data collected by the main speaker device can be enhanced, so as to obtain the audio data to be transmitted, and transmit the audio data to be transmitted to the terminal device 230. The terminal device 230 saves the received audio data to realize recording so that the user can play the audio data through the terminal device 230.

[0079] The above-mentioned conference scenarios include various conference scenarios, including but not limited to education and training, corporate meetings, product launches, etc., which can be online conference scenarios or local conference scenarios. In addition, in addition to the above-mentioned conference scenarios and recording scenarios, the audio processing method of the embodiment of the present application can also be applied to other scenarios, including but not limited to speech scenarios, performance scenarios, etc. The audio collection process in these scenarios is similar to the audio collection process in the above-mentioned conference scenarios, and will not be repeated here.

[0080] It should be noted that what is shown in FIG2 is only an example. In fact, the number of audio acquisition devices and terminal devices is not limited and is not specifically limited in the embodiments of the present application.

[0081] The following describes the audio processing method provided by the exemplary embodiment of the present application in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.

[0082] Before introducing the audio processing method of the embodiment of the present application, the cascade device for implementing the audio processing method of the embodiment of the present application is first introduced.

[0083] The cascaded devices of the present embodiment include a master audio capture device and at least one slave audio capture device, wherein the master audio capture device and the at least one slave audio capture device are connected. When there are multiple slave audio capture devices, the cascaded devices include various connection methods. Several possible connection methods are exemplarily described below.

[0084] A first possible connection mode: a master audio acquisition device and multiple slave audio acquisition devices may be connected in series in a preset order.

[0085] For example, the master audio acquisition device is connected to the slave audio acquisition device 1, the slave audio acquisition device 1 is connected to the slave audio acquisition device 2, the master audio acquisition device 2 is connected to the slave audio acquisition device 3, and so on.

[0086] In this connection mode, the master audio acquisition device and the slave audio acquisition device 2 can communicate indirectly through the slave audio acquisition device 1, and the master audio acquisition device and the slave audio acquisition device 3 can communicate indirectly through the slave audio acquisition device 1 and the slave audio acquisition device 2. That is, when the master audio acquisition device needs to transmit first data to each slave audio acquisition device, it can transmit the first data to the slave audio acquisition device 1, which in turn transmits the first data to the slave audio acquisition device 2, which in turn transmits the first data to the slave audio acquisition device 3. Correspondingly, when the slave audio acquisition device 2 needs to transmit second data to the master audio acquisition device, it can transmit the second data to the slave audio acquisition device 1, which in turn transmits the second data to the slave audio acquisition device 2, and so on.

[0087] The second possible connection mode is that the master audio acquisition device is directly connected to multiple slave audio acquisition devices, that is, a one-to-many connection.

[0088] For example, the master audio acquisition device is connected to slave audio acquisition device 1, slave audio acquisition device 2, slave audio acquisition device 3, etc. respectively.

[0089] In this connection mode, the master audio acquisition device and each slave audio acquisition device 2 can communicate directly.

[0090] A third possible connection mode is to use both serial connection and one-to-many connection modes between the master audio capture device and multiple slave audio capture devices.

[0091] Specifically, a master audio acquisition device establishes a one-to-many connection with at least two slave audio acquisition devices, and at least two of the multiple slave audio acquisition devices are connected in series. For example, the master audio acquisition device is connected to slave audio acquisition device 1 and slave audio acquisition device 2, and the master audio acquisition device 2 is connected to slave audio acquisition device 3.

[0092] In the various connection methods described above, the master audio acquisition device and the multiple slave audio acquisition devices can be connected via a wired network or a wireless network, that is, the master audio acquisition device and the multiple slave audio acquisition devices are connected via a wired network, or the master audio acquisition device and the multiple slave audio acquisition devices are connected via a wireless network, or there is both a wireless network connection and a wired network connection between the master audio acquisition device and the multiple slave audio acquisition devices. For example, wireless networks include but are not limited to Bluetooth, Wireless Fidelity (WiFi), etc., and wired networks can be implemented through various network cables, such as PoE.

[0093] For example, as shown in FIG3 , taking the example of a master audio capture device and multiple slave audio capture devices connected in series via PoE, the master audio capture device 310 is connected to the slave audio capture device 320 via PoE, the master audio capture device 320 is connected to the slave audio capture device 330 via PoE, the slave audio capture device 330 is connected to the slave audio capture device 340 via PoE, and the slave audio capture device 340 is connected to the slave audio capture device 350 via PoE. The master audio capture device 310 can be connected to a terminal device via a wired network or a wireless network, and the terminal device can be, for example, the terminal device 230 in the above embodiment.

[0094] PoE can transmit data and power simultaneously through Ethernet cables. As shown in Figure 4, a PoE system mainly consists of the following components:

[0095] a. Power supply terminal 410, which is responsible for providing power to connected devices.

[0096] b. The power receiving end 420 is a device that receives power provided through the Ethernet cable 430.

[0097] c. Ethernet cable 430, which is used to transmit data and power between the power supply end and the power receiving end.

[0098] In the above Figure 3, the master audio capture device 310 is the power supply end of the slave audio capture device 320, and the slave audio capture device 320 is the power receiving end of the master audio capture device 310; the slave audio capture device 320 is the power supply end of the slave audio capture device 330, and the slave audio capture device 330 is the power receiving end of the slave audio capture device 320; the slave audio capture device 330 is the power supply end of the slave audio capture device 340, and the slave audio capture device 340 is the power receiving end of the slave audio capture device 330.

[0099] Referring to FIG5 , which is a flowchart of an implementation of an audio processing method provided in an embodiment of the present application, the execution subject is a master audio acquisition device, such as the master audio acquisition device 110 in FIG1 or the master audio acquisition device 210 in FIG2 . The master audio acquisition device is connected to at least one slave audio acquisition device to form a cascade device. The specific implementation process of the method includes the following steps S51-S53:

[0100] S51. In response to a speaker mode activation operation triggered on a first device, the first device is set as a speaker device and a speaker identifier of the speaker device is set; wherein the first device is any one of a master audio acquisition device and at least one slave audio acquisition device.

[0101] The user can trigger speaker mode for any cascaded device. The triggering method can be set as needed and is not limited. For example, triggering a setting button on any device, including but not limited to long pressing, short pressing, and rotating the setting button; another example, clicking a setting control displayed on any device, including but not limited to single-clicking and double-clicking the setting control; another example, sliding a setting control displayed on any device; and another example, sending a voice control command to any device, such as "turn on speaker mode."

[0102] The first device may be a main audio acquisition device, or any audio acquisition device.

[0103] In one embodiment, a user can trigger a speaker mode activation operation for any slave audio capture device. In this case, the slave audio capture device is referred to as a first slave audio capture device. As the first device, the slave audio capture device can respond to the speaker mode activation operation and send a speaker activation request to the master audio capture device.

[0104] In the cascade mode, when the first slave audio acquisition device is directly connected to the master audio acquisition device, the first slave audio acquisition device can send the start-speaker request to the master audio acquisition device; when the first slave audio acquisition device is indirectly connected to the master audio acquisition device through other slave audio acquisition devices, the first slave audio acquisition device can forward the start-speaker request to the master audio acquisition device through other slave audio acquisition devices.

[0105] In another embodiment, the user can trigger the speaker mode start operation for the main audio acquisition device. At this time, the main audio acquisition device, as the above-mentioned first device, responds to the triggered speaker mode start operation, sets the main audio acquisition device as the speaker device, and uses the device identifier of the main audio acquisition device as the speaker identifier.

[0106] In the embodiment of the present application, the master audio acquisition device and the slave audio acquisition device may be connected directly or indirectly via a network.

[0107] In one embodiment, the device identifier of the master audio collection device and the device identifier of the slave audio collection device may be a network address, that is, an Internet Protocol Address (IP).

[0108] The master audio capture device and each slave audio capture device communicate via network addresses. Specifically, when a slave audio capture device sends data to the master audio capture device, the destination address of the data can be set to the network address of the master audio capture device; correspondingly, when the master audio capture device sends data to a slave audio capture device, the destination address of the data can be set to the network address of the slave audio capture device.

[0109] In addition, the device identifier may also be other identifiers that can uniquely identify the master audio acquisition device or the slave audio acquisition device, and this is not limited.

[0110] S52: Acquire multiple audio data collected by a master audio collection device and at least one slave audio collection device, and filter out first audio data collected by the master audio collection device from the multiple audio data based on the master speaker identifier.

[0111] In this step, if the main audio acquisition device obtains the audio data collected by at least one device, the main speaker device is filtered out from at least one device based on the speaker identifier and the device identifier carried by at least one audio data; wherein the at least one device includes: at least one slave audio acquisition device and part or all of the main audio acquisition device.

[0112] When either device is in speaker mode, the primary and secondary audio capture devices can each collect audio data, such as the speaker's voice, other people's voices, and other possible noise. The speaker is the person who needs to speak individually in various activity scenarios.

[0113] Each slave audio capture device can send the collected audio data to the master audio capture device. This audio data can carry its own device identification. The master audio capture device can obtain the audio data collected by at least one device and the device identification of at least one device. It can then compare the speaker identification with the obtained device identification to filter out the speaker from the at least one device, thereby obtaining the audio data collected by the master device.

[0114] In one embodiment, the main audio collection device serves as the main speaker device, and thus, the audio data collected by the main audio collection device can be directly determined as the first audio data.

[0115] In another embodiment, the first slave audio capture device serves as the master audio capture device, and the audio data captured by the first slave audio capture device is the first audio data. The first audio data is included in the multiple audio data captured by the master audio capture device. The master audio capture device can indicate that the master audio capture device is the first slave audio capture device based on the master speaker identifier, and filter out the first audio data from the mixed multiple audio data.

[0116] S53. Process the multiple audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, and suppressing the other audio data in the multiple audio data except the first audio data.

[0117] In this step, the processing may include: only enhancing the first audio data collected by the main speaker device, or only suppressing the other audio data except the first audio data collected by the main speaker device, or both enhancing the first audio data collected by the main speaker device and suppressing the other audio data except the first audio data collected by the main speaker device.

[0118] The enhancement process is used to increase the volume of the audio data collected by the main speaker device. Specifically, the enhancement process includes enhancing the amplitude of the audio data collected by the main speaker device. For example, the amplitude of the audio data collected by the main speaker device can be multiplied by a coefficient greater than 1. The value of the coefficient can be set as needed.

[0119] Suppression is used to eliminate or reduce the volume of other audio data (except the audio data collected by the speaker device). Specifically, suppression includes filtering out other audio data and reducing the amplitude of other audio data. For example, the amplitude of other audio data can be multiplied by a coefficient less than 1. The value of this coefficient can be set as needed.

[0120] Several possible implementations of S53 are described below.

[0121] In one possible implementation, other audio data except the first audio data collected by the speaker device can be filtered out from the multiple audio data. The filtered audio data includes the first audio data and a small amount of noise. Then, the first audio data is extracted from the filtered audio data as the audio data to be transmitted.

[0122] For example, as shown in FIG6 , a master device (i.e., a master audio capture device) captures audio data 610, where the master device's device identifier is known. Simultaneously, audio data 620 captured by a slave device (i.e., a slave audio capture device) is received via a network (wired or wireless), including the slave device's device identifier. In speaker mode 630, the master device selects the audio data captured by the speaker device from among the acquired audio data, specifically performing channel selection. Specifically, the master device selects the channel corresponding to the speaker identifier 640 from among the channels corresponding to the various device identifiers. Finally, the audio data 650 captured by the speaker device is output to a terminal device 660 as the audio data to be transmitted for playback.

[0123] In another possible implementation, during the mixing process of multiple audio data, the first audio data is enhanced and the other audio data are suppressed to obtain the audio data to be transmitted.

[0124] The audio mixing process is a process of combining multiple audio data into a single audio output stream. In this process, the above-mentioned processing can be performed on multiple audio data.

[0125] In another possible implementation, the first audio data is enhanced, the other audio data except the first audio data among the multiple audio data are suppressed, and the processed multiple audio data are mixed to obtain the audio data to be transmitted.

[0126] In the embodiments of the present application, to reduce interference of other audio data with the first audio data of the main speaker, the other audio data can be suppressed, or the audio data of the main speaker can be enhanced, or both the audio data collected by the main speaker can be enhanced and the other audio data can be suppressed. In this way, the quality of the main speaker's voice in the audio data to be transmitted can be improved, thereby improving the clarity and volume of the audio when played back by the terminal device.

[0127] S54: Transmit the audio data to be transmitted to the terminal device for playback.

[0128] The terminal device may be the terminal device 230 or other terminal device 240 in the above-mentioned embodiment of the present application. The main audio acquisition device may transmit the audio data to be transmitted directly or indirectly to the terminal device. In other words, the terminal device may include a terminal device connected to the main audio acquisition device, or may include other terminal devices connected to the terminal device, without limitation.

[0129] For example, in a local conference scenario, the terminal device may specifically be a local conference terminal, and the main audio acquisition device may be connected to the local conference terminal to transmit the audio data to be transmitted to the local conference terminal for playback.

[0130] For another example, in an online conference scenario, the main audio acquisition device is connected to the local conference terminal, and the local conference terminal and other conference terminals are connected to the online conference at the same time. After the main audio acquisition device transmits the audio data to be transmitted to the local conference terminal, the local conference terminal transmits the received audio data to other conference terminals, and the other conference terminals play the audio data. That is, the terminal device in the above S53 can be other conference terminals, and at the same time, the local conference terminal can also play the audio data.

[0131] It should be noted that the master audio acquisition device and the slave audio acquisition device in the embodiment of the present application can be located in the same space or in different spaces. For example, the master audio acquisition device and the slave audio acquisition device are located in different conference rooms.

[0132] The overall logic of the audio processing method according to the embodiment of the present application is exemplarily introduced below with reference to FIG7 .

[0133] 7 , the slave audio capture device is referred to as the slave device, and the master audio capture device is referred to as the master device. Assuming that the user triggers a speaker mode activation operation on slave device 2, in step 710, slave device 2 responds to the speaker mode activation operation by sending a speaker activation request to the master device.

[0134] In 720 , the master device determines to start the speaker mode, and sets slave device 2 as the speaker device, and does not save the device identifier of slave device 2 as the speaker identifier. In this way, the entire cascade device enters the speaker mode.

[0135] In 730, in the speaker mode, slave device 1 can send the collected audio data 1 to the master device, slave device 2 can send the collected audio data 2 to the master device, slave device 3 can send the collected audio data 3 to the master device, slave device 4 can send the collected audio data 4 to the master device, slave device 5 can send the collected audio data 5 to the master device, and at the same time, the master device can also collect audio data 6.

[0136] In 740, after acquiring the various audio data, the master device may filter out the audio data 2 collected by the slave device 2 (i.e., the main speaker device) and perform the above-mentioned processing on the various audio data to obtain the audio data to be transmitted. For example, the extracted audio data 2 may be transmitted to the terminal device.

[0137] Any device in the embodiments of the present application (the master audio capture device or the slave audio capture device) can trigger the speaker mode. In the speaker mode, the speaker can speak through the master audio capture device. When the master audio capture device obtains audio data collected by multiple devices, it can filter out the master device from the multiple devices based on the speaker identifier, and then suppress the audio data collected by non-master devices and / or enhance the audio data collected by the master device. In this way, even if the non-master speaker makes a sound, the impact on the speaker's voice is relatively small, which can reduce the interference of other sounds on the speaker's voice, thereby improving the clarity and volume of the sound collected by the cascaded devices.

[0138] At the same time, in the embodiment of the present application, any device can be set as the main speaking device according to the location of the speaker. The person who needs to speak can use any device to speak as the speaker, thereby improving the flexibility of the cascaded devices.

[0139] It should be noted that when any device triggers the speaker mode and turns on the speaker mode, other devices can also trigger the speaker mode, that is, multiple devices can serve as speaker devices at the same time. The following embodiments are described by taking one device as the speaker device as an example.

[0140] The following describes the specific implementation process of triggering the speaker mode from the audio capture device.

[0141] In some embodiments, when any slave audio capture device triggers the speaker mode, the first device is the first slave audio capture device. In S51 of the above embodiment, the master audio capture device responds to the speaker mode activation operation triggered on the first device, sets the first device as the speaker device, and sets the speaker identifier of the speaker device, specifically including:

[0142] In response to receiving the speaker start request sent by the first slave audio acquisition device, setting the first slave audio acquisition device as the speaker device, and setting the device identifier carried in the speaker start request as the speaker identifier;

[0143] The speaker start request is sent by the first slave audio acquisition device in response to the speaker mode start operation triggered on the first slave audio acquisition device.

[0144] Specifically, users can trigger the speaker mode for any slave audio capture device. The triggering method can be set as needed and is not limited to this. For example, triggering a setting button on any slave audio capture device, including but not limited to long pressing the setting button, short pressing the setting button, rotating the setting button, etc.; another example, clicking a setting control displayed on any slave audio capture device, including but not limited to single-clicking the setting control, double-clicking the setting control, etc.; another example, sliding the setting control displayed on any device; another example, sending a voice control command to any slave audio capture device, such as "turn on speaker mode."

[0145] In response to the triggered speaker mode activation operation, any slave audio capture device sends a speaker activation request carrying the device identification of the slave audio capture device to the master audio capture device. In response to the speaker activation request, the master audio capture device sets the slave audio capture device as the speaker device and saves the device identification of the slave audio capture device as the speaker identification.

[0146] Specifically, any slave audio capture device can display an open microphone state or a closed microphone state, that is, any slave audio capture device can switch from an open microphone state to a speaker mode, or from a closed microphone state to a speaker mode. The operation method for triggering the speaker mode in these two cases can be the same or different.

[0147] Among them, the open microphone state can indicate that the slave audio capture device is available, that is, the device is turned on, and the audio data collected by the device needs to be transmitted; the closed microphone state can indicate that the slave audio capture device is unavailable. On the one hand, unavailable can mean that the device is turned off, and on the other hand, unavailable can also mean that the device is turned on, but the audio data collected by the device is not transmitted.

[0148] Switching between the microphone-on state, the microphone-off state, and the speaker mode of any device in the embodiments of the present application may specifically include the following three scenarios:

[0149] (1) The microphone status switches to speaker mode.

[0150] (2) Switch from microphone-mute mode to speaker mode.

[0151] (3) The speaker mode is switched to the microphone open state.

[0152] The triggering operation for switching from the open mic state to the speaker mode can be the same as or different from the triggering operation for switching from the closed mic state to the speaker mode. The triggering operation for switching from the speaker mode to the open mic state can be different from the triggering operation for switching from the open mic state to the speaker mode, and different from the triggering operation for switching from the closed mic state to the speaker mode.

[0153] For example, as shown in Figure 8, assuming that each device is provided with a mute button, the operation to trigger the speaker mode may be a long press of the mute button, and the operation to exit the speaker mode may be a short press of the mute button. Then, when any device is in the microphone-off state or the microphone-on state, the mute button 810 of the device can be long pressed to trigger the speaker mode; after the speaker mode is turned on for the device, the mute button 820 of the device can be short pressed to switch from the speaker mode to the microphone-on state, that is, to exit the speaker mode.

[0154] The microphone on and microphone off states can be displayed in different ways, which can be set as needed and are not limited to this. For example, the microphone on and microphone off states can be indicated by displaying different colors of indicator lights, such as a green light indicating an on state and a red light indicating a off state. Another example is the display of different patterns indicating the microphone on and microphone off states. Another example is the display of different text indicating the microphone on and microphone off states, such as "microphone on" indicating an on state and "microphone off" indicating a off state.

[0155] In an embodiment of the present application, when the speaker mode is triggered for any slave audio capture device, any slave audio capture device can send a speaker start request to the master audio capture device to start the speaker mode, so as to facilitate the speaker to speak using any slave audio capture device, thereby improving the flexibility of the cascaded device.

[0156] In some embodiments, after receiving a speaker start request from any slave audio collection device and setting any slave audio collection device as the speaker device, the master audio collection device may further perform the following steps A1-A2:

[0157] A1. Set and display the microphone muting status of the main audio capture device.

[0158] Among them, after determining that any slave audio collection device is the main speaker device, the main audio collection device determines itself as a non-main speaker device, sets and displays the microphone-mute status. The display method of the microphone-mute status can be found in the above embodiment of this application and will not be repeated here.

[0159] A2. Send a first notification message to at least one slave audio collection device, where the first notification message carries a speaker identifier. The first notification message is used to instruct any of the at least one slave audio collection device to turn on a speaker mode, so that

[0160] The first slave audio capture device displays the microphone open status according to the speaker ID;

[0161] The other slave audio collection devices other than the first slave audio collection device display a muted microphone status according to the speaker ID.

[0162] After the master audio acquisition device determines that the speaker mode is enabled, it can send a first notification message to each slave audio acquisition device. Specifically, the device identifier of each slave audio acquisition device can be a network address, and the destination address of the first notification message can include the network address of each slave audio acquisition device. In this way, when slave audio acquisition device 2 is indirectly connected to the master audio acquisition device through slave audio acquisition device 1, after receiving the first notification message sent by the master audio acquisition device, slave audio acquisition device 1 can determine, based on the destination address of the first notification message, that the first notification message needs to be sent to the connected slave audio acquisition device 2, thereby achieving the broadcast of the first notification message.

[0163] Upon receiving the first notification message, any slave audio capture device that triggers the speaker mode determines that the speaker mode is enabled and identifies itself as the speaker device based on the speaker identifier, and may display the microphone-on status. Upon receiving the first notification message, other slave audio capture devices determine that the speaker mode is enabled and identify themselves as non-speaker devices based on the speaker identifier, and may display the microphone-off status.

[0164] In one embodiment, when the master audio capture device and each slave audio capture device are in the open microphone state, when any slave audio capture device triggers the speaker mode, the slave audio capture device remains in the open microphone state, and the other slave audio capture devices switch from the open microphone state to the closed microphone state.

[0165] For example, as shown in Figure 9, the slave audio capture device is referred to as the slave device, and the master audio capture device is referred to as the master device. Assume that the master device is connected in series with slave devices 1, 2, 3, 4, and 5 in sequence via PoE, and the master device is connected to the terminal device via a wired or wireless network.

[0166] When the master device and each slave device are in the microphone-on state, in 910 , the user triggers the speaker mode for slave device 2 , for example, by long pressing the mute button of slave device 2 .

[0167] In 920 , in response to the trigger operation (ie, the above-mentioned speaker mode activation operation), slave device 2 sends a speaker activation request carrying a device identifier to the master device.

[0168] In 930 , the master device determines to start the speaker mode according to the speaker start request, and sets the device identifier of the slave device 2 as the speaker identifier for storage.

[0169] In 940 , the master device may switch from the open microphone state to the closed microphone state and broadcast a first notification message to each slave device, so that slave device 2 remains in the open microphone state and the other slave devices except slave device 2 switch from the open microphone state to the closed microphone state.

[0170] In another embodiment, when the master audio acquisition device and each slave audio acquisition device are in the mute state, when any slave audio acquisition device triggers the speaker mode, the slave audio acquisition device switches from the mute state to the open state, and the other slave audio acquisition devices remain in the mute state.

[0171] For example, as shown in FIG10, still taking the master device and each slave device in FIG9 as an example, when the master device and each slave device are both in the mute state,

[0172] In 1010 , the user triggers the speaker mode for the slave device 2 , for example, by long pressing the mute button of the slave device 2 .

[0173] In 1020 , in response to the trigger operation, slave device 2 sends a request to start speaking with the device identification to the master device.

[0174] In 1030 , the master device determines to start the speaker mode according to the speaker start request, and sets the device identifier of the slave device 2 as the speaker identifier for storage.

[0175] In 1040 , the master device may maintain the microphone-mute state and broadcast a first notification message to each slave device, so that slave device 2 switches from the microphone-mute state to the microphone-open state, and the slave devices other than slave device 2 maintain the microphone-mute state.

[0176] In an embodiment of the present application, after the speaker mode is triggered to be turned on by the slave audio capture device, the master audio capture device sends a first notification message to each slave audio capture device, so that the speaker device displays the microphone-on state and the non-speaker device displays the microphone-off state, so that the user can intuitively determine the speaker device and the non-speaker device, which is convenient for the speaker to accurately use the speaker device to speak, so that the speaker's audio data can be collected and transmitted to the terminal device for playback.

[0177] In some embodiments, after the master audio capture device sets any slave audio capture device as the master device, the speaker can use the master device to speak. When the speech is finished, the master device can be triggered to exit the master mode. At this time, the master audio capture device can perform the following steps B1-B2:

[0178] B1. In response to receiving a request to exit the main speaker sent by the first slave audio acquisition device, switching the microphone-off state of the master audio acquisition device to the microphone-on state.

[0179] Among them, any slave audio acquisition device can respond to the trigger operation for exiting the speaker mode and send an exit speaker request to the master audio acquisition device. In this way, the master audio acquisition device determines to exit the speaker mode and switches its own closed microphone state to open microphone state.

[0180] Among them, the triggering operation for exiting the speaker mode can be set as needed and is not limited to this. For example, when any setting button on the audio acquisition device is triggered to turn on the speaker mode, the setting button can also be triggered to exit the speaker mode, such as long pressing the setting button to turn on the speaker mode, short pressing the setting button to exit the speaker mode, or short pressing the setting button to turn on the speaker mode, long pressing the setting button to exit the speaker mode; for another example, when any setting control displayed by the audio acquisition device is triggered to turn on the speaker mode, the setting control can also be triggered to exit the speaker mode, such as single-clicking the setting control to turn on the speaker mode, double-clicking the setting control to exit the speaker mode, or double-clicking the setting control to turn on the speaker mode, single-clicking the setting control to exit the speaker mode. The above-mentioned triggering operation for turning on the speaker mode and the triggering operation for exiting the speaker mode are only exemplary and are not limited to the embodiments of the present application.

[0181] B2. Send a second notification message to at least one slave audio collection device, where the second notification message is used to instruct the other slave audio collection devices to switch from a mute state to an open state.

[0182] The second notification message is used to instruct each slave audio collection device to exit the speaker mode, so that the other slave audio collection devices except the first slave audio collection device are switched from the closed microphone state to the open microphone state.

[0183] After determining to exit the speaker mode, the master audio collection device may send a second notification message to each slave audio collection device to broadcast a message of exiting the speaker mode.

[0184] Specifically, the device identifier of each slave audio acquisition device can be a network address, and the destination address of the second notification message can include the network address of each slave audio acquisition device. In this way, when the slave audio acquisition device 2 is indirectly connected to the master audio acquisition device through the slave audio acquisition device 1, after the slave audio acquisition device 1 receives the second notification message sent by the master audio acquisition device, it can determine that the second notification message needs to be sent to the connected slave audio acquisition device 2 based on the destination address of the second notification message, thereby realizing the broadcast of the second notification message.

[0185] For example, as shown in FIG11 , still taking the master device and each slave device in FIG9 as an example,

[0186] In 1110 , after the speaker mode is turned on on the slave device 2 , the speaker can trigger the exit of the speaker mode on the slave device 2 after finishing speaking on the slave device 2 , for example, by short pressing the mute button of the slave device 2 .

[0187] In 1120 , in response to the trigger operation, slave device 2 sends a request to exit the presentation to the master device.

[0188] In 1130 , the master device determines to exit the speaker mode according to the exit speaker request.

[0189] In 1140 , the master device may switch from the mute state to the open state, and broadcast a second notification message to each slave device, so that the slave devices other than slave device 2 switch from the mute state to the open state.

[0190] In an embodiment of the present application, after any slave audio acquisition device exits the speaker mode, the master audio acquisition device sends a second notification message to each slave audio acquisition device to switch the non-speaker device from the closed microphone state to the open microphone state, so that the user can intuitively determine the exit from the speaker mode. Since all devices are in the open microphone state, that is, in normal mode, it is convenient for the user to use any device to speak, thereby ensuring the convenience of use of cascaded devices.

[0191] The following describes the specific implementation process of the main audio acquisition device triggering the main speaker mode.

[0192] In some embodiments, when the speaker mode is triggered for the primary audio capture device, the first device is the primary audio capture device. In response to the speaker mode activation operation triggered for the first device in S51 of the above embodiment, the first device is set as the speaker device and the speaker identifier of the speaker device is set. Specifically, the following steps may be performed:

[0193] In response to a speaker mode start operation triggered on the primary audio acquisition device, the primary audio acquisition device is set as a speaker device, and a microphone-on state of the primary audio acquisition device is set.

[0194] Specifically, the method for triggering the speaker mode on the primary audio capture device can be set as needed and is not limited to this. For example, triggering a setting button on the primary audio capture device includes, but is not limited to, long pressing the setting button, short pressing the setting button, rotating the setting button, etc.; another example is clicking a setting control displayed on the primary audio capture device, including, but not limited to single-clicking the setting control, double-clicking the setting control, etc.; another example is sliding a setting control displayed on any device; another example is sending a voice control command to the primary audio capture device, such as "turn on speaker mode."

[0195] Furthermore, the master audio collection device sends a third notification message to at least one slave audio collection device, where the third notification message is used to instruct the at least one slave audio collection device to display a microphone-mute state.

[0196] Specifically, the third notification message may include a speaker identifier, which is used to instruct at least one slave audio collection device to turn on the speaker mode, so that the at least one slave audio collection device displays a microphone-mute state according to the speaker identifier.

[0197] In an embodiment of the present application, when the main audio acquisition device triggers the speaker mode, the main audio acquisition device sends a third notification message to each slave audio acquisition device to make each slave audio acquisition device display a muted microphone state, so that the user can intuitively determine the main speaker device and the non-main speaker device, so that the speaker can accurately use the main speaker device to speak, so that the speaker's audio data can be collected and transmitted to the terminal device.

[0198] In some embodiments, in response to the speaker mode activation operation triggered on the primary audio capture device in step C1, the primary audio capture device is set as the speaker device, and the microphone-on status of the primary audio capture device is displayed, specifically including the following two situations:

[0199] In the first case, when the main audio acquisition device currently displays the open microphone state, in response to the first speaker mode start operation triggered for the main audio acquisition device, the main audio acquisition device is set as the speaker device and the open microphone state of the main audio acquisition device is maintained.

[0200] In the second case, when the main audio acquisition device currently displays the closed microphone state, in response to the second speaker mode start operation triggered for the main audio acquisition device, the main audio acquisition device is set as the speaker device, and the closed microphone state of the main audio acquisition device is switched to the open microphone state.

[0201] Among them, the main audio acquisition device can display the microphone-on state or the microphone-off state, that is, the main audio acquisition device can switch from the microphone-on state to the speaker mode, and can also switch from the microphone-off state to the speaker mode. The operation method of triggering the speaker mode in these two cases can be the same or different, that is, the above-mentioned first speaker mode activation operation and the second speaker mode activation operation can be the same or different. For details, please refer to the operation method of triggering the speaker mode in the above embodiment of this application.

[0202] In the embodiment of the present application, the main audio acquisition device can switch to the main speaker mode in the microphone-on state, and can also switch to the main speaker mode in the microphone-off state, thereby realizing multiple switching scenarios of the main speaker mode and improving the switching flexibility of the main speaker mode.

[0203] In some embodiments, after the primary audio capture device sets itself as the main speaker device, the speaker can use the main speaker device to speak. When the speech is finished, the main speaker device can be triggered to exit the main speaker mode. At this time, the primary audio capture device can perform the following steps:

[0204] In response to the speaker mode exit operation triggered for the master audio acquisition device, a fourth notification message is sent to at least one slave audio acquisition device, where the fourth notification message is used to instruct the at least one slave audio acquisition device to switch from a closed microphone state to an open microphone state.

[0205] Specifically, the fourth notification message is used to instruct at least one slave audio collection device to exit the speaker mode, so that the at least one slave audio collection device switches from a closed microphone state to an open microphone state.

[0206] In an embodiment of the present application, after any slave audio acquisition device exits the speaker mode, the master audio acquisition device sends a second notification message to each slave audio acquisition device to switch the non-speaker device from the closed microphone state to the open microphone state, so that the user can intuitively determine the exit from the speaker mode. Since all devices are in the open microphone state, that is, in normal mode, it is convenient for the user to use any device to speak, thereby ensuring the convenience of use of cascaded devices.

[0207] In some embodiments, when the main speaker device exits speaker mode, the main audio capture device and at least one slave audio capture device may both be in open microphone mode. In this case, the main audio capture device and the slave audio capture device may capture audio data, and the slave audio capture device may transmit the captured audio data to the main audio capture device. Thus, the main audio capture device may obtain the audio data captured by at least one device. The at least one device includes at least one slave audio capture device and some or all of the main audio capture device.

[0208] In some embodiments, after exiting the speaker mode, when multiple audio data collected by the master audio collection device and the at least one slave audio collection device are obtained, the multiple audio data are mixed to obtain mixed data;

[0209] The mixed audio data is transmitted to the terminal device for playback.

[0210] Specifically, in the process of mixing the audio data of multiple devices, the amplitudes of the audio data of the multiple devices can be adjusted according to the weight coefficients of the multiple devices. For example, the amplitude of the audio data of each device is multiplied by the corresponding weight coefficient. Among them, the weight coefficients of the multiple devices can be the same or different, and the sizes of the weight coefficients of the multiple devices can be set as needed, and there is no limitation on this. For example, the weight coefficient of the main audio acquisition device is set to be greater than the weight coefficient of the slave audio acquisition device to enhance the audio data collected by the main audio acquisition device and weaken the audio data collected by the slave audio acquisition device.

[0211] For example, as shown in FIG12 , the master audio capture device is referred to as the master device, and the slave audio capture devices are referred to as slave devices. After the master device exits the master mode, the master device enters normal mode 1210 . The master device can capture audio data 1220 and receive audio data 1230 (including device identifiers) collected by the slave devices via the network. The master device then mixes the audio data 1220 collected by the master device with the audio data 1231, ..., 1235 collected by each slave device 1240 and outputs the mixed data to the terminal device 1250 .

[0212] In addition, if the master audio acquisition device obtains audio data collected by a device, it transmits the audio data to the terminal device for playback; wherein the device is a slave audio acquisition device or a master audio acquisition device.

[0213] In an embodiment of the present application, after the main speaker device exits the main speaker mode, the normal mode is turned on, all devices are in the microphone-on state, and the user can use any device to speak. After the main audio acquisition device obtains the audio data collected by itself and the audio data collected by each slave audio acquisition device, it can mix the audio data and transmit it to the terminal device, thereby realizing the audio data transmission between the main audio acquisition device and each slave audio acquisition device in the normal mode.

[0214] Based on the same inventive concept, an embodiment of the present application also provides an audio processing method, which is executed by a slave audio acquisition device. The principle of solving the problem by this method is similar to that of the method on the master audio acquisition device side in the above embodiment. Therefore, the implementation of this method can refer to the implementation of the method on the master audio acquisition device side, and the repeated parts will not be repeated.

[0215] Referring to FIG. 13 , an audio processing method provided in an embodiment of the present application is executed by a first slave audio acquisition device among at least one slave audio acquisition device, for example, the slave audio acquisition devices 120-140 in FIG. 1 and the master audio acquisition device 220 in FIG. At least one slave audio acquisition device is connected to the master audio acquisition device. The specific implementation process of the method includes the following steps S131-S132:

[0216] S131. In response to a speaker mode start operation triggered for the first slave audio capture device, a speaker start request is sent to the master audio capture device, where the speaker start request includes a device identifier of the first slave audio capture device, so that the master audio capture device sets the first slave audio capture device as the speaker device and sets the device identifier of the first slave audio capture device to a speaker identifier.

[0217] The first slave audio collection device may be any slave audio collection device. After the master audio collection device sets the speaker identifier, it is saved.

[0218] S132. Collect first audio data and send the first audio data to a primary audio capture device, so that the primary audio capture device filters out the first audio data from multiple received audio data based on the speaker identifier, processes the multiple audio data, and transmits the processed audio data to a terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

[0219] Each slave audio collection device collects audio data and sends the audio data carrying the device identification to the master audio collection device, so that the master audio collection device determines any slave audio collection device as the master device based on the device identification.

[0220] In an embodiment of the present application, any slave audio acquisition device can trigger the speaker mode, the master audio acquisition device uses any slave audio acquisition device as the speaker device, and sets the device identifier of any slave audio acquisition device as the speaker identifier. The speaker device sends the collected audio data to the master audio acquisition device, so that the master audio acquisition device can enhance the audio data collected by the speaker device in the speaker mode, and / or suppress the audio data collected by the non-speaker device. In this way, even if the non-speaker makes a sound, the impact on the speaker's voice is relatively small, which can reduce the interference of other sounds on the speaker's voice and improve the effect of the sound collected by the cascaded devices, including clarity and volume. At the same time, the embodiment of the present application can set any device as the speaker device according to the location of the speaker, and people who need to speak can take turns to speak as the speaker, thereby improving the flexibility of the cascaded devices.

[0221] In some embodiments, in response to the speaker mode activation operation triggered on the first slave audio capture device in S131, sending a speaker activation request to the master audio capture device may include the following two situations:

[0222] Case 1: when the first slave audio acquisition device currently displays the microphone-on state, in response to the third speaker mode start operation triggered for the first slave audio acquisition device, the speaker start request is sent to the master audio acquisition device.

[0223] Case 2: when the first slave audio acquisition device currently displays a mute state, in response to a fourth speaker mode start operation triggered for the first slave audio acquisition device, the speaker start request is sent to the master audio acquisition device.

[0224] Among them, any slave audio acquisition device can display the microphone-on state or the microphone-off state, that is, any slave audio acquisition device can switch from the microphone-on state to the speaker mode, and can also switch from the microphone-off state to the speaker mode. The operation method of triggering the speaker mode in these two cases can be the same or different, that is, the above-mentioned third speaker mode activation operation and the fourth speaker mode activation operation can be the same or different. For details, please refer to the operation method of triggering the speaker mode in the above embodiment of this application.

[0225] In some embodiments, after any slave audio capture device triggers the speaker mode, the following steps may be performed:

[0226] In response to receiving a first notification message sent by the primary audio acquisition device, the microphone open state is displayed according to the speaker identifier carried in the first notification message; wherein the first notification message is used to indicate that the speaker mode is turned on.

[0227] In some embodiments, after any slave audio capture device turns on the speaker mode, when the speaker finishes speaking using the slave audio capture device, the slave audio capture device can be triggered to exit the speaker mode. The slave audio capture device can also perform the following steps:

[0228] In response to a speaker mode exit operation triggered for the first slave audio acquisition device, sending an exit speaker request to the master audio acquisition device, so that the master audio acquisition device sends a second notification message to the at least one slave audio acquisition device;

[0229] The second notification message is used to instruct other slave audio collection devices except the first slave audio collection device to switch from a closed microphone state to an open microphone state.

[0230] Based on the same inventive concept, an embodiment of the present application also provides an audio processing device. The principle of solving the problem by this device is similar to the method of the above embodiment. Therefore, the implementation of this device can refer to the implementation of the above method, and the repeated parts will not be repeated.

[0231] As shown in FIG14 , it is a schematic diagram of the structure of an audio processing device 1400. The device can be set in a master audio acquisition device, which is connected to at least one slave audio acquisition device. The device includes:

[0232] A setting unit 1401 is configured to, in response to a speaker mode activation operation triggered on a first device, set the first device as a speaker device and set a speaker identifier of the speaker device; wherein the first device is any one of the master audio acquisition device and the at least one slave audio acquisition device;

[0233] The screening unit 1402 is configured to obtain a plurality of audio data collected by the master audio collection device and the at least one slave audio collection device, and based on the speaker identifier, screen out the first audio data collected by the master audio collection device from the plurality of audio data;

[0234] The processing unit 1403 is configured to process the multiple audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, suppressing other audio data in the multiple audio data except the first audio data; and

[0235] The first transmission unit 1404 is configured to transmit the audio data to be transmitted to the terminal device for playback.

[0236] In one embodiment, the first device is a first slave audio acquisition device, and the setting unit 1401 is specifically configured to:

[0237] In response to receiving the speaker start request sent by the first slave audio acquisition device, setting the first slave audio acquisition device as the speaker device, and setting the device identifier carried in the speaker start request as the speaker identifier;

[0238] The speaker start request is sent by the first slave audio acquisition device in response to the speaker mode start operation triggered on the first slave audio acquisition device.

[0239] In one embodiment, in response to a speaker mode activation operation triggered for the primary audio acquisition device, the primary audio acquisition device is set as the speaker device, and the microphone open state of the primary audio acquisition device is set, the setting unit 1401 is specifically configured to:

[0240] When the primary audio acquisition device currently displays the microphone-on state, in response to a first speaker mode start operation triggered on the primary audio acquisition device, setting the primary audio acquisition device as the speaker device and maintaining the microphone-on state of the primary audio acquisition device;

[0241] When the primary audio acquisition device currently displays the closed-microphone state, in response to a second speaker mode activation operation triggered for the primary audio acquisition device, the primary audio acquisition device is set as the speaker device, and the closed-microphone state of the primary audio acquisition device is switched to the open-microphone state.

[0242] In one embodiment, the apparatus 1400 further includes:

[0243] The display unit 1405 is used to set and display the microphone muting status of the main audio acquisition device;

[0244] The first notification unit 1406 is configured to send a first notification message to the at least one slave audio collection device, wherein the first notification message carries the speaker identifier, and the first notification message is used to instruct any of the at least one slave audio collection device to turn on the speaker mode, so that

[0245] The first slave audio acquisition device displays a microphone-on status according to the speaker identifier;

[0246] The other slave audio collection devices other than the first slave audio collection device display the microphone-mute status according to the speaker identifier.

[0247] In one embodiment, the apparatus 1400 further includes:

[0248] The switching unit 1407 is configured to switch the microphone-off state of the master audio acquisition device to the microphone-on state in response to receiving the exit-speaker request sent by the first slave audio acquisition device;

[0249] The second notification unit 1408 is configured to send a second notification message to the at least one slave audio collection device, where the second notification message is used to instruct the other slave audio collection devices to switch from the mute microphone state to the open microphone state.

[0250] In one embodiment, the first device is the primary audio acquisition device, and the setting unit 1401 is specifically configured to:

[0251] In response to a speaker mode start operation triggered for the primary audio acquisition device, the primary audio acquisition device is set as the speaker device, and a microphone-opening state of the primary audio acquisition device is set.

[0252] In one embodiment, the apparatus 1400 further includes:

[0253] The third notification unit 1409 is configured to send a third notification message to the at least one slave audio collection device, where the third notification message is used to instruct the at least one slave audio collection device to display a microphone-mute state.

[0254] In one embodiment, the apparatus further comprises:

[0255] The fourth notification unit 1410 is configured to send a fourth notification message to the at least one slave audio capture device in response to a speaker mode exit operation triggered on the master audio capture device, wherein the fourth notification message is used to instruct the at least one slave audio capture device to switch from a mute microphone state to the open microphone state.

[0256] In one embodiment, the apparatus 1400 further includes:

[0257] The audio mixing unit 1411 is configured to, after exiting the speaker mode, obtain multiple audio data collected by the master audio collection device and the at least one slave audio collection device, and then mix the obtained multiple audio data to obtain mixed data;

[0258] The second transmission unit 1412 is configured to transmit the mixed audio data to the terminal device for playback.

[0259] In one embodiment, the processing unit 1403 is specifically configured to:

[0260] The other audio data are filtered out from the multiple audio data, and the first audio data is extracted as the audio data to be transmitted.

[0261] In one embodiment, the processing unit 1403 is specifically configured to:

[0262] In the process of mixing the multiple audio data, the first audio data is enhanced and the other audio data are suppressed to obtain the audio data to be transmitted.

[0263] In one embodiment, the processing unit 1403 is specifically configured to:

[0264] Based on the enhancement processing performed on the first audio data, the other audio data except the first audio data in the multiple audio data are suppressed, and the processed multiple audio data are mixed to obtain the audio data to be transmitted.

[0265] With the same inventive concept, an embodiment of the present application further provides an audio processing device, the principle of which is similar to the method from the audio acquisition device side of the above embodiment. Therefore, the implementation of the device can refer to the implementation of the above method from the audio acquisition device side, and the repeated parts will not be repeated.

[0266] 15 , an audio processing device 1500 is provided in an embodiment of the present application. The device can be provided in any of at least one slave audio acquisition device, where at least one slave audio acquisition device is connected to a master audio acquisition device. The device includes:

[0267] The first sending unit 1501 is configured to, in response to a speaker mode start operation triggered on the first slave audio acquisition device, send a speaker start request to the master audio acquisition device, wherein the speaker start request includes a device identifier of the first slave audio acquisition device, so that the master audio acquisition device sets the first slave audio acquisition device as the speaker device and sets the device identifier of the first slave audio acquisition device to a speaker identifier; and

[0268] The collection unit 1502 is used to collect first audio data and send the first audio data to the primary audio collection device, so that the primary audio collection device can filter out the first audio data from the multiple audio data received based on the main speaker identifier, process the multiple audio data, and transmit the processed audio data to the terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

[0269] In one embodiment, the first sending unit 1501 is specifically configured to:

[0270] When the first slave audio acquisition device currently displays a microphone-on state, in response to a third speaker mode start operation triggered on the first slave audio acquisition device, sending the speaker start request to the master audio acquisition device;

[0271] When the first slave audio acquisition device currently displays a mute state, in response to a fourth speaker mode start operation triggered on the first slave audio acquisition device, the speaker start request is sent to the master audio acquisition device.

[0272] In one embodiment, the apparatus 1500 further includes:

[0273] The receiving unit 1503 is configured to, in response to receiving a first notification message sent by the primary audio acquisition device, display a microphone-on status according to the speaker identifier carried in the first notification message; wherein the first notification message is used to indicate that the speaker mode is turned on.

[0274] In one embodiment, the apparatus 1500 further includes:

[0275] The second sending unit 1504 is configured to, in response to the speaker mode exit operation triggered on the first slave audio acquisition device, send an exit speaker request to the master audio acquisition device, so that the master audio acquisition device sends a second notification message to the at least one slave audio acquisition device;

[0276] The second notification message is used to instruct other slave audio collection devices except the first slave audio collection device to switch from a closed microphone state to an open microphone state.

[0277] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0278] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0279] After introducing the audio processing method and apparatus according to an exemplary embodiment of the present application, an electronic device according to another exemplary embodiment of the present application is introduced next.

[0280] Based on the same inventive concept as the above-mentioned method embodiment, an electronic device is also provided in an embodiment of the present application. This electronic device can be the master audio acquisition device or the slave audio acquisition device in the above-mentioned embodiment. In this embodiment, the structure of the electronic device can be as shown in Figure 16, including a memory 1601, a communication module 1603, a bus 1604, and one or more processors 1602.

[0281] Memory 1601 is used to store computer programs executed by processor 1602. Memory 1601 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.

[0282] Memory 1601 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1601 may be a combination of the aforementioned memories.

[0283] The processor 1602 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1602 is configured to implement the above-mentioned audio processing method when calling the computer program stored in the memory 1601 .

[0284] The communication module 1603 is used to communicate with terminal devices and other audio collection devices.

[0285] The specific connection medium between the memory 1601, communication module 1603, and processor 1602 is not limited in the embodiments of the present application. In Figure 16, the embodiment of the present application shows that the memory 1601 and the processor 1602 are connected via a bus 1604. The bus 1604 is depicted in bold in Figure 16. The connection methods between other components are merely schematic and are not intended to be limiting. The bus 1604 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 16 only uses a single bold line, but does not represent a single bus or a single type of bus.

[0286] The memory 1601 stores a computer storage medium, which stores computer executable instructions for implementing the audio processing method of the embodiment of the present application. The processor 1602 is used to execute the audio processing method of the above embodiment, as shown in Figure 5 or Figure 13.

[0287] In another embodiment, the electronic device may also be other electronic devices. In this embodiment, the structure of the electronic device may be as shown in FIG17 , including: a communication component 1710 , a memory 1720 , an audio circuit 1740 , a Bluetooth module 1750 , a processor 1730 and other components.

[0288] The communication component 1710 is used to communicate with other electronic devices. In some embodiments, it can include a WiFi module. The WiFi module is a short-range wireless transmission technology, and the electronic device can transmit data with other electronic devices through the WiFi module.

[0289] The memory 1720 can be used to store software programs and data. The processor 1730 executes various functions of the electronic device and processes data by running the software programs or data stored in the memory 1720. The memory 1720 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1720 stores an operating system that enables the electronic device to operate. In the present application, the memory 1720 can store the operating system and various application programs, and may also store a computer program that executes the audio processing method of the embodiment of the present application.

[0290] The audio circuit 1740 and the microphone 1741 can provide audio communication between the user and the electronic device. 口The microphone 1741 converts the collected sound signals into electrical signals, which are then received by the audio circuit 1740 and converted into audio data. The audio data is then output to the communication component 1710 for transmission to, for example, another electronic device, or to the memory 1720 for further processing.

[0291] The Bluetooth module 1750 is used to exchange information with other electronic devices having Bluetooth modules via the Bluetooth protocol. For example, an electronic device can establish a Bluetooth connection with another electronic device also having a Bluetooth module via the Bluetooth module 1750 to exchange data.

[0292] The processor 1730 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire terminal. By running or executing software programs stored in the memory 1720 and calling data stored in the memory 1720, it performs various functions of the terminal device and processes data. In some embodiments, the processor 1730 may include one or more processing units; the processor 1730 may also integrate an application processor and a baseband processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the baseband processor mainly processes wireless communications. It is understandable that the above-mentioned baseband processor may not be integrated into the processor 1730. In this application, the processor 1730 can run the operating system, application programs, user interface display and touch response, as well as the audio processing method of the embodiment of the application.

[0293] In some possible implementations, various aspects of the audio processing method provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the audio processing method according to various exemplary implementations of the present application described above in this specification. For example, the electronic device can execute the steps shown in Figure 5 or Figure 13.

[0294] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0295] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0296] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.

[0297] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0298] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The computer program can be executed entirely on the user electronic device, partially on the user electronic device, as a separate software package, partially on the user electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (for example, using an Internet service provider to connect through the Internet).

[0299] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0300] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0301] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain a computer-usable computer program.

[0302] The present application is described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program commands. These computer program commands can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the command executed by the processor of the computer or other programmable data processing device produces a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0303] These computer program commands may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the commands stored in the computer-readable memory produce a manufactured product including a command device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0304] These computer program commands can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the commands executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0305] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0306] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. An audio processing method, performed by a master audio acquisition device, the master audio acquisition device being connected to at least one slave audio acquisition device, the method comprising: In response to a speaker mode activation operation triggered on a first device, setting the first device as a speaker device and setting a speaker identifier of the speaker device; wherein the first device is any one of the master audio acquisition device and the at least one slave audio acquisition device; Acquire multiple audio data collected by the master audio collection device and the at least one slave audio collection device, and based on the speaker identifier, filter out first audio data collected by the master audio collection device from the multiple audio data; Processing the plurality of audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, suppressing other audio data in the plurality of audio data except the first audio data; and The audio data to be transmitted is transmitted to the terminal device for playback.

2. The method according to claim 1, wherein The first device is a first slave audio acquisition device, and in response to a speaker mode activation operation triggered on the first device, setting the first device as a speaker device and setting a speaker identifier of the speaker device includes: In response to receiving the speaker start request sent by the first slave audio acquisition device, setting the first slave audio acquisition device as the speaker device, and setting the device identifier carried in the speaker start request as the speaker identifier; The speaker start request is sent by the first slave audio acquisition device in response to the speaker mode start operation triggered on the first slave audio acquisition device.

3. The method according to claim 2, further comprising: Setting and displaying the microphone muting status of the primary audio acquisition device; A first notification message is sent to the at least one slave audio collection device, wherein the first notification message carries the speaker identifier, and the first notification message is used to instruct any of the at least one slave audio collection device to turn on the speaker mode, so that The first slave audio acquisition device displays a microphone-on status according to the speaker identifier; The other slave audio collection devices other than the first slave audio collection device display the microphone-mute status according to the speaker identifier.

4. The method according to claim 3, further comprising: In response to receiving the exit-speaker request sent by the first slave audio acquisition device, switching the closed-microphone state of the master audio acquisition device to the open-microphone state; A second notification message is sent to the at least one slave audio collection device, where the second notification message is used to instruct the other slave audio collection devices to switch from the microphone-mute state to the microphone-open state.

5. The method according to claim 1, wherein The first device is the main audio acquisition device, and in response to the speaker mode activation operation triggered on the first device, setting the first device as the speaker device and setting the speaker identifier of the speaker device includes: In response to a speaker mode start operation triggered for the primary audio acquisition device, the primary audio acquisition device is set as the speaker device, and a microphone-opening state of the primary audio acquisition device is set.

6. The method according to claim 5, further comprising: A third notification message is sent to the at least one slave audio collection device, where the third notification message is used to instruct the at least one slave audio collection device to display a microphone-mute state.

7. The method according to claim 5 or 6, wherein: The step of setting the primary audio acquisition device as the primary audio acquisition device and setting the microphone-on state of the primary audio acquisition device in response to a speaker mode activation operation triggered on the primary audio acquisition device includes: When the primary audio acquisition device currently displays the microphone-on state, in response to a first speaker mode start operation triggered on the primary audio acquisition device, setting the primary audio acquisition device as the speaker device and maintaining the microphone-on state of the primary audio acquisition device; When the primary audio acquisition device currently displays the closed-microphone state, in response to a second speaker mode activation operation triggered for the primary audio acquisition device, the primary audio acquisition device is set as the speaker device, and the closed-microphone state of the primary audio acquisition device is switched to the open-microphone state.

8. The method according to any one of claims 5 to 7, further comprising: In response to a speaker mode exit operation triggered for the master audio acquisition device, a fourth notification message is sent to the at least one slave audio acquisition device, where the fourth notification message is used to instruct the at least one slave audio acquisition device to switch from a closed microphone state to the open microphone state.

9. The method according to claim 4 or 8, further comprising: After exiting the speaker mode, when multiple audio data collected by the master audio collection device and the at least one slave audio collection device are obtained, mixing the obtained multiple audio data to obtain mixed data; The mixed audio data is transmitted to the terminal device for playback.

10. The method according to any one of claims 1 to 9, wherein The processing of the plurality of audio data to obtain the audio data to be transmitted includes: The other audio data are filtered out from the multiple audio data, and the first audio data is extracted as the audio data to be transmitted.

11. The method according to any one of claims 1 to 9, wherein The processing of the plurality of audio data to obtain the audio data to be transmitted includes: In the process of mixing the multiple audio data, the first audio data is enhanced and the other audio data are suppressed to obtain the audio data to be transmitted.

12. The method according to any one of claims 1 to 9, wherein The processing of the plurality of audio data to obtain the audio data to be transmitted includes: The first audio data is enhanced, the other audio data except the first audio data among the multiple audio data are suppressed, and the processed multiple audio data are mixed to obtain the audio data to be transmitted.

13. An audio processing method, performed by a first slave audio acquisition device among at least one slave audio acquisition device, the at least one slave audio acquisition device being connected to a master audio acquisition device, the method comprising: In response to a speaker mode activation operation triggered for the first slave audio acquisition device, sending a speaker activation request to the master audio acquisition device, wherein the speaker activation request includes a device identifier of the first slave audio acquisition device, so that the master audio acquisition device sets the first slave audio acquisition device as the speaker device and sets the device identifier of the first slave audio acquisition device to a speaker identifier; and, Collect first audio data and send the first audio data to the main audio acquisition device, so that the main audio acquisition device filters out the first audio data from the multiple audio data received based on the main speaker identifier, processes the multiple audio data, and transmits the processed audio data to the terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

14. The method according to claim 13, wherein: The step of sending a speaker start request to the master audio acquisition device in response to a speaker mode start operation triggered on the first slave audio acquisition device includes: When the first slave audio acquisition device currently displays a microphone-on state, in response to a third speaker mode start operation triggered on the first slave audio acquisition device, sending the speaker start request to the master audio acquisition device; When the first slave audio acquisition device currently displays a mute state, in response to a fourth speaker mode start operation triggered on the first slave audio acquisition device, the speaker start request is sent to the master audio acquisition device.

15. The method according to claim 13 or 14, further comprising: In response to receiving a first notification message sent by the primary audio acquisition device, the microphone open state is displayed according to the speaker identifier carried in the first notification message; wherein the first notification message is used to indicate that the speaker mode is turned on.

16. The method according to any one of claims 13 to 15, further comprising: In response to a speaker mode exit operation triggered for the first slave audio acquisition device, sending an exit speaker request to the master audio acquisition device, so that the master audio acquisition device sends a second notification message to the at least one slave audio acquisition device; The second notification message is used to instruct other slave audio collection devices except the first slave audio collection device to switch from a closed microphone state to an open microphone state.

17. An audio processing device, the device being connected to at least one secondary audio acquisition device in a cascade manner, the device comprising: A setting unit, configured to, in response to a speaker mode activation operation triggered on a first device, set the first device as a speaker device and set a speaker identifier of the speaker device; wherein the first device is any one of the master audio acquisition device and the at least one slave audio acquisition device; a screening unit, configured to obtain a plurality of audio data collected by the master audio collection device and the at least one slave audio collection device, and based on the speaker identifier, screen out the first audio data collected by the master audio collection device from the plurality of audio data; a processing unit, configured to process the plurality of audio data to obtain audio data to be transmitted; wherein the processing includes at least one of the following: enhancing the first audio data, suppressing other audio data in the plurality of audio data except the first audio data; and The first transmission unit is configured to transmit the audio data to be transmitted to a terminal device for playback.

18. An audio processing device, comprising at least one slave audio acquisition device including the device, the at least one slave audio acquisition device being connected to a master audio acquisition device, the device comprising: a first sending unit, configured to, in response to a speaker mode start operation triggered on the first slave audio acquisition device, send a speaker start request to the master audio acquisition device, wherein the speaker start request includes a device identifier of the first slave audio acquisition device, so that the master audio acquisition device sets the first slave audio acquisition device as the speaker device and sets the device identifier of the first slave audio acquisition device to a speaker identifier; and, The acquisition unit is configured to acquire first audio data and transmit the first audio data to the primary audio acquisition device, so that the primary audio acquisition device filters out the first audio data from the received multiple audio data based on the speaker identifier, processes the multiple audio data, and transmits the processed audio data to the terminal device for playback; wherein the processing includes at least one of the following: enhancing the first audio data and suppressing other audio data in the multiple audio data except the first audio data.

19. An electronic device comprising a processor and a memory, wherein: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 16.

20. A computer-readable storage medium comprising a computer program, wherein when the computer program is run on an electronic device, the computer program is configured to cause the electronic device to execute the steps of the method according to any one of claims 1 to 16.

21. A computer program product, comprising a computer program, wherein the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of the method according to any one of claims 1 to 16.