Microphone system, sound pickup panel, and audio processing method
By physically separating the pickup panel from the main unit, the pickup panel picks up and converts audio signals, while the main unit processes them, thus solving the sound quality problem of traditional all-in-one microphone systems in complex environments and achieving high-fidelity audio transmission and flexible layout.
Patent Information
- Application Number
- CN202610813731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-25
AI Technical Summary
Traditional all-in-one microphone systems suffer from poor sound quality in complex conference room environments. This is due to the physical contradiction between the coexistence of the pickup unit and the processing circuit, which leads to problems such as electromagnetic interference, noise interference, and pickup blind spots.
The microphone panel is physically separated from the host unit. The microphone panel picks up audio signals and converts them into standard network signals, which are then sent to the host unit for processing via a transmission cable. The host unit is responsible for complex algorithm processing, enabling flexible layout and high-fidelity audio transmission.
It improves the stability and fidelity of audio transmission, reduces blind spots in sound pickup, enhances the scalability and deployment flexibility of the microphone system, adapts to complex acoustic environments, and improves sound quality.
Smart Images

Figure CN122640657A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, specifically to a microphone system, a pickup panel, and an audio processing method. Background Technology
[0002] In modern remote collaboration and communication, conference rooms have become a core setting for high-quality audio interaction. Conference room audio systems need to be able to stably pick up the speech of each participant in complex indoor acoustic environments, effectively eliminate echoes and suppress noise, and ensure pure and natural sound to provide an immersive communication experience for local and remote participants.
[0003] In related technologies, integrated microphones typically integrate the pickup unit and signal processing circuitry into the same housing. While this highly integrated structure is compact, it is limited by the small internal space and the complexity of the electromagnetic environment, making it difficult to fully utilize the computing power of the audio processing chip and advanced algorithms (such as deep noise reduction, adaptive echo cancellation, and high-precision beamforming). At the same time, the close coexistence of the pickup unit with the circuitry and power module can easily introduce circuit noise, power supply ripple interference, and mechanical vibration coupling, thereby compromising the original purity of the audio signal. Furthermore, the fixed integrated design cannot flexibly optimize the position, angle, and directivity of the pickup unit according to the actual acoustic structure of the conference room (such as room size, reverberation time, and the location of background noise sources), which can easily lead to problems such as blind spots, insufficient far-field sensitivity, or over-collection of environmental reverberation.
[0004] The aforementioned factors collectively result in integrated microphones failing to meet the demands of high-fidelity audio interaction in complex conference room environments. Therefore, improving the sound quality performance of conference room audio systems in complex acoustic environments has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a microphone system designed to address the technical problem of poor sound quality in microphone systems of the related art.
[0006] To achieve the above objectives, a first aspect of this application provides a microphone system, including at least one pickup panel, a first host, and a transmission cable connecting the pickup panel and the first host, wherein the pickup panel is physically separated from the first host. At least one of the microphone panels is used to pick up audio signals and convert the audio signals into standard network signals before sending them to the first host via the transmission cable; The transmission cable is used to transmit the standard network signal from the microphone panel to the first host. The first host is used to receive and process the standard network signal to obtain an output signal.
[0007] In some embodiments, the first host includes a second processing device; The second processing device is used to send a clock synchronization signal to the pickup panel to synchronize the sampling clock of each pickup panel with the system clock of the first host, and to receive and process the standard network signal to obtain an output signal.
[0008] In some embodiments, the pickup panel includes: The audio acquisition module is used to pick up audio signals and convert them into digital signals; A network conversion module is used to convert the digital signal into a standard network signal.
[0009] In some embodiments, the audio acquisition module includes a microphone array and an analog-to-digital conversion module, and the network conversion module includes a first processing device and a first physical layer device; The microphone array is used to pick up the audio signal; The analog-to-digital converter module is electrically connected to the microphone array and is used to convert the audio signal into a first digital signal; The first processing device is used to package the first digital signal into a first encapsulation frame signal; and The first physical layer device is used to convert the first encapsulated frame signal into a standard network signal and send it to the first host via a transmission cable; The analog-to-digital conversion module is connected between the microphone array and the first processing device, and the first processing device is connected between the analog-to-digital conversion module and the first physical layer device.
[0010] In some embodiments, the first processing device includes: An audio serial interface acquisition controller is used to receive the first digital signal transmitted by the analog-to-digital conversion module; A media access controller, connected to the audio serial interface acquisition controller, is used to package the received first digital signal into a first encapsulated frame signal. A custom packet sending module is connected to the media access controller and is used to control the timing of the media access controller's encapsulation and the transmission of the first encapsulated frame signal; A precision time protocol program module is used to execute a precision time protocol to synchronize the clock of the pickup panel with the clock of the first host.
[0011] In some embodiments, the media access controller packages data according to a preset data packet length, the preset data packet length corresponding to the sampling period of the analog-to-digital conversion module.
[0012] Secondly, a pickup panel is provided, including: The shell thickness is no more than 10mm; A microphone array, disposed on the pickup surface of the housing, is used to pick up audio signals; A first processing unit is used to convert the audio signal into a standard network signal; and The first module interface is used to connect a transmission cable to transmit the standard network signal to an external host; wherein the microphone panel is physically separated from the external host and is connected to the external host via a transmission cable.
[0013] Thirdly, an audio processing method is provided, applied to the microphone system described above, the audio processing method comprising: Audio signals are picked up through a microphone panel and converted into standard network signals. The standard network signal is transmitted to the first host via a transmission cable; The first host receives and processes the standard network signal to obtain the output signal.
[0014] In some embodiments, audio signals from the same sound source are acquired synchronously through at least two pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The first host receives standard network signals sent by each microphone panel and extracts audio signals collected by each microphone panel. The first host calculates the time difference between the audio signals collected by each pickup panel, and determines the location information of the sound source based on the time difference.
[0015] In some embodiments, the audio signal of the target sound source is acquired synchronously through at least two pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The first host receives standard network signals sent by each microphone panel and extracts audio signals collected by each microphone panel. The first host dynamically adjusts the weighting coefficients and delay parameters of the audio signals collected by each pickup panel according to the location information of the target sound source, thereby forming a virtual beam pointing towards the target sound source.
[0016] Fourthly, an audio processing method is provided, applied to the microphone system described above, comprising: Audio signals are collected by multiple pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The first host receives standard network signals sent by each pickup panel, obtains the audio signal quality or sound source location of each pickup panel, dynamically selects a portion of the pickup panels to form an effective pickup array, and adjusts the signal weight of the selected pickup panels.
[0017] Fifthly, a computer storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implements the audio processing method described above.
[0018] In a sixth aspect, a control device is provided, on which a computer program or instructions are stored, which, when executed by a processor, implement the steps of the audio processing method described above.
[0019] In summary, the above technical solution offers several advantages. First, it allows for independent deployment of the pickup panel at the desired pickup location, while the primary host can be placed in another location, such as being hidden, thus adapting to complex environments. Second, through this physically separated setup, the pickup panel converts the picked-up audio signal into a standard network signal and transmits it to the primary host via a cable. The primary host then processes the standard network signal to obtain the output signal, effectively avoiding distortion and attenuation caused by interference during long-distance analog transmission, thereby improving the stability and fidelity of audio transmission. Third, the pickup panel is located away from potential interference sources such as the primary host's processing circuitry and power module, physically isolating the circuit noise from affecting the front-end analog signal and ensuring high fidelity and a high signal-to-noise ratio of the original audio signal. On the one hand, physically separating the pickup panel from the main unit allows for flexible adjustment of the number of pickup panels according to usage needs, thereby expanding the pickup range of the microphone system, meeting the pickup requirements of multiple scenarios and areas, and improving the scalability and deployment flexibility of the microphone system. On the other hand, it can flexibly optimize the position, angle, and directivity of the pickup unit according to the actual acoustic structure of the conference room (such as room size, reverberation time, and background noise source location), reducing pickup blind spots, improving far-field sensitivity, and avoiding reverberation in the acquisition environment, effectively improving the overall sound quality of the microphone system. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.
[0022] Figure 1 This is a schematic diagram of the uplink / downlink between the microphone system's pickup panel and the first host in an exemplary embodiment of this application; Figure 2 This is a schematic diagram showing the position of the microphone panel relative to the ceiling in an exemplary embodiment of the microphone system provided in this application; Figure 3 This is a schematic diagram of a microphone system provided in an exemplary embodiment of this application, in which multiple pickup panels are connected to a first host and a second host to form a star topology. Figure 4 This is a schematic diagram of the locking member of the installation component provided in an exemplary embodiment of this application being retracted; Figure 5 This is a schematic diagram of the locking structure of the mounting component provided in an exemplary embodiment of this application; Figure 6 yes Figure 5 Enlarged view of point A in the middle; Figure 7 This is a schematic diagram of the structure of the mounting components provided in an exemplary embodiment of this application, which are retracted to pass through the guide channel of the pickup panel; Figure 8 This is a schematic diagram of the structure of the pickup panel provided in the exemplary embodiment of this application, which is attached to the ceiling by the mounting components; Figure 9 This is a schematic diagram of the overall structural architecture of the microphone system provided in an exemplary embodiment of this application; Figure 10 This is a schematic diagram of the overall structural architecture of the pickup panel provided in an exemplary embodiment of this application; Figure 11 This is a schematic diagram of the overall structural architecture of the first host, including the switch, provided in an exemplary embodiment of this application; Figure 12 This is yet another schematic diagram of the overall structural architecture of the microphone system provided in the exemplary embodiments of this application; Figure 13 This is yet another schematic diagram of the overall structural architecture of the microphone system provided in the exemplary embodiments of this application; Figure 14 This is a flowchart illustrating the audio processing method provided in an exemplary embodiment of this application; Figure 15 This is another flowchart illustrating the audio processing method provided in an exemplary embodiment of this application.
[0023] Explanation of reference numerals in the attached figures: 10. Microphone system; 11. Pickup panel; 11a. Audio acquisition module; 111. Microphone array; 112. Analog-to-digital conversion module; 11b. Network conversion module; 113. First processing unit; 1131. Audio serial interface acquisition controller; 1132. Media access controller; 1133. Custom packet sending program module; 1134. Precision time protocol program module; 114. First physical layer device; 115. Speaker unit; 116. Signal switching module; 117. Audio mixing module Block; 12, First host; 121, Second physical layer device; 122, Second processing device; 1221, First processor; 123, Second processor; 124, Switch; 13, Transmission cable; 14, Second host; 15, Mounting assembly; 151, Mounting base; 152, Locking element; 153, Reset element; 20, Ceiling; 20a, Guide channel; M1, Output port; M11, First output port; M12, Second output port; M2, First module interface; M3, Second module interface. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] In the following description, specific embodiments of this application will be illustrated with reference to steps and symbols performed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be referred to several times as being performed by a computer. Computer performance as referred to in this application includes operations performed by a computer processing unit on electronic signals represented by data in a structured format. This operation transforms the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise alter the operation of the computer in a manner well known to those skilled in the art. The data structure maintained by the data is the physical location of the memory, which has specific characteristics defined by the data format. However, the principles of this application are illustrated with specific embodiments and are not intended to be limiting. Those skilled in the art will understand that many of the steps and operations described below can also be implemented in hardware.
[0026] The terms "module" or "unit" as used in this application can be considered as software objects executing on the computing system. The different components, modules, engines, and services described in this application can be considered as implementation objects on the computing system. While the apparatus and methods described in this application are preferably implemented in software, they can also be implemented in hardware, both of which are within the scope of protection of this invention.
[0027] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used in the embodiments of this application may also include the plural forms. It should be further understood that the term “comprising” as used in the specification of this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when an element is “connected” or “coupled” to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein may include wireless connection or wireless coupling. The term “and / or” as used herein includes all or any unit and all combinations of one or more associated listed items.
[0028] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0029] Based on the technical issues mentioned in the background, conference rooms have become a core scenario for high-quality audio interaction in modern remote collaboration and communication. Conference room audio systems need to be able to stably capture the speech of each participant in complex indoor acoustic environments, effectively eliminate echoes and suppress noise, ensuring pure and natural sound, and providing an immersive communication experience for both local and remote participants.
[0030] In the field of conference audio equipment, traditional all-in-one microphone designs have long been considered the natural choice for meeting the sound pickup needs of conference rooms due to their compact structure and ease of deployment. The industry generally agrees that the challenges in improving sound quality primarily stem from the performance of the microphone unit itself, the precision of the algorithm processing, or achieving better circuit layout and shielding within limited space. Therefore, common technological improvement paths focus on selecting more sensitive microphone chips, integrating more powerful processors, or optimizing audio algorithms within a given all-in-one housing, attempting to achieve breakthroughs in sound quality under the constraints of high integration.
[0031] However, through in-depth analysis and practice, the inventors of this patent recognized a fundamental, yet long-overlooked, premise in the aforementioned common technical approaches: that "sound pickup" and "processing" must share the same physical space. It is this seemingly self-evident design constraint, rather than a specific component or algorithm, that constitutes the core root of the sound quality bottleneck. The inventors observed that within the integrated structure, the electromagnetic radiation generated by the high-performance audio processing chip, the ripple noise of the power supply circuit, and the extremely sensitive analog pickup unit create an irreconcilable physical contradiction. This contradiction cannot be completely resolved by improving a single component, as it stems from the inherent flaw at the system architecture level where "strong interference sources" and "weak signal sources" are forced to coexist closely.
[0032] The discovery process of this problem was further clarified and deepened in specific application scenarios. Taking ceiling-mounted installation as an example, traditional integrated ceiling microphones are often bulky in order to accommodate the processing circuitry, imposing strict requirements on installation space and ceiling height. The industry usually regards this as an unavoidable trade-off and focuses its research and development on how to make the circuit board smaller and the shielding better. However, the inventors broke out of this mindset and realized that "large size" and "limited sound quality" are not two independent problems, but two external manifestations of the same architectural problem: the root cause of the large size is to accommodate the processing circuitry, and the root cause of the limited sound quality is that these circuits cause near-field interference to the pickup unit. Therefore, the real technical challenge is not how to "integrate better", but how to "separate reasonably".
[0033] In view of this, a microphone system 10 that can be flexibly arranged and effectively improves audio acquisition quality and transmission stability is provided, aiming to solve the space and sound quality limitations caused by the integrated microphone system 10. This application overcomes the problem of unbalanced space and sound quality caused by the integrated microphone system 10 by physically separating the sound pickup function from the first host 12 in terms of physical structure.
[0034] Figure 1 This is a schematic diagram of the uplink / downlink between the microphone system 10 pickup panel 11 and the first host 12 in an exemplary embodiment of this application. Figure 2 This is a schematic diagram showing the position of the microphone system 10 pickup panel 11 relative to the ceiling 20 in an exemplary embodiment of this application.
[0035] The first aspect of this application, with reference to Figure 1 and Figure 2A microphone system 10 is provided. The microphone system 10 includes at least one pickup panel 11, a first host 12, and a transmission cable 13 connecting the pickup panel 11 and the first host 12. In this embodiment, the pickup panel 11 and the first host 12 are physically separated. The at least one pickup panel 11 is used to pick up audio signals, convert the audio signals into standard network signals, and then transmit them to the first host 12 via the transmission cable 13. The transmission cable 13 is used to transmit the standard network signals from the pickup panel 11 to the first host 12. The first host 12 is used to receive and process the standard network signals to obtain an output signal.
[0036] It should be noted that, in this embodiment, the microphone panel 11 refers to the front-end part of the device used to pick up audio signals and convert them into standard network signals. Its main function is to receive audio signals from the environment and convert them into standard network signals. It is understood that the microphone panel 11 is only responsible for performing analog-to-digital conversion on the audio signals collected by the microphone array 111, encapsulating them into data packets according to standard network protocols (such as Ethernet), and sending them to the first host 12 via the transmission cable 13. The microphone panel 11 does not perform any substantial processing on the audio signals themselves, nor does it change the original acoustic characteristics of the audio signals. That is, the standard network signal received by the first host 12 carries unprocessed raw audio data, preserving complete sound field information (including the relative time difference, amplitude difference, and phase difference between each microphone channel), providing the highest quality basic data for high-precision sound source localization, adaptive beamforming, and three-dimensional spatial audio rendering. Compared to traditional all-in-one microphones, which often integrate dozens or hundreds of microphone units, an internal DSP processor, and complex algorithms (including those for beamforming, noise suppression, automatic mixing, and sound source localization), once the audio stream processed by the internal DSP is output, no subsequent device can obtain the original audio data of the microphone array 111, thus preventing secondary localization or more advanced spatial audio rendering. In contrast, the pickup panel 11 of this application only outputs the original multi-channel audio data of the picked-up audio. It does not contain a digital signal processor (DSP) or algorithm module for performing mixing, noise reduction, sound source localization, beamforming, or echo cancellation on the picked-up audio. The first host 12 can dynamically select to use simple automatic mixing, run complex audio algorithms, or even forward the original audio data to a third-party processing platform, demonstrating great application flexibility.
[0037] In some implementations, the pickup panel 11 may be designed to have a small volume and thickness to facilitate mounting on various surfaces, such as the ceiling 20 or a desktop.
[0038] The first host 12 refers to the back-end device used to receive and process standard network signals from the microphone panel 11. The first host 12 typically integrates a high-performance processor and various interfaces, enabling it to perform complex algorithmic processing on the received signals and manage the operation of the entire system. The first host 12 can be deployed in a location physically separate from the microphone panel 11, such as above a low-voltage box, cabinet, or ceiling 20.
[0039] The transmission cable 13 refers to the physical medium that connects the pickup panel 11 and the first host 12 to transmit data and / or power. Exemplarily, the transmission cable 13 can be of various types, such as Ethernet cable, fiber optic cable or other specialized cable. The configuration of the transmission cable 13 is usually adapted according to the required transmission distance, bandwidth and anti-interference capability, which will not be elaborated in the embodiments of this application.
[0040] Audio signals refer to the electrical signals corresponding to the sound information acquired through the pickup panel 11. Standard network signals refer to data signals that conform to a specific network communication protocol (such as the Ethernet protocol). This signal format facilitates transmission and processing in a network environment and can utilize existing network infrastructure.
[0041] The output signal refers to the audio signal that, after being processed by the first host 12, can be used by other devices (such as power amplifiers, recording equipment, audio processors, audio codecs, or video conferencing systems). The output signal can be an analog signal, a digital signal, or a network audio stream. Figure 3 Taking the output signal being processed by the first host 12 and then transmitted to the second host 14 as an example.
[0042] Specifically, the microphone system 10 in this embodiment includes at least one pickup panel 11, a first host 12, and a transmission cable 13 connecting the pickup panel 11 and the first host 12. The pickup panel 11 can be one or more, used to collect sound in different areas. The first host 12, as the core processing unit of the system, is responsible for centrally processing the standard network signals transmitted from the pickup panel 11. The transmission cable 13 serves as the data and / or power transmission channel between the two.
[0043] The microphone panel 11 and the first host unit 12 are physically separated. This separate design allows the microphone panel 11 to be independently optimized in terms of structure and size, such as by designing it as an ultra-thin form, to facilitate installation in various confined spaces. Simultaneously, the first host unit 12 can be placed in a concealed location away from the microphone panel 11, thus avoiding any impact on the room's aesthetics and facilitating maintenance and heat dissipation. For example, the microphone panel 11 can be installed on the ceiling 20 surface of the conference room, while the first host unit 12 can be placed above the ceiling 20 inside the conference room or in an equipment room or low-voltage distribution box outside the conference room.
[0044] The primary host 12 is physically separated and placed in a concealed location next to the conference room's electrical control box or above the suspended ceiling 20. This separation allows the primary host 12 to have a larger size and stronger processing power, while avoiding any impact on the conference room's interior space and aesthetics. For example, the primary host 12 can integrate a high-performance digital signal processor for executing complex audio processing algorithms.
[0045] For example, during a meeting, the microphone panel 11 is used to pick up audio signals within the meeting room. These audio signals are converted into standard network signals internally within the microphone panel 11. For instance, the microphone panel 11 can capture the speaker's voice in real time and convert it into data packets conforming to the Ethernet protocol. Subsequently, these standard network signals are transmitted to the first host 12 via transmission cable 13. Transmission cable 13 can be a standard Ethernet cable, which not only transmits data but also provides power to the microphone panel 11 via Power over Ethernet (PoE) technology, thus simplifying cabling. For example, a single Cat5e Ethernet cable can fulfill the data transmission and power supply requirements of the microphone panel 11. Then, after receiving the standard network signals from transmission cable 13, the first host 12 processes the standard network signals to obtain an output signal. Finally, the first host 12 sends the processed audio signal as an output signal to the meeting room's amplifier, recording equipment, or video conferencing terminal for subsequent use.
[0046] Figure 3 This is a schematic diagram of a microphone system 10 provided in an exemplary embodiment of this application, in which multiple pickup panels 11 are connected to a first host 12 and a second host 14 to form a star topology.
[0047] It should be understood and referenced. Figure 3 In this example, there can be multiple microphone panels 11, which can be distributed in different locations, such as different locations on the ceiling 20, and each can be connected to the first host 12 through its own transmission cable 13. In this way, on the one hand, the complexity of wiring can be reduced; on the other hand, a single first host 12 can integrate and process the audio signals of multiple microphone panels 11 to cover a wider pickup area and enhance the uniformity of sound field perception. When the speaker moves from the coverage area of one microphone panel 11 to the coverage area of another microphone panel 11, the first host 12 can perform seamless and high-precision sound source localization across the entire macro array range based on the time difference of arrival (TDOA) of the audio signals received by all microphone panels 11. This enables the microphone system 10 to continuously track moving sound sources and avoid pickup blind spots.
[0048] In an alternative embodiment, the microphone system 10 in this example can be implemented as a ceiling-mounted microphone 20 to pick up audio signals. That is, the pickup panel 11 is attached to the side of the ceiling 20 facing the ambient space, while the first host 12 is located on the side of the ceiling 20 away from the ambient space. In this case, only a small hole needs to be made in the ceiling 20 for the transmission cable 13 to pass through, and the microphone system 10 can be installed by passing the transmission cable 13 through the small hole. In this way, damage to the structure of the ceiling 20 can be greatly reduced, while improving the overall aesthetics.
[0049] In an alternative embodiment, the microphone system 10 in this example also takes the form of a ceiling-mounted microphone 20. Figures 4 to 8 A structural schematic diagram of the pickup panel 11, mounting assembly 15, and ceiling 20 in this example is shown. In this example, reference... Figure 7 A guide channel 20a can be opened on the ceiling 20, and the pickup panel 11 is detachably connected to the ceiling 20 via mounting assembly 15. Please continue to refer to the figure. Figures 4 to 6 In this example, the mounting assembly 15 includes a mounting base 151, a reset element 153, and a locking element 152, with a guide channel 20a formed in the ceiling 20. In this example, the mounting base 151 is connected to the pickup panel 11, and the reset element 153 and the locking element 152 are coaxially mounted on the mounting base 151. Please refer to [reference needed]. Figure 4 , Figure 7 and Figure 8 During installation, the pickup panel 11 with mounting assembly 15 is lifted to the guide channel 20a of the ceiling 20, and the end of the locking member 152 slides into the inner wall of the guide channel 20a. During the sliding process, the locking member 152 is compressed by the inner wall of the channel and undergoes elastic deformation, retracting inward. After the mounting base 151 passes through the guide channel 20a, the locking member 152 springs outward under the elastic force of the reset member 153, and the reset member 153 abuts against the ceiling 20. At this time, the upper surface of the pickup panel 11 is in close contact with the lower surface of the ceiling 20.
[0050] If maintenance or replacement of the microphone panel 11 is required, simply press the locking member 152 inward from both sides or a specific position of the microphone panel 11 to overcome the resistance of the reset member 153 and retract inward, thereby freeing it from the restriction of the guide channel 20a of the ceiling 20. Then the microphone panel 11 can be moved to separate it from the ceiling 20.
[0051] In this example, the locking element 152 is constructed as a plate-shaped reset plate, which increases the contact area with the ceiling 20. The reset element 153 is constructed as a torsion spring, with one of its two torsion arms abutting against the locking element 152 and the other against the pickup panel 11. The mounting base 151 is constructed as a columnar frame. The rotation axes of the torsion spring and the reset plate are coaxially arranged and both are mounted on the columnar frame. In an optional embodiment, the microphone system 10 in this example also adopts the form of a ceiling 20 microphone. Multiple pickup panels 11 can be distributed in different areas within the ceiling 20. The first host 12 can be placed in other easily concealed locations such as a cabinet or a low-voltage box. By having the distributed pickup panels 11 work in conjunction with the concealed first host 12, the installation of the microphone system 10 can be completed without damaging the structure of the ceiling 20, eliminating the dependence on the internal installation space of the ceiling 20 and making the deployment of the microphone system 10 more flexible. Through the above technical solution, on the one hand, the pickup panel 11 can be independently deployed at the required pickup location according to the usage scenario, while the first host 12 can be placed in another location, such as being hidden, thus achieving a layout that adapts to complex environments. On the other hand, through physically separated setup, the pickup panel 11 converts the picked-up audio signal into a standard network signal and then transmits it to the first host 12 via the transmission cable 13. The first host 12 then processes the standard network signal to obtain the output signal, effectively avoiding distortion and attenuation problems caused by interference in long-distance analog transmission of audio signals, thereby improving the stability and fidelity of audio transmission. Furthermore, the pickup panel 11 is far away from potential interference sources such as the processing circuit and power module of the first host 12, physically isolating the influence of circuit noise on the front-end analog signal, ensuring high fidelity and high signal-to-noise ratio of the original audio signal. On the other hand, by physically separating the pickup panel 11 from the first host 12, the number of pickup panels 11 can be flexibly set according to usage needs, thereby expanding the pickup range of the microphone system 10, meeting the pickup needs of multiple scenarios and multiple areas, and improving the scalability and deployment flexibility of the microphone system 10. On the other hand, the position, angle and directivity of the pickup unit can be flexibly optimized according to the actual acoustic structure of the conference room (such as room size, reverberation time, and background noise source location), reducing pickup blind spots, improving far-field sensitivity, and avoiding reverberation in the acquisition environment, effectively improving the overall sound quality of the microphone system 10.
[0052] Figure 9 This is a schematic diagram of the overall structural architecture of the microphone system 10 provided in an exemplary embodiment of this application.
[0053] To improve system reliability, in some embodiments, please refer to Figure 9The first host 12 includes a second processing device 122. In this embodiment, the second processing device 122 is used to send a clock synchronization signal to the pickup panel 11 so that the sampling clock of each pickup panel 11 is synchronized with the system clock of the first host 12, and to receive and process standard network signals to obtain an output signal.
[0054] It should be understood that the second processing device 122 in this embodiment is the core processing unit within the first host 12, and its main function is to coordinate the time base of the entire microphone system 10 and perform audio signal processing. The second processing device 122 can be specifically implemented as a processor capable of processing a large number of audio data streams in parallel. Specific examples of the processor included in the second processing device 122 are described below and will not be repeated here.
[0055] Through the above technical solution, in the microphone system 10, the pickup panel 11 is physically separated from the first host 12 and connected via a transmission cable 13. The pickup panel 11 is used to pick up audio signals and convert them into standard network signals, which are then sent to the first host 12 via the transmission cable 13. The first host 12 is used to receive and process these standard network signals. To solve the problem of asynchronous sampling clocks that may occur when multiple pickup panels 11 work independently, this application further proposes that a second processing device 122 be installed inside the first host 12. This second processing device 122 serves as the core clock source and signal processing center of the entire system. Its working principle is as follows: First, the second processing device 122 can actively send clock synchronization signals to all connected pickup panels 11. These clock synchronization signals are transmitted to each pickup panel 11 via the transmission cable 13. After receiving the signals, the pickup panel 11 adjusts its own clock signal according to the received clock synchronization signals, so that the clocks of each pickup panel 11 and the microphone system 10 composed of the first host 12 are synchronized. Preferably, the pickup panel 11 adjusts its sampling clock according to the received clock synchronization signal to keep it synchronized with the system clock of the first host 12. Therefore, regardless of the number of pickup panels 11, they will all sample audio at the same time, ensuring that all acquired audio data is aligned on the timeline. Subsequently, the second processing unit 122 receives time-synchronized standard network signals from each pickup panel 11. Since these signals are synchronized during acquisition, the second processing unit 122 can directly process the time-aligned standard network signals to obtain a high-quality output signal.
[0056] To avoid redundancy in the internal circuitry and chaotic signal processing of the microphone panel 11 due to unclear audio signal conversion efficiency and structural division of labor, thereby increasing panel thickness and complexity, in some embodiments, please refer to... Figure 9The microphone panel 11 includes an audio acquisition module 11a and a network conversion module 11b. The audio acquisition module 11a is used to acquire audio signals and convert them into digital signals. The network conversion module 11b is used to convert digital signals into standard network signals.
[0057] For example, the main function of the audio acquisition module 11a is to convert the sound wave signals in the environment into digital signals that can be processed by the electronic system. The network conversion module 11b is responsible for encapsulating and modulating the digital signals output by the audio acquisition module 11a, converting them into standard network signals that conform to standard network protocols, so that they can be transmitted through the transmission cable 13.
[0058] By clearly dividing the functions of the pickup panel 11 into an audio acquisition module 11a and a network conversion module 11b, the signal processing flow can be modularized and professionally divided. Specifically, the audio acquisition module 11a is used to pick up audio signals and convert them into digital signals. This process avoids interference that analog signals may encounter during transmission and provides data packets for subsequent digital signal processing. The network conversion module 11b can receive the digital signals transmitted from the audio acquisition module 11a and convert them into standard network signals for transmission to the first host 12 via the transmission cable 13. With this technical solution, the clear division of functions among the modules within the pickup panel 11 allows each module to be independently optimized. On the one hand, it simplifies the signal processing path, reducing redundant components and wiring complexity on the circuit board, thereby effectively reducing the overall thickness and internal space occupied by the pickup panel 11, making it easier to achieve an ultra-thin design. On the other hand, the modular design also improves the reliability and consistency of signal conversion, ensuring that the entire process from audio pickup to network transmission is efficient and stable.
[0059] In the process of audio acquisition module 11a and network conversion module 11b picking up and converting audio signals, there is insufficient efficiency and reliability in audio signal conversion, which may lead to increased latency, degraded signal quality, and the risk of data loss during transmission, thereby affecting the real-time performance and stability of the entire microphone system 10. Therefore, in some embodiments, please refer to... Figure 9The audio acquisition module 11a includes a microphone array 111 and an analog-to-digital converter (ADC) module 112, while the network conversion module 11b includes a first processing unit 113 and a first physical layer device 114. The microphone array 111 of the audio acquisition module 11a is used to pick up audio signals; the ADC module 112 of the audio acquisition module 11a is electrically connected to the microphone array 111 and is used to convert the audio signals into first digital signals. The first processing unit 113 of the network conversion module 11b is used to package the first digital signals into first encapsulated frame signals, and the first physical layer device 114 of the network conversion module 11b is used to convert the first encapsulated frame signals into standard network signals and transmit them to the first host 12 via a transmission cable.
[0060] The analog-to-digital conversion module 112 is connected between the microphone array 111 and the first processing device 113, and the first processing device 113 is connected between the analog-to-digital conversion module 112 and the first physical layer device 114.
[0061] For example, the audio acquisition module 11a includes a microphone array 111, which is a collection of sensors for picking up audio signals. Each microphone array 111 can be composed of multiple independent miniature microphone units arranged in a specific geometry, such as a linear array, a circular array, or a planar array. The multiple independent microphone units included in the microphone array 111 work together to achieve a wider pickup range, higher sensitivity, and the ability to sense the direction of the sound source. As another implementation, the microphone array 111 can also be a MEMS (Micro-Electro-Mechanical Systems) microphone array 111 integrated on a single chip to achieve a smaller size and higher integration. The audio acquisition module 11a includes an analog-to-digital converter module 112 for converting the audio signals picked up by the microphone array 111 into digital signals. This analog-to-digital converter module 112 typically includes one or more analog-to-digital converters (ADCs), which convert continuously varying analog voltage signals into discrete digital bitstreams according to a certain sampling rate and quantization precision. For example, a high-precision Sigma-Delta ADC chip or an ADC function module integrated into an audio codec can be used.
[0062] The network conversion module 11b includes a first processing unit 113 for receiving the first digital signal output by the analog-to-digital conversion module 112 and packaging it into a first encapsulated frame signal. Exemplarily, the first processing unit 113 in this embodiment can be a system-on-a-chip (SoC), which integrates a processing core, memory, and various peripheral interfaces, enabling data packaging functionality through hardware logic. The network conversion module 11b includes a first physical layer device 114 for converting the first encapsulated frame signal into a standard network signal and transmitting it to the first host 12 via a transmission cable. As one feasible implementation, the first physical layer device 114 can be an Ethernet physical layer transceiver (PHY chip), responsible for implementing physical layer data encoding, decoding, modulation, demodulation, and media access control. For example, an Ethernet PHY chip conforming to the IEEE 802.3 standard can be used, supporting transmission rates of 100Mbps or 1Gbps.
[0063] The technical solution of this application embodiment refines the audio acquisition module 11a and the network conversion module 11b, and clarifies their internal components and connections, constructing an efficient and reliable audio signal processing link. Specifically, the microphone array 111 is responsible for high-fidelity pickup of the original audio signal, ensuring the quality of the input signal. Subsequently, the analog-to-digital conversion module 112 is closely connected to the microphone array 111, rapidly converting the analog audio signal into a first digital signal. This early digitization process effectively avoids the problem of analog signals being susceptible to interference and distortion during transmission, thereby improving the purity of the signal. Next, after receiving the first digital signal, the first processing device 113 efficiently packages it to form a first encapsulated frame signal. This packaging process can be optimized according to network transmission characteristics to reduce data transmission latency. Finally, the first physical layer device 114 converts the first encapsulated frame signal into a standard network signal and sends it to the first host 12 through the transmission cable 13. This standardized network signal transmission method ensures the stability and compatibility of data transmission. Throughout the process, the connections between the analog-to-digital converter module 112, the microphone array 111, and the first processing device 113, as well as the connection between the first processing device 113 and the first physical layer device 114, together form a seamless signal processing path, minimizing signal loss and delay between each stage. This structured design enables the pickup panel 11 to efficiently and reliably complete the acquisition, digitization, packaging, and network transmission of audio signals, providing high-quality input signals to the first host 12, thereby ensuring that the entire microphone system 10 still achieves excellent real-time performance and stability even under a physically separated architecture.
[0064] The following is a concrete example: the microphone array 111 can consist of 128 MEMS microphone units arranged in a circular array to achieve 360-degree omnidirectional sound pickup. The analog-to-digital converter (ADC) module 112 can be a multi-channel audio ADC chip, which is connected to the first processing device 113 via an I2S interface, converting the analog signal output from the microphone array 111 into a first digital signal at a 48kHz sampling rate and 24-bit precision. The first processing device 113 can be a SoC chip, capable of receiving the first digital signal output from the ADC module 112 and packaging it into a first encapsulated frame signal. The first physical layer device 114 can be a Gigabit Ethernet PHY chip conforming to the IEEE 802.3ab standard, capable of converting the first encapsulated frame signal into a standard network signal and transmitting it to the first host 12 via a transmission cable 13 connected to an RJ45 interface. In this configuration, the ADC module 112 is electrically connected to the first processing device 113, and the first processing device 113 is connected to the first physical layer device 114, forming a high-efficiency, low-latency audio data link.
[0065] Figure 10 This is a schematic diagram of the overall structural architecture of the pickup panel 11 provided in an exemplary embodiment of this application.
[0066] To ensure real-time transmission and synchronization of audio signals, in some embodiments, please refer to... Figure 9 and Figure 10 The first processing device 113 includes an audio serial interface acquisition controller 1131, a media access controller 1132, a custom packet sending program module 1133, and a precision time protocol program module 1134. The audio serial interface acquisition controller 1131 receives a first digital signal from the analog-to-digital converter module 112. The media access controller 1132 is connected to the audio serial interface acquisition controller 1131 and is used to package the received first digital signal into a first encapsulated frame signal. The custom packet sending program module 1133 is connected to the media access controller 1132 and is used to control the encapsulation and transmission timing of the first encapsulated frame signal by the media access controller 1132. The precision time protocol program module 1134 executes a precision time protocol to synchronize the clock of the pickup panel 11 with the clock of the first host 12.
[0067] Specifically, the Serial Audio Interface controller 1131 (SAI controller) refers to a hardware or software module dedicated to processing audio data streams. Its main function is to receive the first digital signal from the analog-to-digital converter module 112. This SAI controller 1131 can be a dedicated peripheral interface integrated into a system-on-a-chip (SoC), such as an I2S (Inter-IC Sound) interface controller, ensuring that the digital audio signal can be efficiently and losslessly transmitted from the analog-to-digital converter module 112 to the first processing device 113 for corresponding processing operations in the correct format and timing.
[0068] Media access controller 1132 refers to a hardware or software module responsible for data link layer functions. Its main function is to package the received first digital signal into a first encapsulated frame signal according to a specific protocol (such as Ethernet protocol). The media access controller 1132 can be a standard Ethernet media access controller 1132 (MAC controller), or it can be a MAC controller customized for a specific network protocol. The specific choice depends on the actual needs, and should be able to achieve frame assembly, address management, error detection, and flow control to ensure that digital audio data can be correctly encapsulated into network-transmittable data packets.
[0069] The custom packet sending module 1133 refers to a software or firmware module that controls the encapsulation behavior of the media access controller 1132 and the timing of the transmission of the first encapsulated frame signal according to a preset strategy. This custom packet sending module 1133 can be an application running on a microcontroller or embedded processor, or a state machine integrated into the hardware logic. Through this custom packet sending module 1133, the size of the data packets, the transmission frequency, and the transmission interval can be flexibly adjusted to optimize network transmission latency and bandwidth utilization. For example, the number of audio sample points contained in each data packet can be precisely controlled according to the audio sampling rate and the required real-time performance, thereby achieving extremely low transmission latency.
[0070] The Precision Time Protocol (PTP) module 1134 refers to a software or hardware module that implements high-precision clock synchronization. It executes the Precision Time Protocol (PTP, IEEE 1588), for example, to align all microphone arrays 111 to the clock of the first host 12 via a network (with the first host 12 as the master clock and the microphone panels 11 as slave clocks; the PTP algorithm calculates the time difference of the timestamps for clock calibration of the microphone panels 11). Exemplarily, the Precision Time Protocol module 1134 can be a PTP daemon running on a System-on-Chip (SOC) to ensure that the internal clock of the microphone panels 11 is highly synchronized with the system clock of the first host 12, typically achieving nanosecond (ns) level accuracy. Through the PTP protocol, the microphone panels 11 can receive clock synchronization messages from the first host 12 via the transmission cable 13 and adjust their own clocks according to the timestamp information in the messages, thereby ensuring that the audio data collected by all microphone panels 11 are strictly aligned on the timeline.
[0071] In one specific implementation, the first processing device 113 can be integrated into a low-power embedded processor, such as a microcontroller based on the ARM Cortex-M series. In this microcontroller, the audio serial interface acquisition controller 1131 can be a built-in I2S peripheral for receiving a digital audio stream of the first digital signal from the analog-to-digital converter module 112 (e.g., a multi-channel PCM186x series ADC chip). The media access controller 1132 can be a built-in Ethernet MAC controller responsible for packaging the first digital signal acquired by the I2S interface into a first encapsulated frame signal (Ethernet frame). The custom packet sending module 1133 can be a firmware program running on the microcontroller that calculates the number of audio sampling points contained in each first encapsulated frame signal (Ethernet frame) based on a preset audio sampling rate (e.g., 48kHz) and a desired delay (e.g., 20 microseconds) (e.g., 512 bytes of data for a 48kHz sampling rate), and controls the MAC controller to encapsulate and send data packets according to this cycle. Meanwhile, the precision time protocol program module 1134 can be a lightweight PTP protocol stack implementation, also running on the microcontroller. Through PTP message interaction with the first host 12, it precisely calibrates the system clock of the pickup panel 11, ensuring nanosecond-level synchronization with the clock of the first host 12. In this way, the pickup panel 11 can convert the acquired audio signal into a standard network signal with extremely low latency and high-precision synchronization.
[0072] Through the above technical solution, the first processing device 113 can accurately control the packaging timing of digital audio signals and ensure high-precision clock synchronization between the pickup panel 11 and the first host 12, thereby improving the real-time performance and accuracy of audio signals during transmission and effectively avoiding problems such as audio distortion, data packet loss, or phase misalignment caused by inaccurate timing and clock asynchrony. In this way, high-quality, highly synchronized raw data can be provided to the first host 12 for audio processing (such as beamforming and sound source localization), ensuring the audio processing capability of the entire microphone system 10.
[0073] In order to avoid affecting the real-time transmission of audio signals, in some embodiments, the media access controller 1132 packages data according to a preset data packet length, which corresponds to the sampling period of the analog-to-digital conversion module 112.
[0074] Specifically, the preset data packet length refers to a fixed or configurable size pre-set for each data packet before data transmission to ensure the standardization and efficiency of data transmission. For example, it can be set to a fixed number of bytes or dynamically adjusted according to system requirements. Packetization refers to the process of encapsulating the raw digital signal data according to the preset data packet length and a specific protocol format to form a data packet (or encapsulated frame signal) that can be transmitted over the network. This process is usually completed by the media access controller 1132 to optimize data transmission efficiency, reduce transmission overhead, and facilitate decapsulation processing at the receiving end.
[0075] The sampling period refers to the time interval required for the analog-to-digital converter module 112 to sample the analog audio signal once. For example, when the sampling frequency is 48kHz, the sampling period is 1 / 48000 of a second, which is approximately 20 microseconds. The sampling period determines the time resolution and fidelity of the digital audio signal.
[0076] It should be understood that corresponding the preset data packet length to the sampling period of the analog-to-digital conversion module 112 indicates a direct and matching relationship between the two. For example, the amount of data contained in the preset data packet length corresponds to the amount of data generated within one sampling period, or the amount of data contained in the preset data packet length is an integer multiple of the amount of data generated within multiple sampling periods. This correspondence ensures the synchronization of data generation and data encapsulation, avoids data accumulation or waiting, and thus optimizes transmission latency.
[0077] In one specific implementation, the analog-to-digital converter module 112 in the pickup panel 11 can be configured to sample the audio signal at a sampling frequency of 48kHz, i.e., a sampling period of 20 microseconds. The media access controller 1132 in the first processing device 113 can be programmed to package data according to a preset data packet length of 512 bytes. Since 20 microseconds of sampled data corresponds to 512 bytes at a sampling rate of 48kHz, the media access controller 1132 can immediately package these data into a first encapsulated frame signal after the analog-to-digital converter module 112 completes one sampling (i.e., generates 512 bytes of data). The custom packet sending module 1133 can control the media access controller 1132 to precisely complete the packaging and prepare for transmission at the end of each sampling period, thereby realizing a "sample once, send once" mechanism. In this way, it can be ensured that the delay in the entire process from the sampling of the audio signal to its encapsulation into a network data packet and preparation for transmission is minimized, for example, an extremely low delay of only 20 microseconds can be achieved.
[0078] By using the above technical solution, and by setting the packetization behavior of the Media Access Controller 1132 (MAC controller) to correspond with the sampling period of the Analog-to-Digital Converter 112, low-latency optimization of signal transmission is achieved. Specifically, the Analog-to-Digital Converter 112 converts analog audio signals into first digital signals and generates data at a fixed sampling period. The Media Access Controller 1132 is configured to packetize data according to a preset data packet length corresponding to this sampling period. This correspondence ensures that whenever the Analog-to-Digital Converter 112 completes the generation of data for one or more sampling periods, the Media Access Controller 1132 can immediately encapsulate this data into a complete data packet. For example, if the Analog-to-Digital Converter 112 generates sampled data every 20 microseconds, the Media Access Controller 1132 will packetize the data according to a preset length that can accommodate 20 microseconds of sampled data. In this way, buffering and waiting of data within the Media Access Controller 1132 can be avoided, thereby significantly reducing the overall latency from sampling to data packet transmission. By adopting this technical solution, the first processing device 113 can package the first digital signal into a first encapsulated frame signal, and convert it into a standard network signal through the first physical layer device 114, and send it to the first host 12 via the transmission cable 13. This allows the entire microphone system 10 to maintain the real-time performance of the audio signal even under a physically separated setup architecture, providing high-quality, low-latency input for the first host 12 to process the standard network signal, thereby ensuring the overall audio processing performance and user experience of the system.
[0079] To avoid inefficient signal processing or degraded signal quality, in some embodiments, please refer to [the relevant documentation / reference]. Figure 9The first host 12 includes a second physical layer device 121 and a second processing device 122. In this embodiment, the second physical layer device 121 is used to convert standard network signals into second encapsulated frame signals. The second processing device 122 is connected to the second physical layer device 121 to receive the second encapsulated frame signals and perform processing operations to obtain an output signal.
[0080] The second physical layer device 121 refers to a hardware unit that implements physical layer communication functions. It is capable of converting standard network signals (e.g., Ethernet electrical or optical signals) transmitted from the transmission cable 13 into second-encapsulated frame signals that can be recognized and processed by internal digital logic circuits. It is also responsible for physical layer operations such as signal modulation / demodulation, encoding / decoding, clock recovery, and media access control. As one possible implementation, the second physical layer device 121 can be a standalone Ethernet PHY (Physical Layer Transceiver) chip, such as a common Gigabit Ethernet PHY chip, which connects to the upper-layer MAC controller via a standard interface (such as RGMII or GMII).
[0081] The second encapsulated frame signal refers to the digital signal extracted from the standard network signal and conforming to a specific data link layer protocol format after processing by the second physical layer device 121. In this embodiment, the second encapsulated frame signal is typically in the form of a data frame, containing the original audio signal (i.e., the audio signal picked up by the pickup panel 11) and necessary protocol headers and checksum information, which can be directly parsed by upper-layer processing logic (such as the media access controller 1132 MAC). As one possible implementation, the second encapsulated frame signal can be a digital signal in Ethernet MAC frame format, which is transmitted between the second physical layer device 121 and the second processing device 122 through interfaces such as RGMII or MII.
[0082] The second processing device 122 is the main processing device in the first host 12, used to receive and parse the second encapsulated frame signal, and perform processing operations to obtain an output signal. The processing operations performed on the second encapsulated frame signal can optimize audio quality, implement specific functions, and ultimately generate an output signal that meets system requirements. As one possible implementation, the second processing device 122 can be a high-performance system-on-a-chip (SoC).
[0083] Connecting the second physical layer device 121 to the second processing device 122 establishes a data channel for transmitting the second encapsulated frame signal, ensuring that the second encapsulated frame signal converted by the second physical layer device 121 can be accurately transmitted to the second processing device 122 for processing. Regarding the connection interface between the second physical layer device 121 and the second processing device 122, a standardized digital interface can be used to ensure the integrity of data transmission and the accuracy of timing. As a feasible implementation, the connection between the two can use the RGMII (Reduced Gigabit Media Independent Interface), a parallel digital interface widely used in Gigabit Ethernet devices, which enables high-speed data transmission with a relatively low pin count.
[0084] By introducing a second physical layer device 121 and a second processing device 122 into the first host 12, the above technical solution avoids the need for efficient conversion of standard network signals before processing, as well as the inefficiency and signal quality degradation caused by format mismatch. On one hand, the second physical layer device 121 can convert standard network signals into second encapsulated frame signals, providing a standardized input format for the second processing device 122, thus simplifying the signal reception and parsing complexity of the second processing device 122. Utilizing the direct and efficient conversion mechanism of the second physical layer device 121, in conjunction with the signal processing function of the second processing device 122, ensures the stability of audio signal transmission from the network to the first host 12 for processing, reduces signal processing latency, and improves the accuracy and reliability of signal processing. On the other hand, based on the technical solution of the second processing device 122 sending a clock synchronization signal to the pickup panel 11 to synchronize the sampling clock of each pickup panel 11 with the system clock of the first host 12, the physical layer processing of the standard network signal by the second physical layer device 121 further ensures the accurate transmission and reception of the clock synchronization signal and the precise timing of the audio data itself. In this way, it can also provide accurate data for the second processing device 122 to perform processing operations, thereby enabling the output of higher quality and more spatial audio signals, so as to improve the processing performance and output audio quality of the entire microphone system 10 in complex audio environments.
[0085] In some embodiments, please refer to Figure 9 The second processing device 122 includes a first processor 1221. In this embodiment, the first processor 1221 is used to receive a second encapsulation frame signal and convert the second encapsulation frame signal into a second digital signal.
[0086] It should be understood that the first processor 1221 is an electronic component or module used to perform specific data processing tasks. In the embodiments of this application, the first processor 1221 can be a digital signal processor (DSP) integrated into a system-on-a-chip (SoC). The DSP is capable of performing in-depth processing of digital signals according to digital signal processing algorithms. Digital signal processing algorithms include, but are not limited to, complex digital filtering, adaptive beamforming, acoustic echo cancellation, noise suppression, and speech enhancement algorithms, which achieve high-precision processing of audio signals.
[0087] In this embodiment, the first processor 1221 receives the second encapsulation frame signal to ensure that it can obtain the demodulated standard network signal from the second physical layer device 121, thereby guaranteeing data integrity and timeliness. This receiving process allows the first processor 1221 to interact with the second physical layer device 121 at the data link layer or physical layer via a high-speed serial interface (e.g., SPI, I2S, EthernetMAC interface).
[0088] Through the above technical solution, by introducing a first processor 1221 into the second processing device 122, the first processor 1221 is specifically responsible for receiving the second encapsulated frame signal converted by the second physical layer device 121 and converting the second encapsulated frame signal into a second digital signal. This allows the second processing device 122 to effectively separate network layer data processing from subsequent audio algorithm processing, avoiding low signal conversion efficiency or increased processing latency. This provides high-quality, standardized digital audio input for subsequent audio processing operations, improving the smoothness and reliability of data processing in the entire microphone system 10. Specifically, the second physical layer device 121 is responsible for demodulating the standard network signal from the physical medium and converting it into a second encapsulated frame signal, while the first processor 1221 acts as the first processing unit of the data stream to parse the second encapsulated frame signal, remove network protocol overhead, and extract audio digital data. This ensures a smooth transition of data from the network transmission format to the processable digital audio format, avoiding resource conflicts and efficiency degradation that might occur if a single processing unit simultaneously undertakes complex network protocol parsing and high-performance audio algorithm calculations. On the other hand, by using the first processor 1221 to convert the second packaged frame signal into a second digital signal, the allocation of system resources can be optimized, and the overall processing performance and response speed can be improved.
[0089] To meet the needs of high-quality conferencing or audio applications, in some embodiments, the first processor 1221 is further configured to perform processing operations on the second digital signal to obtain a third digital signal. These processing operations include at least one of the following: noise reduction, automatic mixing, echo cancellation, and audio enhancement.
[0090] It should be understood that by setting the first processor 1221 to perform noise reduction, automatic mixing, echo cancellation, and enhancement on the second digital signal, the clarity and intelligibility of the audio signal can be effectively improved, meeting the auditory needs of different application scenarios and thus improving the quality of meetings or audio applications. The third digital signal is the processed digital audio signal, that is, the result of the first processor 1221 performing noise reduction, automatic mixing, echo cancellation, or audio enhancement on the second digital signal, and has higher clarity, lower noise, or better mixing effects.
[0091] Noise reduction refers to the process of identifying and suppressing unwanted noise components in audio signals using algorithms, while preserving the clarity of speech or target audio as much as possible. For example, spectral subtraction can be used to estimate the noise spectrum and subtract it from the noisy signal spectrum to reduce noise; alternatively, a deep learning-based neural network model can be trained to identify and suppress noise patterns; or an adaptive filter can be used to estimate the noise signal and eliminate it from the input signal. This improves the signal-to-noise ratio of speech, making the spoken content clearer and more intelligible.
[0092] Automatic mixing refers to dynamically adjusting the gain of each signal based on the activity level of multiple audio input sources and preset rules to ensure the balance and clarity of all speakers' voices. For example, a gating auto mixer can be used, which activates a channel when the signal strength exceeds a certain threshold and adjusts the total gain according to the number of channels; or a gain-sharing auto mixer can be used, where the total gain of all channels remains constant, and the active channel receives more gain. Of course, algorithms based on speech activity detection (VAD) and intelligent gain control can also be used to manage the mixing more finely, so as to avoid the problem of multiple speakers' voices masking each other when speaking at the same time, or the problem of a single speaker's voice being too loud / too soft, and to provide a smooth and natural mixing effect.
[0093] Echo cancellation refers to the use of algorithms to identify and eliminate echoes in audio signals caused by the microphone picking up sound from a speaker, preventing sound feedback and distortion. For example, adaptive filters (such as NLMS and RLS algorithms) can be used to estimate the echo path and generate echo copies for cancellation; or nonlinear processing (such as nonlinear processing units, NLP) can be combined to process residual echoes, further improving the cancellation effect. This ensures the clarity of two-way communication and avoids the "feedback" or "hollow feeling" commonly seen in meetings.
[0094] Audio enhancement refers to improving the perceived quality, clarity, or loudness of an audio signal by adjusting its frequency response, dynamic range, or other characteristics. For example, an equalizer (EQ) can be used to adjust the gain of different frequencies to optimize timbre; automatic gain control (AGC) or a compressor can be used to smooth volume fluctuations and make the sound more stable; and an exciter can be used to increase harmonic components, making the sound more penetrating or immersive. Its purpose is to improve the overall listening experience of audio, making it more expressive or easier to understand.
[0095] Through the above technical solution, when the pickup panel 11 picks up the audio signal and converts it into a standard network signal, and then sends it to the first host 12 via the transmission cable 13, the second physical layer device 121 in the first host 12 converts the standard network signal into a second encapsulated frame signal. The first processor 1221 in the second processing device 122 receives the second encapsulated frame signal and converts it into a second digital signal. Based on this, the first processor 1221 further processes the second digital signal, performing noise reduction, automatic mixing, echo cancellation, or audio enhancement, thereby obtaining a high-quality third digital signal. Using this technical solution, the first processor 1221 undertakes the basic task of signal conversion on the one hand, and on the other hand, is responsible for optimizing the audio quality, ensuring that the audio signal remains in optimal condition throughout the entire link from the original audio pickup to the final output. Furthermore, based on the aforementioned situation where the microphone system 10 includes multiple pickup panels 11 and the sampling clock of each pickup panel 11 is synchronized with the system clock of the first host 12, the first processor 1221 can utilize the second digital signal synchronized by the multiple pickup panels 11 to perform operations such as automatic mixing, noise reduction, and echo cancellation. For example, by analyzing the synchronization signals from different pickup panels 11, the location of the sound source can be identified more accurately, speech and noise can be distinguished, or a more effective echo cancellation model can be constructed, thereby improving the audio processing effect in multi-microphone scenarios.
[0096] The following is a concrete example: In a conference room equipped with multiple microphone panels 11, a first host 12 receives synchronous standard network signals from each microphone panel 11. A second physical layer device 121 in the first host 12 converts these standard network signals into second encapsulated frame signals and transmits them to a second processing unit 122. Upon receiving the second encapsulated frame signals, a first processor 1221 in the second processing unit 122 converts them into second digital signals and performs a series of processes on them to obtain a third digital signal. For example, when there is background noise from the air conditioner or projector in the conference room, the first processor 1221 can use a spectral subtraction-based noise reduction algorithm to identify and filter out noise components in the second digital signal, thereby obtaining a cleaner third digital signal. When multiple participants speak simultaneously, the first processor 1221 can use a gain-sharing automatic mixing algorithm to dynamically adjust the gain of each second digital signal, ensuring that the voices of all speakers remain balanced and clear in the mixed third digital signal. If the conference room speakers are playing the voice of a remote participant, and that voice is picked up again by the local audio pickup panel 11, the first processor 1221 will activate an adaptive echo cancellation algorithm. This algorithm estimates the echo path and generates an echo copy to cancel it out, thereby eliminating the echo in the third digital signal. Furthermore, the first processor 1221 can also adjust the frequency response of the second digital signal using an equalizer or smooth volume fluctuations using automatic gain control to enhance audio and improve the listening experience of the output third digital signal.
[0097] In some embodiments, please refer to Figure 9 The first host 12 also includes a second processor 123 and at least one output port M1. The second processor 123 is used to receive a third digital signal and distribute it to the corresponding output port M1.
[0098] It should be understood that the second processor 123 is an independent computing unit whose primary responsibility is to perform data routing, management, and distribution tasks. Specifically, the second processor 123 can be implemented as a central processing unit integrated into a system-on-a-chip (SoC) with sufficient input / output (I / O) interfaces and real-time processing capabilities to efficiently process and distribute audio data.
[0099] At least one output port M1 is a physical interface for transmitting a third digital signal to an external device. For example, the output interface may include an analog audio output port M1, such as an RCA, XLR, or 3.5mm TRS connector, for connecting analog audio devices such as amplifiers or active speakers. It may also include a digital audio output port M1, such as an S / PDIF (optical or coaxial) or AES / EBU connector, for connecting a digital audio processor or recording equipment. Alternatively, it may include a network audio output port M1, such as a Dante, AVB (Audio Video Bridging), or AES67 connector, for transmitting audio streams over Ethernet to compatible network audio devices.
[0100] By having a second processor 123 mounted on the first host 12 directly obtain the processed third digital signal from the first processor 1221, the continuity and accuracy of signal transmission are ensured. Furthermore, the second processor 123 efficiently routes the third digital signal to at least one output port M1 according to system configuration or real-time instructions. This allows the first processor 1221 to execute the core audio processing algorithm, while the second processor 123 handles output management. This optimizes the resource allocation of the microphone system 10, improves overall response speed and output management efficiency, and ensures that the microphone system 10 maintains high performance and stability even when faced with multi-output requirements or complex routing logic.
[0101] Please continue to refer to this. Figure 9 In this embodiment, the output port M1 includes a first output port M11 and a second output port M12. The first output port M11 converts the third digital signal into an analog signal and outputs it, while the second output port M12 packages the third digital signal into a network protocol packet and outputs it. With this technical solution, on the one hand, the second processor 123 connects to an external audio device through the first output port M11, transmitting the processed audio signal to an external speaker or headphones, achieving high-quality audio output; on the other hand, the second processor 123 can also connect to a network device through the second output port M12. Thus, the second output port M12 packages the third digital signal into a network protocol packet and outputs it to the network device, supporting remote transmission and distribution of audio signals.
[0102] For example, the processed third digital signal can be converted into an analog signal for local playback. Simultaneously, the processed third digital signal can be packaged into network protocol packets and output to a network device for storage, such as for recording, or for remote playback, meeting the needs of scenarios like remote conferencing and online teaching. In this example, the first output port M11 can be a line output port M1 (i.e., line out output port M1), which can output a standard line-level analog audio signal without power amplification. The processed digital signal is then converted back into an analog signal using the line out port M1 for local playback. The second output port M12 can be a Dante output port M1. In this example, the Dante output port M1 follows the Dante protocol, packaging the third digital signal into network protocol packets for low-latency transmission to the Dante network via Ethernet.
[0103] Through the above technical solution, on the one hand, the second processor 123 is responsible for handling the output distribution task, thereby effectively reducing the burden on the first processor 1221, freeing up the computing power of the first processor 1221, and enabling the first processor 1221 to have sufficient computing power to execute complex audio processing algorithms, so as to improve the response speed and output management efficiency of the microphone system 10. On the other hand, by providing at least one output port M1, it can be ensured that the processed high-quality audio signal is flexibly and efficiently transmitted to various external devices (such as local playback or online playback).
[0104] In some embodiments, the second processor 123 is further configured to manage the operating state of the microphone system 10 and to perform parameter configuration of the microphone system 10.
[0105] Specifically, managing the operating status of the microphone system 10 refers to real-time monitoring, recording, and analysis of various indicators, events, and resource usage of the microphone system 10 during operation, in order to ensure the stable and efficient operation of the microphone system 10 and to promptly detect and warn of potential problems. This includes, but is not limited to, monitoring the online status, connection quality, audio data stream integrity, processor load, memory usage, network bandwidth usage, and internal system errors or abnormal events of each pickup panel 11.
[0106] The second processor 123 can periodically send heartbeat packets or status query commands to each pickup panel 11, and receive and parse the operating status information returned by the panels, such as power supply voltage, temperature, and the health status of the microphone array 111. Simultaneously, the second processor 123 can monitor its own and the connected output ports M1's operating status, such as the output signal level and the online status of connected devices. Furthermore, the second processor 123 can maintain a system status table or log, either internally or through interaction with external memory, recording key events and performance data. When an anomaly is detected, the second processor 123 can trigger an early warning mechanism, such as sending alarm information to the management platform via a network interface or storing error logs locally. In this embodiment, the management of the microphone system 10's operating status ensures the long-term stable and reliable operation of the microphone system 10, guaranteeing its availability and stability.
[0107] Parameter configuration of the microphone system 10 refers to setting, modifying, and updating various adjustable parameters within the microphone system 10 based on actual application scenarios, user needs, or system optimization goals. Parameter configuration of the microphone system 10 can affect various stages of audio acquisition, processing, transmission, and output, such as audio gain, noise reduction intensity, the adaptability of echo cancellation algorithms, automatic mixing thresholds, beamforming direction, QoS settings for network transmission protocols, and routing rules for output port M1. After receiving the parameter configuration instructions, the second processor 123 parses them and applies them to the corresponding hardware modules or software algorithms, such as updating the register settings of the digital signal processor (DSP) or modifying the parameter variables of the internal algorithm. Furthermore, the second processor 123 can preset multiple configuration modes, which users can switch between with simple commands. Each mode corresponds to a predefined set of parameters. Parameter configuration is a key function that enables the microphone system 10 to adapt to different environments and needs, allowing the system to be optimized based on factors such as conference room size, acoustic characteristics, number of speakers, and noise levels, thereby ensuring the audio performance of the microphone system 10.
[0108] Through the above technical solution, the operation status management and parameter configuration functions of the microphone system 10 are integrated into the second processor 123, thereby improving the manageability and adaptability of the microphone system 10 on the basis of its original high-quality audio processing and distribution capabilities.
[0109] In some embodiments, the pickup panel 11 has a thickness of 3 mm to 10 mm and is configured to be mounted on a ceiling 20 or a desktop.
[0110] The pickup panel 11 in this embodiment includes a microphone array 111, an analog-to-digital converter module 112, and a first processing device 113 for converting audio signals into standard network signals. In the split-type microphone system 10, the pickup panel 11 is physically separated from the first host 12, so the pickup panel 11 can perform audio acquisition and signal digitization and networking at the front end.
[0111] In this embodiment, the thickness of the microphone panel 11 is limited to 3mm to 10mm, giving it an ultra-thin characteristic. This allows the microphone panel 11 to blend into the installation environment with minimal visual presence, avoiding the obtrusiveness of traditional bulky equipment. In this embodiment, the microphone panel 11 is configured to be installed in the ceiling 20; for example, it can be fixed inside the ceiling 20. In this embodiment, the microphone panel 11 is configured to be installed on a desktop, meaning it can be placed or fixed on a flat surface such as a conference table or office desk. For example, it can be embedded in the desktop through an opening, making the panel surface flush with the desktop; alternatively, it can be fixed to the edge of the desktop using a desktop clamp or bracket. No limitation is imposed; the actual configuration will prevail.
[0112] Please continue to refer to this. Figure 9 In some embodiments, the pickup panel 11 further includes a first module interface M2, the first host 12 further includes a second module interface M3, and the transmission cable 13 is connected between the first module interface M2 and the second module interface M3.
[0113] It should be understood that the first module interface M2 refers to a physical connection port located on the pickup panel 11, used for electrical and mechanical connection with the transmission cable 13. This first module interface M2 can be a standardized interface, such as an RJ45 interface, with a snap-fit mechanism to ensure a secure connection. The first module interface M2 provides a clearly defined, reusable access point for the transmission cable 13.
[0114] The second module interface M3 refers to a physical connection port located on the first host 12, which matches the first module interface M2 on the microphone panel 11, and is used for the connection of the transmission cable 13. Similar to the first module interface M2, the second module interface M3 can be a standardized RJ45 interface, facilitating connection with general Ethernet cables. The second module interface M3 serves as a standardized entry point for the first host 12 to receive standard network signals from the microphone panel 11 and provide power.
[0115] In this embodiment, the transmission cable 13 is connected between the first module interface M2 and the second module interface M3, thus establishing a physical link between the pickup panel 11 and the first host 12. Connecting the transmission cable 13 through the first module interface M2 and the second module interface M3 ensures the integrity of signal transmission and the stability of power supply. Both ends of the transmission cable 13 are equipped with connectors that match the first module interface M2 and the second module interface M3, for example, if the interface is RJ45, then both ends of the cable are RJ45 crystal heads, allowing the cable to be securely fixed in the two module interfaces and reducing the risk of connection interruption due to accidental pulling or vibration.
[0116] Taking the example where both the first module interface M2 and the second module interface M3 are configured as RJ45 interfaces, during installation, the operator only needs to insert one end of the Ethernet cable into the RJ45 interface of the pickup panel 11 and the other end into the RJ45 interface of the first host 12. The internal locking mechanism of the interface ensures that the cable is securely locked. When maintenance or equipment replacement is required, simply press the clip on the RJ45 connector to easily pull out the cable, achieving quick disassembly and reconnection.
[0117] In this embodiment, the clock of the second module interface M3 of the first host 12 is synchronized with the clock of the network audio output port M1. If both the second module interface M3 and the output port M1 are network audio ports, such as Dante, AVB (Audio Video Bridging), or AES67 interfaces, the audio stream is transmitted via Ethernet to a compatible network audio device (such as a speaker, recording device, audio processor, audio codec, or video conferencing host).
[0118] Please continue to refer to this. Figure 9 In some embodiments, the first host 12 is further provided with a second power supply port for receiving power from an external device and for enabling data interaction with the external device. For example, the external device may also be a network switch 124 capable of providing DC power. The second power supply port and the connection ports included in the external device can be connected by a transmission cable to achieve power and data interaction. In this example, the second power supply port may be a Power over Ethernet (PoE) port.
[0119] It should be noted that in this embodiment of the application, the external devices connected to the pickup panel 11 and the first host 12 can be the same external device, such as the same network switch 124, or different external devices, thereby realizing distributed deployment.
[0120] In this embodiment, the first host 12 is connected to the second host 14 via a second power supply port. In this structure, one end of the first host 12 is connected to multiple microphone panels 11, and the other end is connected to the second host 14, forming a star topology. Compared to the traditional method of individually wiring each microphone to the second host 14, this greatly simplifies system wiring, reduces wiring complexity and cost, and enables centralized management and intelligent control of the devices, facilitating centralized management and fault isolation. In this example, the second host 14 can be a conference host. The second host 14 is used to interface with the conference system platform to process, forward, and store audio data. It can also send control commands to the first host 12 to manage the working status of the microphone panels 11 and set configuration parameters. For example, refer to... Figure 9 The microphone panel 11 and the first host 12 shown are connected by a transmission cable 13 to transmit signals and exchange electrical energy.
[0121] Figure 11 This is a schematic diagram of the overall structural architecture of the first host 12, including the switch 124, provided in an exemplary embodiment of this application.
[0122] When the first host 12 of the microphone system 10 needs to connect multiple pickup panels 11, the first host 12 may lack an effective interface and switch 124, causing each pickup panel 11 to need to be connected separately to the second processing device 122. This results in redundant wiring, complex signal management, and limited system scalability. Therefore, please refer to... Figure 11 In some embodiments, the first host 12 further includes a switch 124 and multiple second module interfaces M3. The multiple second module interfaces M3 are used to connect transmission cables 13 corresponding to the microphone panels 11. The microphone panels 11 connected to the multiple module interfaces M3 are synchronized with the clock of the first host 12. These clock synchronization signals are transmitted to each microphone panel 11 via the transmission cables 13. After receiving the signals, the microphone panels 11 adjust their own clock signals according to the received clock synchronization signals, so that the clocks of each microphone panel 11 and the microphone system 10 composed of the first host 12 are synchronized. The switch 124 is connected between the second module interfaces M3 and the second processing device 122, so that all transmission cables 13 are connected to the same switch 124.
[0123] Multiple second module interfaces M3 refer to the ports that physically connect the first host 12 to the external transmission cable 13, which can provide an independent access point for each independent pickup panel 11, thereby allowing the first host 12 to connect to and manage multiple pickup panels 11 at the same time.
[0124] It should be understood that the switch 124 in this embodiment receives standard network signals carried by transmission cables 13 from different pickup panels 11 and forwards the standard network signals to the second processing device 122 according to the target address. The switch 124 in this embodiment integrates multiple second physical layer devices 121, connecting each transmission cable 13 to the interface of the corresponding second physical layer device 121. Signals from the various second module interfaces M3 connected to each transmission cable 13 are first converged to the switch 124, where they undergo preliminary routing and management before being uniformly transmitted to the second processing device 122 for further audio processing. This avoids the second processing device 122 directly facing multiple independent physical connections, thus simplifying the design complexity of the second processing device 122. By concentrating all transmission cables 13 (i.e., connections to all pickup panels 11) onto the same switch 124, the microphone system 10 can form a unified network topology, facilitating centralized management and unified scheduling of data flows to achieve efficient signal routing and resource allocation.
[0125] As a specific implementation, the example is a switch 124 configured as a Gigabit Ethernet switching chip. In a large conference room, two microphone panels 11 may need to be deployed to ensure comprehensive audio coverage. (Reference) Figure 11 In this embodiment, the microphone system 10's first host 12 can integrate a switch 124 with two RJ45 interfaces. Each pickup panel 11 is connected to an RJ45 interface on the first host 12 via a Cat5e Ethernet cable. These RJ45 interfaces are internally connected to a Gigabit Ethernet switching chip, which is then connected to a second processing device 122 inside the first host 12 via a high-speed bus. When the two pickup panels 11 simultaneously pick up audio signals and convert them into standard network signals, these signals are transmitted to the RJ45 interfaces of the first host 12 via their respective Ethernet cables. After receiving these two independent network signal streams, the switching chip efficiently forwards these data packets to the second processing device 122 according to its internal MAC address table or VLAN configuration. Upon receiving the converged data stream, the second processing device 122 can perform unified audio processing, such as noise reduction, automatic mixing, and beamforming or sound source localization in conjunction with clock synchronization information.
[0126] It should be noted that the second processing device 122 described in this specific embodiment can be the second processing device 122 described above, that is, the second processing device 122 includes a first processor 1221 and a second processor 123, where the first processor 1221 is a DSP and the second processor 123 is a central processing unit.
[0127] The above technical solution achieves the effect of expanding multiple pickup panels 11, thus solving the problem that the first host 12 may lack effective interfaces and switches 124 in multi-pickup panel 11 deployment scenarios, leading to redundant wiring, complex signal management, and limited system scalability. On one hand, by integrating multiple second module interfaces M3 and a switch 124 into the first host 12, centralized and standardized connections of multiple pickup panels 11 are achieved, reducing wiring complexity and cost, avoiding redundancy caused by each pickup panel 11 being individually connected to the second processing device 122, and simplifying signal management through signal aggregation and routing via the switch 124 to improve data transmission efficiency. On the other hand, this architecture enhances the scalability of the microphone system 10, allowing for flexible increases or decreases in the number of pickup panels 11 according to actual needs without large-scale modifications to the core processing logic of the first host 12. Furthermore, the second processing device 122 can receive and process synchronous audio data from all pickup panels 11, simplifying the overall architecture of the first host 12.
[0128] In some embodiments, the pickup panel 11 also integrates a speaker unit 115. The first host 12 in this embodiment is further configured to: receive external audio signals, convert the external audio signals into standard network signals, and then send them to the pickup panel 11 via a transmission cable 13. The pickup panel 11 in this embodiment is further configured to: receive standard network signals via the transmission cable 13, convert them into analog audio signals, and then drive the speaker unit 115 to play them.
[0129] It should be understood that speaker unit 115 refers to a sound playback device (such as a speaker) integrated inside the pickup panel 11, used to convert audio signals back into sound signals and play them. External audio signals refer to audio signals input to the first host 12 that need to be played, such as voices from remote participants in a remote conferencing terminal, music or notification tones from locally connected media playback devices such as computers and mobile phones, or notification tones or processed audio echo cancellation reference signals generated by the first host 12 itself.
[0130] In this embodiment, the speaker unit 115 can be directly disposed inside the housing of the pickup panel 11, coexisting in the same cavity with the audio acquisition module 11a and network conversion module 11b included in the pickup panel 11. Alternatively, it can occupy a separate acoustic chamber. The speaker unit 115 transforms the pickup panel 11 from a simple audio acquisition device into an integrated input / output terminal with audio playback functionality. This eliminates the need for separate ceiling-mounted speakers and corresponding amplifiers, as is common in conference rooms. The first host 12 is also used to receive external audio signals, convert them into standard network signals, and send them to the pickup panel 11 via the transmission cable 13. Optionally, the second processing device 122 inside the first host 12 decodes, mixes, and reduces noise for these audio signals, and then encapsulates them into network data packets conforming to the Ethernet protocol. These network data packets are modulated into standard network signals by the second physical layer device 121 (i.e., the PHY chip) inside the first host 12.
[0131] It should be noted that the pickup panel 11 does not make any substantial changes or feature extraction to the content of the audio signal; it only sends the raw audio data to the first host 12 via the network. The first host 12 executes a digital signal processor (DSP) or algorithm module for mixing, noise reduction, sound source localization, beamforming, echo cancellation, or audio enhancement. After the first host 12 converts the external audio signal into a standard network signal, it sends it to the pickup panel 11 via the transmission cable 13. Before playing the external audio signal, the pickup panel 11 can perform relevant sound effect processing on the audio signal to be played, such as equalization, gain adjustment, and echo addition.
[0132] In this technical solution, the transmission cable 13 in this embodiment carries not only standard network signals in the uplink direction but also standard network signals in the downlink direction. Uplink and downlink data streams can be transmitted simultaneously on different wire pairs within the same twisted pair without interference, achieving multiplexing of the transmission cable 13 and avoiding the cumbersome process of separate wiring for downlink sound reinforcement. The pickup panel 11 is also used to receive standard network signals via the transmission cable 13, convert them into analog audio signals, and then drive the speaker unit 115 to play them. Specifically, after receiving the downlink standard network signal, the first physical layer device 114 (i.e., the PHY chip) inside the pickup panel 11 performs demodulation and unpacking operations to restore the digital audio data. The first processing device 113 sends the digital audio data to the analog-to-digital converter module 112 (i.e., the ADC chip) to convert the digital signal into an analog electrical signal. After being amplified by the power amplifier, the analog electrical signal drives the diaphragm of the speaker unit 115 to vibrate, thereby producing sound.
[0133] Through the above technical solutions and the above two-way audio transmission architecture, a complete, closed-loop audio interaction system can be constructed.
[0134] Figure 12 This is another schematic diagram of the overall structural architecture of the microphone system 10 provided in the exemplary embodiment of this application. Please refer to... Figure 12In some embodiments, the microphone panel 11 further includes a signal switching module 116. The signal switching module 116 in this embodiment is used to selectively transmit the audio signal picked up by the microphone array 111 or the external audio signal received by the transmission cable 13 to the speaker unit 115.
[0135] Specifically, the signal switching module 116 is a circuit or logic unit used to select from multiple input audio signals and route the selected signal to a single output, ensuring that only one audio source is transmitted to the speaker unit 115 at any given time, thereby avoiding signal conflicts and potential echo interference. In practical implementation, the signal switching module 116 can take various forms. For example, at the hardware level, it can be an analog multiplexer chip that selects different analog audio input channels through control signals; it can also be a digital audio switch that routes audio streams in the digital domain. At the software level, if the audio signal has been digitized, dynamic selection and routing of different digital audio streams can be implemented through programming logic in the first processing device 113 or elsewhere.
[0136] Figure 13 This is another schematic diagram of the overall structural architecture of the microphone system 10 provided in the exemplary embodiment of this application. Please refer to... Figure 13 In some embodiments, the microphone panel 11 further includes an audio mixing module 117. The audio mixing module 117 in this embodiment is used to mix the audio signal picked up by the microphone array 111 with the external audio signal received by the transmission cable 13 and then transmit it to the speaker unit 115.
[0137] Specifically, the audio mixing module 117 is a circuit or logic unit capable of superimposing (mixing) two or more input audio signals into a composite audio signal, allowing simultaneous playback of sounds from different audio sources to achieve coordinated output or a richer auditory experience. In practical implementation, the audio mixing module 117 can be an analog mixing circuit that uses a resistor network or operational amplifier to weightedly superimpose analog audio signals; or it can be a digital mixer that performs weighted summation of digital audio samples in the digital domain. In digital systems, this function is typically implemented by the first processing unit 113 or a digital signal processor through software algorithms, enabling independent gain control and phase adjustment of each audio signal to optimize the mixing effect.
[0138] For example, when the pickup panel 11 includes a signal switching module 116, the signal switching module 116 can selectively transmit either the audio signal picked up by the microphone array 111 or the external audio signal received by the transmission cable 13 to the speaker unit 115 for playback, according to preset control commands or user input. This selective transmission mechanism ensures that at any given time, the speaker unit 115 plays only audio from a single source, thereby effectively avoiding the conflict and echo problems that may be caused by the simultaneous output of different audio signals, and ensuring the clarity and stability of the audio output.
[0139] For example, when the microphone panel 11 includes an audio mixing module 117, the audio mixing module 117 can superimpose and mix the audio signal picked up by the microphone array 111 with the external audio signal received by the transmission cable 13, and then transmit the mixed composite audio signal to the speaker unit 115 for playback. This mixing mechanism allows locally picked-up audio (such as the voice of a conference speaker) to be output simultaneously with external audio (such as the voice of a remote conference participant or background music), thereby achieving a more immersive and collaborative audio experience. For example, in a video conferencing scenario, the voice of the local speaker can be played simultaneously with the voice of the remote participant through the speaker unit 115, allowing the local participant to clearly hear the remote voice while also hearing their own voice, enhancing the sense of communication.
[0140] In this embodiment, both the signal switching module 116 and the audio mixing module 117 work closely with the existing audio acquisition module 11a and network conversion module 11b in the microphone panel 11. Through the transmission cable 13, the first host 12 can convert external audio signals into standard network signals and send them to the microphone panel 11. The microphone panel 11 receives these signals and converts them into analog audio signals for playback by the speaker unit 115. Based on this, the signal switching module 116 or the audio mixing module 117 manages and processes these potential audio sources to ensure that the output of the speaker unit 115 conforms to the expected audio strategy. Thus, the microphone panel 11 not only has audio input functionality but also, by integrating the speaker unit 115 and the intelligent audio management module, achieves two-way audio interaction, greatly enhancing the functional integrity and application flexibility of the microphone system 10.
[0141] For example, in conference speaking mode, the first processing device 113 can select to transmit only the audio signal picked up by the microphone array 111; while when playing background music or remote conference audio, it switches to the external audio signal received by the transmission cable 13. As another specific implementation, when the pickup panel 11 is equipped with an audio mixing module 117, this audio mixing module 117 can also be implemented by a software algorithm within the first processing device 113. The first processing device 113 receives the audio signal from the microphone array 111 and the external audio signal from the transmission cable 13. The first processing device 113 executes a digital mixing algorithm to perform a weighted summation of the two digital audio signals, generating a mixed digital audio signal. For example, a gain coefficient can be applied to the signal from the microphone array 111, and another gain coefficient can be applied to the external audio signal, and then the two can be superimposed. The mixed digital audio signal is then transmitted to the digital-to-analog converter driving the speaker unit 115, thereby achieving synchronized playback of local speech and remote audio. For example, in a video conference, the voices of local participants and remote participants can be played simultaneously through speaker unit 115 to provide a more natural communication experience. For instance, in a video conference scenario, local participants can clearly hear the voices of remote participants while also hearing their own voice, creating a more natural dialogue environment. This flexible audio management mechanism allows microphone system 10 to adapt to more diverse application scenarios, optimize audio output modes, and improve the overall quality and efficiency of audio interaction.
[0142] In some embodiments, the microphone system 10 includes a plurality of pickup panels 11. The first host 12 is further configured to: extract the audio signals collected by each pickup panel 11 based on the standard network signals sent by each pickup panel 11, and dynamically select a portion of the pickup panels 11 to form an effective pickup array.
[0143] Specifically, the microphone system 10 is equipped with multiple pickup panels 11, which are physically separated and can be distributed and deployed in a preset area. The first host 12 receives standard network signals from each pickup panel 11. The standard network signal is a data stream conforming to the Ethernet protocol, which is encapsulated and modulated by the network conversion module 11b and transmitted via the transmission cable 13. The first host 12, through its second physical layer device 121 and second processing unit 122, can identify and receive these independent network data streams from different pickup panels 11. After receiving the standard network signal, the second processing unit 122 (e.g., the first processor 1221) within the first host 12 is responsible for decapsulating and demodulating these standard network signals, restoring them to the original digital audio signal, ensuring that the audio data transmitted from each pickup panel 11 can be accurately identified and processed by the first host 12.
[0144] Dynamic selection refers to the first host 12 intelligently deciding to enable or disable some pickup panels 11 based on preset strategies or real-time analysis results to form an optimal pickup array for the current environment. Dynamic selection can combine the following factors, such as: the first host 12 can analyze the signal-to-noise ratio, loudness, or clarity of the audio signals collected by each pickup panel 11 in real time, prioritizing panels with high signal quality; combined with sound source localization algorithms, the first host 12 can identify the location of the current main sound source and select the panel closest to the sound source or with the best sound pickup effect. The first host 12 can also identify noise sources in specific areas and temporarily disable or reduce the weight of pickup panels 11 in that area to reduce noise introduction. Alternatively, the first host 12 allows users to manually select or exclude certain panels through the control interface.
[0145] An effective pickup array refers to a virtual microphone array 111 composed of selected pickup panels 11 that provides optimal audio pickup performance. By dynamically adjusting the array composition, the microphone system 10 can adapt to changing acoustic environments and speaker positions, thereby optimizing overall pickup performance.
[0146] Through the above technical solution, the microphone system 10 can overcome the limitations of the fixed array of traditional multi-microphone systems 10. In dynamically changing conference environments, such as when the speaker moves, local noise occurs, or the conference area is reconstructed, the first host 12 can identify and select the pickup panel 11 that best suits the current acoustic conditions, thus forming an "effective pickup array". This technical solution improves the focus and clarity of audio acquisition, effectively suppresses unnecessary environmental noise and reverberation, and avoids redundant information and processing burdens that may be introduced by using all panels in a fixed manner. As a result, the microphone system 10 can provide higher quality output signals, ensuring that remote participants receive a clearer and more professional listening experience. At the same time, it optimizes the resource utilization of the microphone system 10, enabling it to exhibit good adaptability and performance in various complex and dynamic application scenarios.
[0147] In some embodiments, the first host 12 is further configured to: extract the spatial features of the audio signals collected by each pickup panel 11, and generate a spatial audio signal based on the spatial features; output the spatial audio signal along with the processed audio signal to a remote video conferencing host for remote rendering of three-dimensional immersive audio.
[0148] Among them, spatial characteristics include at least one of the following: sound source arrival direction, sound field diffusion, surround angle, and distance information.
[0149] For example, extracting the spatial features of the audio signals collected by each pickup panel 11 refers to identifying and quantifying attributes related to the sound source location and sound field environment from the audio data received by multiple pickup panels 11 using specific algorithms and analysis methods. This may include determining the relative position of the sound source by analyzing the time difference, phase difference, or amplitude difference of the same sound source signal received by different pickup panels 11, or evaluating the sound field diffusion by analyzing the reverberation characteristics of the acoustic environment. This process provides the basic data for subsequently generating audio signals with a sense of space.
[0150] For example, generating spatial audio signals based on spatial features refers to encoding or integrating the extracted spatial attribute information into the audio signal in a way that can be recognized and utilized by remote devices, thereby forming an audio data stream capable of simulating a three-dimensional sound field effect. This can be achieved by embedding spatial features as metadata into the audio stream, or through specific multi-channel encoding (such as Ambisonics encoding or object-based audio encoding), so that the signal can reproduce the spatial location of the sound source and the immersive feeling of the environment when played back.
[0151] For example, outputting the spatial audio signal along with the processed audio signal to the remote video conferencing host means that after the first host 12 completes the routine processing of the audio signal (such as noise reduction, automatic mixing, etc.), it packages the processed audio content and the generated spatial audio signal (containing spatial features) into a composite data stream and transmits it to the remote video conferencing host through the network interface. This ensures that the remote host can simultaneously obtain high-quality audio content and metadata for spatial rendering.
[0152] For example, rendering 3D immersive audio remotely refers to the process whereby a remote video conferencing host receives an audio signal containing spatial features and uses its built-in audio rendering engine, combined with the head tracking data or playback device configuration of the remote user, to generate a personalized 3D sound field. This allows the remote user's headphones or multi-channel audio system to reproduce the real acoustic environment of the conference room, enhancing the immersion and realism of the meeting.
[0153] Spatial characteristics include at least one of the following: direction of arrival (DOA), sound field diffusion, surround angle, and distance information. The direction of arrival (DOA) refers to the direction in which sound waves propagate from the sound source to the pickup panel 11, and is typically calculated precisely using signal processing algorithms of the microphone array 111 (such as MVDR, MUSIC, or TDOA-based algorithms). Sound field diffusion refers to the uniformity of energy distribution and reverberation characteristics of the sound field in space, and can be estimated by analyzing room acoustic parameters (such as reverberation time and early decay time) or by combining them with a room geometry model. The surround angle refers to the horizontal and vertical angles of the sound source relative to the listener or reference point, and can be calculated based on the direction of arrival and a preset listener position. Distance information refers to the distance from the sound source to the pickup panel 11 or reference point, and can be calculated using sound wave propagation time combined with sound speed, or estimated using a sound source intensity attenuation model.
[0154] In this embodiment, the first host 12 receives audio signals from each pickup panel 11 and utilizes the differences in these signals to extract various spatial features of the sound source, such as the direction of arrival, sound field diffusion, surround angle, and distance information. Since the microphone system 10 includes multiple pickup panels 11, and the sampling clock of each pickup panel 11 is synchronized with the system clock of the first host 12, spatial features can be extracted. Based on the extracted spatial features, the first host 12 generates a spatial audio signal containing spatial information. The first host 12 outputs the spatial audio signal along with conventionally processed (e.g., noise reduction, automatic mixing) audio signals to a remote video conferencing host. Upon receiving these signals, the remote video conferencing host can utilize the rich spatial information contained in the spatial audio signal to render three-dimensional immersive audio, thereby reproducing the realistic sound field of the meeting at the remote location and providing remote participants with an immersive auditory experience.
[0155] In one specific implementation, in a conference room equipped with multiple microphone panels 11, when a speaker speaks, each microphone panel 11 synchronously acquires its audio signal and converts it into a standard network signal, which is then sent to a first host 12. A second processing unit 122 within the first host 12, such as a digital signal processor (DSP), receives and analyzes the audio signals from all the microphone panels 11. The DSP can run advanced sound source localization algorithms, such as those based on Generalized Cross-Correlation Phase Transform (GCC-PHAT) or Multiple Signal Classification (MUSIC), to calculate the direction of arrival and distance information of the sound source by comparing the time difference (TDOA) and phase difference of the signals received by different microphone panels 11. Simultaneously, the DSP can also analyze the acoustic characteristics of the conference room and estimate the sound field diffusion. Combining these spatial characteristics, the DSP generates a spatial audio signal, which can be encoded using the Ambisonics encoding format, encoding information such as the direction, distance, and ambient reverberation of the sound source into a multi-channel audio stream. Subsequently, the first host 12 transmits the Ambisonics-encoded spatial audio signal, along with the regular audio signal processed by noise reduction and automatic mixing, to the remote video conferencing host via its network interface (e.g., an interface supporting Dante or AES67 protocols). Upon receiving these data signals, the remote video conferencing host's built-in audio rendering engine decodes and renders immersive 3D audio in real time based on the remote user's head posture (if wearing a head-tracking device) or a preset listening position. This allows the remote user to clearly perceive the speaker's specific location and movement trajectory within the conference room, thus achieving a highly realistic sense of presence.
[0156] In summary, the microphone system 10 in this embodiment solves the problems of difficult installation, large size and inconvenient maintenance caused by traditional integrated devices by physically separating the microphones. It has the advantages of easy installation, strong spatial adaptability, beautiful appearance and easy maintenance.
[0157] Secondly, a pickup panel 11 is provided, including a housing, a microphone array 111, a first processing device 113, and a first module interface M2. The thickness of the housing is no more than 10 mm. The microphone array 111 is disposed on the pickup surface of the housing for picking up audio signals. The first processing device 113 is used to convert the audio signals into standard network signals. The first module interface M2 is used to connect a transmission cable 13 to transmit the standard network signals to an external host.
[0158] In this embodiment, the microphone panel 11 is physically separated from the external host and is connected to the external host via a transmission cable 13.
[0159] Through the above technical solutions, the microphone system 10 in this embodiment achieves physical separation of sound pickup and processing functions, making the pickup panel 11 extremely thin and light, easy to install, and not affecting the aesthetics of the room. Furthermore, based on the physical separation between the pickup panel 11 and the first host 12, the first host 12 can be flexibly deployed in a concealed location (such as a low-voltage box or above the ceiling 20), eliminating the dependence on the installation space inside the ceiling 20, and facilitating maintenance and upgrades. On the other hand, the pickup panel 11 and the first host 12 adopt a standard network signal transmission method, realizing the standardization and efficiency of data transmission. Compared with the problems of poor anti-interference and complex wiring that traditional analog signal transmission may face, this provides a more stable and simpler technical solution. It can also achieve simultaneous transmission of data and power through a single cable, further simplifying wiring work and improving the versatility and reliability of the system.
[0160] Thirdly, an audio processing method is provided for use in the microphone system 10 described above. Figure 14 This is a schematic flowchart of an audio processing method provided in an exemplary embodiment of this application. (Reference) Figure 14 The audio processing method includes the following steps: Step S901: Pick up the audio signal through the pickup panel 11 and convert the audio signal into a standard network signal.
[0161] Step S902: Transmit the standard network signal to the first host 12 via the transmission cable 13.
[0162] Step S903: Receive and process standard network signals through the first host 12 to obtain an output signal.
[0163] In steps S901 to S903, the pickup panel 11, transmission cable 13 and first host 12 mentioned are the pickup panel 11, transmission cable 13 and first host 12 described above. Their specific structures have been described above and will not be repeated here.
[0164] In steps S901 and S902, the audio signal has already been converted into a standard network signal in step S901, while in step S902, the standard network signal is transmitted to the first host 12 through the transmission cable 13. Therefore, the transmission cable 13 no longer transmits a weak analog signal that is susceptible to interference, but a standard network signal with a certain degree of anti-interference capability. In this way, the first host 12 can be deployed in a location far away from the pickup panel 11, such as in a computer room or a concealed cabinet, while the pickup panel 11 can be flexibly deployed in the optimal pickup position near the sound source, allowing the pickup panel 11 to be far away from the first host 12 for flexible arrangement.
[0165] In step S903, the first host 12, as the core processing unit of the system, can receive and process standard network signals to obtain an output signal. Since the transmission cable 13 in step S902 transmits standard network signals, even if the transmission cable 13 passes through a complex electromagnetic environment, as long as the network link is normal, the audio data can be transmitted without loss. The standard network signal received by the first host 12 is logically consistent with the standard network signal output by the pickup panel 11, thereby avoiding the signal attenuation and sound quality degradation problems caused by long-distance analog transmission.
[0166] Through the above technical solution, the pickup panel 11 and the first host 12 adopt the standard network signal transmission method, which realizes the standardization and efficiency of data transmission. Compared with the problems of poor anti-interference and complex wiring that may be faced by traditional analog signal transmission, it provides a more stable and simpler technical solution. It can also realize the simultaneous transmission of data and power through a single cable, further simplifying the wiring work and improving the versatility and reliability of the system.
[0167] It should be noted that the order of steps S901 to S903 in the embodiments of this application is not limited to the above steps, and can be repeated repeatedly as needed, without limitation.
[0168] In some embodiments, step S901 can also be implemented through the following steps: Step S9011: Audio signals from the same sound source are simultaneously acquired through at least two pickup panels 11, and the sampling clock of each pickup panel 11 is synchronized with the system clock of the first host 12.
[0169] Step S903 can also be achieved through the following steps: Step S9031: Receive the standard network signal sent by each pickup panel 11 through the first host 12, and extract the audio signal collected by each pickup panel 11.
[0170] Step S9032: The first host 12 calculates the time difference between the audio signals collected by each pickup panel 11, and determines the location information of the sound source based on the time difference.
[0171] Through the aforementioned steps S9011, S9031, and S9032, this embodiment of the application utilizes the layout of the multi-pickup panel 11 combined with a high-precision clock synchronization technology to achieve real-time capture of the sound source location. This not only provides a visual display of the speaker's location but also offers crucial spatial parameters for potential beamforming, automatic tracking, and other functions, thereby enhancing the overall performance of the audio processing system.
[0172] In another feasible implementation, step 901 can also be implemented by the following steps: Step S9012: Acquire the audio signal of the target sound source synchronously through at least two pickup panels 11, and keep the sampling clock of each pickup panel 11 synchronized with the system clock of the first host 12.
[0173] In another feasible implementation, step S903 can also be implemented by the following steps: Step S9033: Receive the standard network signal sent by each pickup panel 11 through the first host 12, and extract the audio signal collected by each pickup panel 11.
[0174] Step S9034: The first host 12 dynamically adjusts the weighting coefficients and delay parameters of the audio signals collected by each pickup panel 11 according to the location information of the target sound source, so as to form a virtual beam pointing to the target sound source.
[0175] It should be understood that a virtual beam is not a physically existing directional transmission and reception area, but rather a digitally simulated directional characteristic similar to that of a radar or directional microphone, achieved through algorithms. Specifically, the DSP module inside the first host 12 executes beamforming algorithms (such as delay summation algorithms or adaptive filtering algorithms).
[0176] It should be understood that the adjustment of the delay parameter is based on phase compensation according to the sound path difference. Since the target sound source arrives at each pickup panel 11 at different distances, the arrival times of the sound waves at each pickup panel 11 differ. If the signals from each pickup panel 11 are directly added together, they may cancel each other out due to phase inconsistency. Therefore, the first host 12 calculates the time difference of the sound waves arriving at each pickup panel 11 based on the known coordinates of the sound source and the coordinates of each pickup panel 11, and applies corresponding delay compensation to each audio signal accordingly. For example, if the sound source is closer to pickup panel 11A and farther from pickup panel 11B, the system applies a certain delay to the signal from pickup panel 11A to align it with the signal from panel B in time. After the delay processing, the signal components from the target sound source are in phase in each channel, and the energy is significantly enhanced after superposition.
[0177] It should be understood that the purpose of adjusting the weighting coefficients is to optimize the beam shape and suppress sidelobe interference. The first host 12 can multiply the signal of each channel by different coefficients based on the signal-to-noise ratio of each pickup panel 11 or a preset weighting algorithm (such as Chebyshev weighting or adaptive weighting). For example, a larger weight is assigned to a panel that is closer to the sound source and has a higher signal strength; a smaller weight is assigned to a panel that is farther away or obstructed. By dynamically adjusting the weighting coefficients, the main lobe of the synthesized beam (i.e., the high-sensitivity region) can be accurately pointed to the target sound source, while a "null" (i.e., a low-sensitivity region) is formed in the direction of the interference source (such as an air conditioner vent or a projector position), thereby greatly filtering out environmental noise.
[0178] Through the aforementioned steps S9012, S9033, and S39034, a highly directional "virtual microphone" is constructed based on the physically omnidirectional pickup panel 11 array. The virtual beam can not only flexibly track moving sound sources but also adjust its directional characteristics in real time according to the distribution of ambient noise, solving the problem of traditional physical microphones having fixed directionality and being unable to cope with complex and changing sound field environments, thus significantly improving the clarity and signal-to-noise ratio of the pickup.
[0179] Fourthly, an audio processing method is provided for use in the microphone system 10 described above. Figure 15 This is another flowchart illustrating the audio processing method provided in an exemplary embodiment of this application. Please refer to... Figure 15 The audio processing method includes the following steps: Step S1001: Audio signals are collected through multiple pickup panels 11, and the sampling clock of each pickup panel 11 is synchronized with the system clock of the first host 12.
[0180] Step S1002: Receive standard network signals sent by each pickup panel 11 through the first host 12, obtain the audio signal quality or sound source location of each pickup panel 11, dynamically select a portion of the pickup panels 11 to form an effective pickup array, and adjust the signal weight of the selected pickup panels 11.
[0181] The audio processing method described in the fourth aspect of this application focuses on the flexible expansion and adaptive array management of the microphone system 10. In large conference halls or distributed sound pickup scenarios, a large number of pickup panels 11 are often deployed, but not all pickup panels 11 need to be in full-power operation at all times. This application embodiment achieves a balance between optimal sound pickup effect and system resources through the dynamic scheduling of multiple pickup panels 11 by the first host 12.
[0182] Through steps S101 and S102 described above, the microphone system 10 overcomes the limitations of fixed arrays in traditional multi-microphone systems 10. In dynamically changing conference environments, such as speaker position changes, the appearance of local noise, or conference area reconstruction, the first host 12 can identify and select the pickup panel 11 most suitable for the current acoustic conditions, thus forming an "effective pickup array." This technical solution improves the focus and clarity of audio acquisition, effectively suppresses unnecessary environmental noise and reverberation, and avoids redundant information and processing burdens that might be introduced by using all panels in a fixed manner. As a result, the microphone system 10 can provide higher quality output signals, ensuring that remote participants receive a clearer and more professional listening experience. Simultaneously, it optimizes the resource utilization of the microphone system 10, enabling it to exhibit good adaptability and performance in various complex and dynamic application scenarios.
[0183] Fifthly, a computer storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implements the audio processing method described above.
[0184] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in any of the above embodiments, such as: picking up an audio signal through a pickup panel 11 and converting the audio signal into a standard network signal; transmitting the standard network signal to a first host 12 through a transmission cable 13; and receiving and processing the standard network signal through the first host 12 to obtain an output signal.
[0185] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0186] Since the computer program or instructions stored in the computer-readable storage medium can execute the steps in any of the above method embodiments provided in the embodiments of this application, the beneficial effects that the methods described in any of the above method embodiments can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0187] It should be noted that, for the methods of the embodiments of this application, those skilled in the art will understand that all or part of the processes of the methods of the embodiments of this application can be implemented by a computer program controlling related hardware. This computer program can be stored in a computer-readable storage medium, such as in the memory of an electronic device, and executed by at least one processor within the electronic device. During execution, it can include the processes of the embodiments of the method. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, etc.
[0188] In a sixth aspect, a control device is provided, on which a computer program or instructions are stored, which, when executed by a processor, implement the steps of the audio processing method described above.
[0189] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.
[0190] It should be noted that, in the data processing stage, the technical solution of this application has strictly limited the scope of data collection to the minimum necessary to achieve the technical objectives, preventing the acquisition of irrelevant information. For any user information to be collected, the data subject will be clearly informed and their consent obtained. Furthermore, technologies such as encrypted storage and access control are employed to strengthen data security and ensure the security and compliance of the entire data processing process. The technical model and decision-making mechanism are based on objective technical parameters and do not introduce unnecessary parameters such as gender or age that may lead to discrimination, resolutely eliminating algorithmic discrimination and upholding public order and good morals. In addition, the specification fully describes the technical implementation methods, application scenarios, and compliance protection details. The claims are consistent with the content of the specification, key compliance designs are clear and verifiable, and the overall technical design is guided by the protection of public interests and adherence to social ethics, without any circumstances that harm public interests or violate public order and good morals.
[0191] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0192] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0193] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.
[0194] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.
Claims
1. A microphone system (10), characterized in that, It includes at least one pickup panel (11), a first host (12), and a transmission cable (13) connecting the pickup panel (11) and the first host (12), wherein the pickup panel (11) and the first host (12) are physically separated. At least one of the microphone panels (11) is used to pick up audio signals and convert the audio signals into standard network signals and send them to the first host (12) through the transmission cable (13). The transmission cable (13) is used to transmit the standard network signal from the pickup panel (11) to the first host (12). The first host (12) is used to receive and process the standard network signal to obtain an output signal.
2. The microphone system (10) according to claim 1, characterized in that, The first host (12) includes a second processing device (122); The second processing device (122) is used to send a clock synchronization signal to the pickup panel (11) so that the sampling clock of each pickup panel (11) is synchronized with the system clock of the first host (12), and to receive and process the standard network signal to obtain an output signal.
3. The microphone system (10) according to claim 2, characterized in that, The pickup panel (11) includes: The audio acquisition module (11a) is used to pick up audio signals and convert them into digital signals; The network conversion module (11b) is used to convert the digital signal into a standard network signal.
4. The microphone system (10) according to claim 3, characterized in that, The audio acquisition module (11a) includes a microphone array (111) and an analog-to-digital conversion module (112), and the network conversion module (11b) includes a first processing device (113) and a first physical layer device (114). The microphone array (111) is used to pick up the audio signal; The analog-to-digital converter module (112) is electrically connected to the microphone array (111) and is used to convert the audio signal into a first digital signal; The first processing device (113) is used to package the first digital signal into a first encapsulation frame signal; and The first physical layer device (114) is used to convert the first encapsulated frame signal into a standard network signal and send it to the first host (12) via a transmission cable. The analog-to-digital conversion module (112) is connected between the microphone array (111) and the first processing device (113), and the first processing device (113) is connected between the analog-to-digital conversion module (112) and the first physical layer device (114).
5. The microphone system (10) according to claim 4, characterized in that, The first processing device (113) includes: An audio serial interface acquisition controller (1131) is used to receive the first digital signal transmitted by the analog-to-digital conversion module (112); The media access controller (1132), connected to the audio serial interface acquisition controller (1131), is used to package the received first digital signal into a first encapsulation frame signal; A custom packet sending module (1133) is connected to the media access controller (1132) and is used to control the timing of the encapsulation and transmission of the first encapsulated frame signal of the media access controller (1132); The precision time protocol program module (1134) is used to execute the precision time protocol to synchronize the clock of the pickup panel (11) with the clock of the first host (12).
6. The microphone system (10) according to claim 5, characterized in that, The media access controller (1132) packages data packets according to a preset data packet length, which corresponds to the sampling period of the analog-to-digital conversion module (112).
7. The microphone system (10) according to claim 2, characterized in that, The first host (12) also includes: The second physical layer device (121) is used to convert the standard network signal into a second encapsulated frame signal; The second processing device (122) is connected to the second physical layer device (121) and is used to receive the second encapsulation frame signal and perform processing operations to obtain an output signal.
8. The microphone system (10) according to claim 7, characterized in that, The second processing device (122) includes: The first processor (1221) is used to receive the second encapsulation frame signal and convert the second encapsulation frame signal into a second digital signal.
9. The microphone system (10) according to claim 8, characterized in that, The first processor (1221) is also used to perform processing operations on the second digital signal to obtain a third digital signal; The processing operation includes at least one of the following: Noise reduction, automatic mixing, echo cancellation, and audio enhancement.
10. The microphone system (10) according to claim 9, characterized in that, The first host (12) also includes a second processor (123) and at least one output port (M1). The second processor (123) is used to receive the third digital signal and distribute it to the corresponding output port (M1).
11. The microphone system (10) according to claim 10, characterized in that, The second processor (123) is also used to manage the operating state of the microphone system (10) and to perform parameter configuration of the microphone system (10).
12. The microphone system (10) according to any one of claims 1 to 11, characterized in that, The pickup panel (11) is 3 mm to 10 mm thick and is configured to be mounted on a ceiling (20) or a desktop.
13. The microphone system (10) according to any one of claims 1 to 9, characterized in that, The pickup panel (11) also includes a first module interface (M2), the first host (12) also includes a second module interface (M3), and the transmission cable (13) is connected between the first module interface (M2) and the second module interface (M3).
14. The microphone system (10) according to claim 4, characterized in that, The first host (12) also includes: Multiple second module interfaces (M3) are used to connect the corresponding transmission cables (13); A switch (124) is connected between the second module interface (M3) and the second processing device (122) so that each transmission cable (13) is connected to the same switch (124).
15. The microphone system (10) according to any one of claims 1 to 14, characterized in that, The pickup panel (11) also integrates a speaker unit (115). The first host (12) is also used to: receive external audio signals, convert the external audio signals into standard network signals, and send them to the pickup panel (11) through the transmission cable (13). The pickup panel (11) is also used to: receive the standard network signal through the transmission cable (13), convert it into an analog audio signal, and drive the speaker unit (115) to play it.
16. The microphone system (10) according to claim 15, characterized in that, The pickup panel (11) also includes: A signal switching module (116) is used to selectively transmit the audio signal picked up by the microphone array (111) or the external audio signal received by the transmission cable (13) to the speaker unit (115); or The audio mixing module (117) is used to mix the audio signal picked up by the microphone array (111) with the external audio signal received by the transmission cable (13) and then transmit it to the speaker unit (115).
17. The microphone system (10) according to any one of claims 1 to 16, characterized in that, The microphone system (10) includes a plurality of the pickup panels (11); The first host (12) is also used to: extract the audio signals collected by each pickup panel (11) according to the standard network signals sent by each pickup panel (11), and dynamically select a portion of the pickup panels (11) to form an effective pickup array.
18. The microphone system (10) according to any one of claims 1 to 16, characterized in that, The first host (12) is also used for: Extract the spatial features of the audio signals collected by each pickup panel (11), and generate a spatial audio signal based on the spatial features; The spatial audio signal, along with the processed audio signal, is output to the remote video conferencing host for remote rendering of 3D immersive audio. The spatial features include at least one of the following: sound source arrival direction, sound field diffusion, surround angle, and distance information.
19. A pickup panel (11), characterized in that, include: The shell thickness is no more than 10mm; A microphone array (111) is disposed on the pickup surface of the housing for picking up audio signals; A first processing device (113) is used to convert the audio signal into a standard network signal; and The first module interface (M2) is used to connect the transmission cable (13) to transmit the standard network signal to the external host; wherein the microphone panel (11) is physically separated from the external host and is connected to the external host through the transmission cable (13).
20. An audio processing method, applied to the microphone system according to any one of claims 1-16, characterized in that, include: Audio signals are picked up through a microphone panel and converted into standard network signals. The standard network signal is transmitted to the first host via a transmission cable; The first host receives and processes the standard network signal to obtain the output signal.
21. The audio processing method according to claim 20, characterized in that: The step of picking up audio signals through a microphone panel and converting the audio signals into standard network signals includes: Audio signals from the same sound source are acquired synchronously by at least two pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The step of receiving and processing the standard network signal through the first host to obtain the output signal includes: The first host receives standard network signals sent by each microphone panel and extracts audio signals collected by each microphone panel. The first host calculates the time difference between the audio signals collected by each pickup panel, and determines the location information of the sound source based on the time difference.
22. The audio processing method according to claim 20, characterized in that: The step of picking up audio signals through a microphone panel and converting the audio signals into standard network signals includes: The audio signal of the target sound source is acquired synchronously by at least two pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The step of receiving and processing the standard network signal through the first host to obtain the output signal includes: The first host receives standard network signals sent by each microphone panel and extracts audio signals collected by each microphone panel. The first host dynamically adjusts the weighting coefficients and delay parameters of the audio signals collected by each pickup panel according to the location information of the target sound source, thereby forming a virtual beam pointing towards the target sound source.
23. An audio processing method applied to the microphone system of claim 17, characterized in that, include: Audio signals are collected by multiple pickup panels, and the sampling clock of each pickup panel is synchronized with the system clock of the first host. The first host receives standard network signals sent by each pickup panel, obtains the audio signal quality or sound source location of each pickup panel, dynamically selects a portion of the pickup panels to form an effective pickup array, and adjusts the signal weight of the selected pickup panels.