Data transmission system and method and storage medium
By introducing network switching equipment for path management in medium and large-scale video conferencing systems, the scalability and stability issues of audio data transmission under multi-microphone deployments are resolved, achieving efficient audio data transmission and high-quality audio processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, medium and large-scale video conferencing systems struggle to achieve stable sound pickup, audio signal consistency, and system scalability when deploying multiple microphones. This is mainly due to the lack of intelligent path management and forwarding control capabilities in branch nodes, leading to accumulated latency, insufficient synchronization accuracy, and rigid topology.
By introducing network switching equipment as branch nodes in the cascaded link for path management and forwarding control, a cascaded audio transmission system consisting of microphone equipment, network switching equipment, and audio processing equipment is constructed to achieve intelligent path management and forwarding of audio data, ensuring data integrity and independence.
It improves the scalability and stability of audio data transmission, supports cross-regional collaborative sound pickup, high-quality speech enhancement and precise echo suppression, and enhances the system's synchronization accuracy and overall processing efficiency.
Smart Images

Figure CN122053542A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio transmission technology, specifically to a data transmission system, method, and storage medium. Background Technology
[0002] In modern medium to large-scale video conferencing environments, audio pickup quality has a decisive impact on the clarity of voice communication and meeting efficiency. With the continuous expansion of conference room physical spaces, the increasing complexity of participant seating arrangements, and the higher demands for audio fidelity in remote collaboration, single-point or limited-microphone pickup solutions in related technologies can no longer meet the stringent requirements of stable audio pickup performance, audio signal consistency, and system scalability in medium to large-scale conference scenarios.
[0003] To cover larger meeting spaces and support multiple independent pickup areas, related technologies typically require the deployment of more microphone devices and centralized aggregation and processing of audio data collected from each microphone. However, in large-scale deployments of multiple microphones, these technologies often rely on cascaded aggregation devices (such as hubs) with audio energy optimization capabilities or simple cascading architectures. While branch nodes can compare the energy of multiple received audio signals and select the optimal path for uploading, they discard the original audio data from other channels and only forward a single signal upwards. This makes it impossible to independently identify, completely preserve, and flexibly schedule multiple audio streams. Consequently, it is difficult to achieve collaborative work between different pickup areas while ensuring the real-time performance and integrity of audio data, especially when it is necessary to simultaneously support advanced audio functions such as cross-area collaborative pickup, sound source localization, high-quality speech enhancement, and precise echo cancellation. Microphone networks in related technologies still have shortcomings in terms of scalability, time synchronization accuracy, data transmission link reliability, and overall architecture processing efficiency. The root cause lies in the lack of intelligent path management and forwarding control capabilities of branch nodes for audio data streams, making it difficult to simultaneously meet the requirements of large-scale deployment and high-quality audio processing. Summary of the Invention
[0004] This application provides a data transmission system, method, and storage medium, aimed at improving the scalability of audio data transmission.
[0005] In a first aspect, this application provides a data transmission system, the system comprising: At least one microphone device, at least one network switching device, and an audio processing device; The at least one microphone device is connected to the network switching device in a cascaded manner to form at least one cascaded link; The microphone device is used to collect local audio data and transmit the audio data upstream in the cascaded link; The network switching device is used to receive audio data streams from one or more downstream devices and transmit the audio data streams upstream of the cascaded link. The network switching device is used as a branch node in the cascaded link to perform path management and forwarding control of audio data streams from different downstream devices. The audio processing device is connected to the upstream end of the at least one cascaded link and is used to receive and process the converged audio data.
[0006] Secondly, this application also provides a data transmission method applied to the aforementioned audio data transmission system; the method includes: Local audio data is collected using a microphone device and transmitted upstream in the cascaded link. The network switching device receives audio data streams from one or more downstream devices and performs path management and forwarding control on the audio data streams to transmit the audio data streams upstream of the cascaded link. The audio data aggregated from the cascaded link is processed by an audio processing device.
[0007] Thirdly, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods described above.
[0008] This application creatively introduces a cascaded audio transmission system composed of microphone devices, network switching devices, and audio processing devices. The network switching device, possessing intelligent path management and traffic scheduling capabilities, is integrated into the hub-and-spoke-based audio cascade architecture of related technologies, serving as a core branch and aggregation node. This design enables audio data streams from multiple downstream microphones or sub-links to be efficiently, reliably, and with low latency aggregated and transmitted to upstream processing devices based on the address identification and independent forwarding mechanism of the network switching device, without broadcast conflicts or relying on stepwise mixing compression. Therefore, when supporting large-scale, multi-branch complex deployment topologies, the system not only ensures the integrity and independence of the original audio data of each channel, laying the data foundation for high-precision acoustic processing at the backend (e.g., sound source localization and speech enhancement), but also significantly improves the stability of the entire audio data transmission link and the synchronization accuracy of global devices through the deterministic transmission mechanism provided by the network switching device. This allows for better implementation of advanced audio functions such as cross-regional collaborative sound pickup, high-quality speech enhancement, and precise echo suppression in practical applications. Attached Figure Description
[0009] Figure 1 A schematic diagram of the structure of a data transmission system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the cascaded topology of the microphone devices provided in the embodiments of this application; Figure 3 This is a schematic diagram of the chain topology of the microphone device provided in the embodiments of this application; Figure 4 This is a schematic diagram of a star topology for a microphone device and a network switching device provided in an embodiment of this application; Figure 5 A schematic diagram of a star-link hybrid topology for a microphone device and a network switching device provided in an embodiment of this application; Figure 6 This is a schematic diagram of a hierarchical aggregation or mixing strategy provided in an embodiment of this application; Figure 7 This is a schematic diagram of a centralized aggregation or mixing strategy provided in an embodiment of this application; Figure 8 A schematic flowchart of a data transmission method provided in an embodiment of this application; Figure 9 This is a schematic diagram of the visual management interface of the data transmission system provided in the embodiments of this application; Figure 10 This is another schematic diagram of the data transmission system provided in the embodiments of this application; Figure 11 This is a schematic diagram of a microphone device visual management interface provided in an embodiment of this application; Figure 12 This is another schematic diagram of the data transmission system provided in the embodiments of this application; Figure 13 This is a schematic diagram of a hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] This application provides a data transmission system, method, and storage medium. Through a system including a microphone device 101, a network switching device 102, and an audio processing device 103, it solves the problem of efficient cascading transmission of audio data in a multi-microphone system, and how to effectively aggregate and transmit the audio data collected by each microphone to the audio processing device 103. In large-scale deployment scenarios, the related solutions lack scalability, and the use of related branch nodes (e.g., a central processing unit) may result in insufficient timeliness, large transmission delays, and difficulty in supporting complex topologies.
[0012] In the field of medium to large-scale video conferencing, the industry's long-term focus on technological evolution and optimization has generally been on the pickup performance of individual microphone units, the efficiency of audio codec algorithms, or local improvements within a given network topology (e.g., a star topology based on a cascaded aggregation device with audio energy optimization capabilities (referred to as a Hub in this paper), or a simple daisy chain). Faced with problems such as increased latency and decreased stability during system expansion, the industry typically attributes these to external factors such as cable quality, device clock accuracy, or network bandwidth, and seeks solutions along these lines. This mindset has led to a common perception that in large-scale, distributed microphone arrays, the conflict between audio data aggregation latency and system scale, as well as the rigidity of the topology, are inherent properties of such systems that must be compromised. For example, while the Hub can compare the energy of multiple received audio signals and select the optimal path for uploading, it discards the original audio data of other channels at branch nodes, only forwarding a single signal upwards, thus failing to preserve the integrity of multiple channels.
[0013] Through in-depth analysis, the inventors of this application discovered that the aforementioned common understanding may have obscured the true root of the problem. The inventors observed that whether it's a hub employing a single-path optimization mechanism or a serial link leading to linear latency accumulation, the core flaw points to the same deep structural problem: the "branch node" in the relevant architecture, responsible for data aggregation and forwarding, is essentially a limited or simple merging point; it lacks the ability to intelligently identify, completely preserve, and flexibly schedule data flows. This deficiency in the capabilities of this critical node fundamentally restricts the determinism of data transmission, the flexibility of the topology, and the integrity of the data itself when the entire system expands.
[0014] However, the problem's insidious nature lies in the fact that it's not simply an audio technology issue, but rather a data network architecture problem hidden within audio application scenarios. In the field of audio system design, engineers are accustomed to seeking answers within the realms of audio and video protocols, acoustic devices, and dedicated audio processing chips. "Network switching equipment" is typically considered information technology (IT) infrastructure, belonging to another technological field used for general data transmission. Its profound connection to the high real-time, multi-channel concurrent professional audio streaming has not been established. It is precisely this gap between fields that has prevented the fundamental solution—"introducing network devices with intelligent addressing and switching capabilities to reshape the core architecture of audio cascaded networks"—from being recognized and attempted by those skilled in the art for a long time.
[0015] In view of this, the embodiments of this application aim to provide a cascaded audio transmission system and method based on network switching equipment, which can solve the problems of limited system topology flexibility, accumulated transmission delay, insufficient synchronization accuracy and rigid topology structure caused by the lack of intelligent scheduling capability of branch nodes in the prior art.
[0016] like Figure 1 As shown, Figure 1 This is a schematic diagram of a data transmission system provided in an embodiment of this application. The system includes at least one microphone device 101, at least one network switching device 102, and an audio processing device 103. The at least one microphone device 101 and the network switching device 102 are cascaded to form at least one cascaded link. The microphone device 101 is used to collect local audio data and transmit the audio data upstream of the cascaded link. The network switching device 102 is used to receive audio data streams from one or more downstream devices and transmit the audio data streams upstream of the cascaded link. The network switching device 102 acts as a branch node in the cascaded link, performing path management and forwarding control on audio data streams from different downstream devices. The audio processing device 103 is connected to the upstream of at least one cascaded link and is used to receive and process the converged audio data.
[0017] The system includes multiple nodes, each equipped with a microphone device (mic) 101, a network switch device 102, or an audio processing device 103. The microphone device 101 can be a device for converting acoustic signals into electrical signals, acquiring initial audio data, and transmitting the audio data to upstream devices (e.g., upstream microphone device 101 or upstream network switch device 102) or audio processing device 103. The microphone device 101 can be a single pickup unit or an array integrating multiple pickup elements. The network switch device 102 can be a device that implements packet forwarding in the network. For example, the network switch device 102 can be a switch (SW), acting as a branch node in a cascaded link, receiving audio data from one or more downstream microphone devices 101 or network switch devices 102, and forwarding the audio data to upstream devices or audio processing device 103. The network switch device 102 aggregates and distributes the audio data. Cascaded connection refers to a connection method in which multiple devices are connected together in series or in a tree structure to form a data transmission chain, allowing data to be transmitted step by step. A cascaded link refers to a complete data transmission path formed by microphone device 101 and network switching device 102 connected in a cascading manner.
[0018] Audio data refers to the raw digital audio signal directly acquired by microphone device 101. The upstream direction refers to the direction in which data is transmitted from the end device to audio processing device 103 in the cascaded link. The downstream direction refers to the direction in which data is transmitted from audio processing device 103 to the end device in the cascaded link. An audio data stream refers to a continuous, real-time transmitted sequence of audio data.
[0019] The audio processing device 103 can be a central control unit for receiving, storing, analyzing, and processing audio data. In this application, the audio processing device 103 is used to receive audio data from microphone devices 101 and / or network switching devices 102, and can further process, store, or transmit the audio data. The audio processing device 103 can be a dedicated audio processor, server, or computer with corresponding processing capabilities. For example, the audio processing device 103 can be a central processing unit, a central audio processing device 103, or a central processing unit. The aggregated audio data refers to all audio data collected by multiple microphone devices 101 in the cascaded link, forwarded by network switching devices 102, and finally concentrated at the audio processing device 103.
[0020] As an example, the system includes upstream network devices and several microphone devices 101, which are installed on nodes of the system to form microphone nodes. The upstream network devices can be Ethernet switching devices (e.g., network switch 102) supporting the Precision Time Protocol (PTP) and a central processing unit, responsible for clock source distribution, control and monitoring, and necessary audio aggregation. Network switch 102 can be a branch point, installed on nodes of the system to form branch nodes. Where an upstream microphone node connects to multiple downstream microphone nodes, the upstream microphone node can also be a branch node.
[0021] The data transmission system of this application constructs a cascaded data transmission network consisting of a microphone device 101, a network switching device 102, and an audio processing device 103. This allows audio data collected by multiple microphone devices to be aggregated to the upstream device through branch nodes, thereby improving the system's scalability and transmission stability when transmitting audio data in medium to large-scale conference scenarios.
[0022] In some embodiments, the audio processing device 103 is further configured to: automatically initiate a device discovery process when the system is powered on or a new device is detected to identify the microphone device 101 and the network switching device 102 in the network; construct and display a physical topology map based on the connection relationship information reported by each device, wherein the physical topology map represents the hierarchical relationship of at least one cascaded link; and monitor the working status of each device in real time, wherein the working status includes at least one of power supply method, audio signal quality, link delay or device temperature.
[0023] Through the above technical solution, the system of this application can realize automated management and real-time monitoring of microphone device 101 and network switching device 102 in the cascaded link. When the system is powered on or a new device is connected, the audio processing device 103 can automatically initiate a device discovery process, identifying all devices in the network without manual intervention, thereby significantly simplifying the system deployment and configuration process. Based on the connection relationship information reported by the devices, the system can construct and display an intuitive physical topology diagram (e.g., a tree diagram) to represent the hierarchical relationship of the cascaded links, greatly improving the administrator's efficiency in understanding the network structure.
[0024] Meanwhile, the audio processing device 103 can monitor the key operating status of each device in real time, including power supply mode, audio signal quality, link latency and device temperature, so that the system can detect and respond to potential abnormal situations in a timely manner, such as power interruption, audio signal degradation or device overheating, effectively solving the problems of complex device configuration and insufficient status monitoring in related multi-microphone systems.
[0025] In some embodiments, the microphone device 101 is further configured to receive first audio data sent by a downstream microphone device 101, process the first audio data according to a preset strategy to obtain second audio data, and transmit the second audio data upstream of the cascaded link; the network switching device 102 is further configured to receive a second audio data stream from one or more downstream devices and transmit the second audio data stream upstream of the cascaded link; the audio processing device 103 is further configured to issue a preset strategy to the cascaded link based on system configuration or environmental status.
[0026] The microphone device 101 processes the received first audio data according to a preset strategy to obtain second audio data. The preset strategy can be configured according to specific application scenarios; for example, it can involve filtering, merging, compressing, noise suppression, echo cancellation, or feature extraction of the first audio data. Through this processing, the generated second audio data is typically more concise, optimized, or information-rich than the original first audio data. The microphone device 101 sends the obtained second audio data to upstream microphone devices 101, upstream network switching devices 102, and / or audio processing devices 103, thus continuing the upstream transmission of the locally pre-processed data. The network switching device 102 can receive and efficiently forward the pre-processed second audio data, ensuring that the second audio data converges upstream along the correct path. The audio processing device 103 no longer needs to process all the original first audio data; instead, it receives the data already processed by the microphone device 101, reducing its own processing burden. As the control center of the entire system, the audio processing device 103 can dynamically issue different preset strategies to the microphone devices 101 in the cascaded link based on the current system configuration (e.g., the meeting mode or speech mode selected by the user) or the real-time environmental status (e.g., the ambient noise detected by the sensor, the activity level of the pickup area, etc.).
[0027] In some embodiments, microphone device 101 includes a first port for receiving Ethernet power and / or data, and a second port for providing Ethernet power and / or data; the second port of upstream microphone device 101 is connected to the first port of downstream microphone device 101; network switching device 102 includes a third port for receiving Ethernet data and at least one fourth port for providing Ethernet power and / or data; the third port is connected to the second port of upstream microphone device 101, and the fourth port is connected to the first port of downstream microphone; wherein, the third port of network switching device 102 is also used to connect to the fourth port of upstream network switching device 102.
[0028] Among them, such as Figure 2 As shown, Figure 2This is a schematic diagram of the cascaded topology of the microphone devices provided in this application embodiment. The first port can be a network port with integrated Power over Ethernet (PoE) Powered Device (PD) functionality, denoted as port A. The first port can draw power from the connected Ethernet cable and simultaneously receive Ethernet data. The first port serves as the data and power input terminal for the microphone device 101 in the connection topology. The second port can be a network port with integrated PoE Power Sourcing Equipment (PSE) functionality, denoted as port B. The second port can provide power to connected downstream devices and forward Ethernet data. The second port serves as the data and power output terminal for the microphone device 101, enabling cascading between the microphone device 101 and / or the network switching device 102. Figure 3 As shown, Figure 3 This is a schematic diagram of a chain topology for microphone devices provided in an embodiment of this application. Through the connection between the second port of the upstream microphone device 101 and the first port of the downstream microphone device 101, single-wire transmission of power and data is achieved between the microphone devices 101, thus constructing a chain topology. Some network switching device nodes 102 or microphone devices 101 support wireless backhaul (e.g., Wi-Fi 6 / 7, 5G private networks). The system forms a hybrid flexible topology with a wired backbone and wireless access, suitable for scenarios where cabling is difficult.
[0029] As an example, each microphone device 101 integrates two Ethernet ports (port A first, port B second) and a DC input interface. Port A acts as a powered device, receiving Power over Ethernet and handling uplink data, while port B acts as a power supply device, providing PoE to downstream devices and handling downlink data. The DC input interface can receive external DC power. The microphone device 101 internally includes a power control component (hardware switch) for selecting between PoE and DC power.
[0030] The third port is an Ethernet data receiving port, used to receive Ethernet data from upstream devices. This port may not have PoE power supply capability or may only serve as a data input port. The fourth port is a network port with integrated PoE PSE functionality. This port can provide power to connected downstream devices and forward Ethernet data. For example... Figure 4 As shown, Figure 4 This is a schematic diagram of a star topology for a microphone device and a network switching device provided in an embodiment of this application. It is configured with multiple fourth ports to form multiple branches. Each branch connects to multiple downstream microphone devices 101 or downstream network switching devices 102, thereby constructing a star topology.
[0031] The third port connects to the second port of the upstream microphone device 101, enabling the network switch 102 to receive data from the upstream microphone device 101. The fourth port connects to the first port of the downstream microphone device 101, enabling the network switch 102 to provide power and data to the downstream microphone device 101. The third port also connects to the fourth port of the upstream network switch 102, enabling cascading of the network switches 102 to further expand the network coverage and connectivity, accommodating larger-scale audio acquisition system deployments. It should be noted that, as Figure 5 As shown, Figure 5 This is a schematic diagram of a star-chain hybrid topology for microphone devices and network switching devices provided in this application embodiment. In terms of power supply and topology strategy, the first microphone device 101 or the microphone device 101 at the branch starting point can be powered by PoE or DC power from the network switching device 102. Downstream microphone devices 101 obtain PoE step by step through port B to form a linear trunk. When coverage needs to be extended, additional branches are opened at the microphone device 101 through the network switching device 102 to form a star-chain hybrid topology.
[0032] As an example, multiple microphone devices 101 connect to an upstream or next-level node via port A and supply power and forward data to the next-level node via port B. Links can be cascaded sequentially along table edges, ceilings, or walls to form a trunk. When larger areas need coverage, several nodes branch out in two or more directions to form a star-chain hybrid topology. This topology achieves wide-area coverage with minimal loopbacks in environments with strong cabling constraints, while retaining the flexibility for branch-based zone management and processing.
[0033] In some embodiments, the microphone device 101 further includes a DC input interface and a power supply control component; wherein the microphone device 101 is configured to receive Ethernet power and data from upstream through its network port; the DC input interface is used to connect to an external DC power supply; the power supply control component is used to control the microphone device 101 to use Ethernet power when the Ethernet power supply meets a first preset condition; and to control the microphone device 101 to use the DC input interface to receive power from the external DC power supply when the Ethernet power supply does not meet the first preset condition; wherein the first preset condition represents a condition in which the Ethernet power supply is sufficient to power the microphone device 101 itself and its downstream devices.
[0034] The microphone device 101 receives Power over Ethernet (PoE) and data from upstream via its network port. This configuration allows the microphone device 101 to simultaneously obtain power and transmit data through a single network cable, simplifying cabling and reducing deployment costs. The DC input interface can be a physical interface on the microphone device 101 for connecting to an external DC power source, such as a standard Direct Current (DC) input interface. Its function is to provide a backup or primary power input path for the microphone device 101 when PoE is insufficient or unavailable, ensuring continuous operation of the device. The power supply control component can be an internal control structure within the microphone device 101 used to manage the power supply path. For example, the power supply control component can be a hardware switch or physical button used to dynamically adjust the power supply mode according to the power supply status. The external DC power supply can be a power supply device independent of PoE. The first preset condition can be a condition indicating that the PoE power, voltage, and other parameters provided by the upstream device are sufficient to support the operation of the microphone device 101 itself and to supply power to downstream devices.
[0035] Specifically, when the Ethernet power supply provided by the upstream microphone device 101 meets the first preset condition, the power supply control component will prioritize receiving Ethernet power. The microphone device 101 obtains power through Ethernet power and can simultaneously distribute power to downstream devices through its own power supply port. When the Ethernet power supply provided by the upstream microphone device 101 does not meet the first preset condition, the power supply control component will switch the power supply path and control the DC input interface to receive power from an external DC power source, ensuring that the microphone device 101 can operate continuously and stably and preventing downstream devices from malfunctioning due to insufficient power supply.
[0036] In some examples, power supply or logic circuitry (OR-ing) and backfeed protection circuitry can be combined to enable priority selection for different power paths, seamless switching during power outages or voltage drops, and prevent backfeeding, supporting safe and reliable chain topologies and star-chain hybrid topologies. The microphone device 101 dynamically assesses the allocable power based on its own load and the number of downstream devices. If upstream PoE is insufficient or link cable voltage drop results in insufficient power margin, the power supply mode is switched from PoE to DC power via a physical button. This strategy ensures power supply safety and predictable scalability under long link or multi-branch conditions.
[0037] In some embodiments, the preset strategy includes a first strategy for merging initial audio data and a second strategy for not merging initial audio data; the merging process includes aggregation processing or weighted mixing processing; wherein, the microphone device 101 is further configured to, based on the first strategy, filter first audio data that meets the second preset condition, and perform aggregation processing on the first audio data that meets the second preset condition to obtain second audio data; or, perform weighted mixing processing on multiple first audio data to obtain second audio data; the microphone device 101 is further configured to, based on the second strategy, determine all first audio data as second audio data; wherein, the preset strategy is configured by the audio processing device 103 or loaded by the microphone device 101 based on a locally stored scene template.
[0038] The preset strategy can be the rules or algorithms followed by the microphone device 101 when processing the received first audio data. The preset strategy can be pre-set according to specific application scenarios, audio processing requirements, or system configurations, and stored and executed internally within the microphone device 101. By introducing different preset strategies, the system can flexibly handle diverse audio data processing tasks. For example, in some cases, audio data needs to be merged to reduce the amount of data transmitted, while in other cases, the independence of the original audio data needs to be maintained.
[0039] The first strategy could be to merge multiple first audio data. The purpose of merging is to integrate audio data from different sources or different time points, thereby reducing the amount of data, simplifying subsequent processing, or achieving specific audio effects. Merging can be further subdivided into aggregation processing or weighted mixing processing, depending on the specific algorithm and requirements. The second strategy could be that after receiving the first audio data, the microphone device 101 does not merge it, but directly identifies all the received first audio data as second audio data, preserving the integrity and details of each independent audio source. For example, the second strategy can be used to maintain the independence of the original audio data when recording multiple channels or analyzing each channel independently.
[0040] The source and loading method of the preset strategy determine the system's flexibility and response speed. Configuration issued by the audio processing device 103 means that the central control unit can dynamically push the latest processing strategy to the microphone devices 101 in the cascaded link based on the global system status or user commands. Remote configuration can be performed via network protocols (such as TCP / IP and UDP). The microphone device 101 loading scene templates based on local storage allows it to make autonomous decisions and processes based on preset local configurations when offline or when the network is interrupted. For example, it can automatically load the default scene mode when the device starts up, or switch scene templates based on environmental changes detected by local sensors.
[0041] When executing the first strategy, the microphone device 101 filters first audio data that meets a second preset condition. The second preset condition may include, but is not limited to, the volume level, frequency range, specific sound source identification, timestamp, data integrity, or data source of the audio data. For example, the condition can be set to process only audio data with a volume exceeding a certain threshold, or only audio data from a specific direction. Filtering effectively removes irrelevant or low-quality audio data, thereby improving processing efficiency and the quality of the final audio data. The microphone device 101 can perform aggregation processing on the first audio data that meets the second preset condition. Aggregation processing can combine or package multiple filtered first audio data according to certain rules to form second audio data. Typically, multiple independent audio data frames or data packets can be logically or physically spliced or encapsulated. Alternatively, the microphone device 101 can perform weighted mixing processing on multiple first audio data. Weighted mixing processing can mix multiple first audio data according to preset weights to generate second audio data. During the mixing process, each input audio data is assigned a weight value, which determines its contribution to the final mixed audio.
[0042] As an example, in terms of algorithms, the system supports distributed processing and layered mixing: each microphone node can implement any of the following processing methods locally: beamforming, noise reduction, and echo cancellation; the upstream microphone node performs weighted mixing on the received downstream multi-channel signals and then outputs only the aggregated stream upwards, thereby significantly reducing the uplink bandwidth and the central processing pressure.
[0043] like Figure 6 As shown, Figure 6 This is a schematic diagram of a hierarchical aggregation or mixing strategy provided in an embodiment of this application. The first strategy can be a local processing and hierarchical aggregation strategy: individual microphone nodes enable beamforming (BF) as needed to improve directivity, enable noise suppression (NS) to suppress steady-state background noise, and perform acoustic echo cancellation (AEC) in conjunction with locally acquired far-end references; upstream microphone nodes perform weighted mixing or selective aggregation of multiple downstream uplink signals, sending only one or a small number of aggregated signals uplink to reduce uplink bandwidth usage and central processing load. The central processing unit generates a far-end reference stream consistent with the program path and distributes it downwards along the link in groups according to microphone nodes; each node completes local AEC under PTP alignment, and if necessary, can also send the locally processed signal along with the reference for centralized AEC.
[0044] like Figure 7 As shown, Figure 7This diagram illustrates a centralized aggregation or mixing strategy provided in an embodiment of this application. The second strategy can be a full upload and centralized processing strategy: all microphone nodes upload the raw or lightly pre-processed full audio data (microphone nodes do not perform mixing) to the central processing unit according to PTP-aligned timestamps. The central processing unit then uniformly completes beamforming, noise reduction, echo cancellation, and final mixing. Even if some processing is performed locally on the microphone nodes, the microphone nodes still upload the full data, allowing the central processing unit to complete the remaining processing and fusion simultaneously. The central processing unit continuously generates and distributes remote references, and each microphone node or the central processing unit performs AEC at the same time to ensure sample-level alignment and consistent output. The two strategies can be selected or used in a mixed manner depending on the scenario requirements.
[0045] In some embodiments, the system is applied to an application scenario containing multiple pickup areas, wherein each cascaded link is assigned to at least one pickup area; when different pickup areas are in different active states, the audio processing device 103 issues different preset strategies to the corresponding cascaded links, including at least one of the following: when the pickup area is in an active state, the microphone device 101 on the corresponding cascaded link executes a first strategy to perform weighted mixing or aggregation processing on multiple audio data; when the pickup area is in an inactive state, the microphone device 101 on the corresponding cascaded link executes a second strategy to transmit the original audio data or output a suppression signal; wherein the allocation of cascaded links is dynamically adjusted by the audio processing device 103.
[0046] Specifically, a pickup area refers to a physical space or logical partition that requires independent audio acquisition and processing in a specific application environment. For example, a conference room can be divided into multiple pickup areas such as a presentation area, a discussion area, and an audience area; a multi-functional hall can be divided into different pickup areas according to different activity needs (e.g., speeches, group discussions, performances). Area division can be based on physical partitions, spatial layout, logical functions, or preset scene modes. Application scenarios can cover a variety of occasions requiring multi-point audio acquisition and processing, such as large conference systems, distance education, smart classrooms, multimedia lecture halls, theater stages, and smart home environments. Assigning each cascade link to at least one pickup area means establishing a relationship between one or more cascade links in the system and one or more specific pickup areas. This can be a pre-configured static assignment, such as assigning the cascade links of conference room A to the presentation area during system initialization; or it can be a dynamically adjusted assignment, such as reassigning different cascade links to different pickup areas based on changes in conference modes. The purpose of this assignment is to achieve effective management and strategic application of audio data from different areas.
[0047] Activity status can be determined in various ways, such as by voice energy thresholds collected by microphone device 101, human image activity detected by camera, manual user commands, or spatial zoning changes detected by other sensors (e.g., infrared sensors, ultrasonic sensors). Sending different preset strategies means that audio processing device 103 selects and sends corresponding audio processing strategies to the microphone devices 101 on the corresponding cascaded link based on the determined activity status of the pickup area. Preset strategies can be templates pre-stored in audio processing device 103 or dynamically generated instructions based on real-time conditions. The sending method can be via network protocols (e.g., TCP / IP, UDP) to ensure that the strategies are accurately and promptly transmitted to the target microphone device 101.
[0048] When the pickup area is active, the microphone device 101 on the corresponding cascaded link executes the first strategy, performing weighted mixing or aggregation processing on the multi-channel audio data. When the pickup area is inactive, the microphone device 101 on the corresponding cascaded link executes the second strategy, either transmitting the original audio data directly or outputting a suppression signal. Transmitting the original audio data means that the microphone device 101 transmits the received original audio data (or after simple encapsulation) upstream without complex processing. Outputting a suppression signal means that the microphone device 101 generates a mute signal, a low-energy noise signal, or directly stops transmitting audio data for that area. Furthermore, the dynamic adjustment of the cascaded link allocation by the audio processing device 103 means that the audio processing device 103 can modify the allocation relationship between the cascaded links and the pickup areas in real time based on system operating status, environmental changes, or user needs. For example, when a change in the activity status of a pickup area is detected, the audio processing device 103 can reassess the load of the current link and adjust the cascaded link used by that area to optimize the data transmission path or load balancing. The above adjustments can be based on preset rules and algorithms, or through manual intervention.
[0049] In application scenarios involving multiple pickup areas, this application can dynamically issue different preset strategies to the corresponding cascaded links based on the activity status of different pickup areas, and dynamically adjust the allocation of cascaded links. This enables the system to flexibly adapt to complex needs in multi-area environments, effectively solving the problems of rigid strategies, resource waste, and low transmission efficiency in traditional solutions.
[0050] In some embodiments, when the system is configured in the main speaker mode, the main speaker area is set to an active state, and other microphone areas are set to an inactive state; when the system is configured in the discussion mode, the discussion area is set to an active state, and the main speaker area is set to an inactive state; when the system is configured in the free discussion mode, each microphone area is dynamically configured to an active or inactive state according to the real-time activity status.
[0051] The presenter mode prioritizes audio acquisition and processing in the presenter area. This mode can be selected by the user through the interface of the audio processing device 103 or triggered by an external control system (e.g., a conference management system) via preset commands. The presenter area is set to active status, causing the microphone devices 101 within this area to merge the acquired audio data (e.g., aggregate or weighted mixing) to ensure clear and complete transmission of the speaker's voice. Other microphone areas are configured to inactive status, allowing the microphone devices 101 in other areas to transmit raw audio data or output suppression signals, thereby reducing background noise and interference from non-presenter areas. The discussion mode prioritizes audio acquisition and processing in the discussion area while lowering the priority of the presenter area. This mode is triggered similarly to the presenter mode. The discussion area is set to active status, causing the microphone devices 101 within this area to merge the acquired audio data to focus on the speaker's voice. The main speaking area is configured to be inactive, so that the microphone device 101 in the area can transmit raw audio data or output suppressed signals to avoid unnecessary interference in the main speaking area during the discussion.
[0052] The free discussion mode dynamically adjusts the activity status of each pickup area based on its real-time activity. The triggering method for this mode is similar to the previously mentioned modes. Each pickup area is dynamically configured to be active or inactive based on its real-time activity status, instructing the audio processing device 103 to continuously monitor the activity status of each pickup area (e.g., voice energy collected by the microphone device 101, human image activity detected by the camera, or external commands issued by the user), and dynamically sends a first or second strategy to the corresponding cascaded link based on the monitoring results, thereby achieving flexible adjustment of the audio processing strategy.
[0053] The main speaking mode of this application ensures priority processing and clear transmission of audio in the main speaking area, effectively suppressing interference from non-main speaking areas; the discussion mode focuses on audio acquisition in the discussion area, avoiding unnecessary interference from the main speaking area during discussions; and the free discussion mode can dynamically adjust the activity status of each sound pickup area according to the real-time activity status, greatly enhancing the system's adaptability and efficiency in complex and changing scenarios.
[0054] In some embodiments, the active or inactive state of the pickup area is determined based on at least one of the following methods: when the voice energy collected by the microphone device 101 continuously exceeds a third preset condition, the corresponding pickup area is configured as active; when the camera detects human image activity in the pickup area and identifies it as a speaker, the corresponding pickup area is configured as active; when an external command is received from the user, the relevant pickup area is configured as active or inactive according to the command content; when the system detects a sensor signal indicating a change in the spatial partitioning state of the pickup area, the cascaded links are reallocated according to the spatial partitioning result and the active state of the corresponding area is updated.
[0055] Speech energy refers to the energy of a sound wave passing through a unit area per unit time, and is typically quantified using parameters such as amplitude, power, or loudness of the audio signal acquired by microphone device 101. The third preset condition is a parameter used to determine whether the speech energy has reached an active state threshold. When the speech energy continuously exceeds this condition, it indicates that there is continuous speech activity in the pickup area, and it should be configured as an active state.
[0056] Human activity refers to the movement or posture changes of a human body or face detected by the camera within a designated area using image processing technology. Determining whether detected human activity is related to speaking behavior can be achieved by combining technologies such as facial recognition, lip reading, head posture analysis, and gesture recognition, or by judging whether the person is facing the microphone device 101 and whether their mouth is opening or closing. When the camera detects human activity and identifies it as a speaker, the pickup area is configured as active.
[0057] User-issued external commands refer to control commands sent by the user to the system through an external interface (such as a control panel, remote control, mobile application, or voice assistant). After receiving the command, the system parses the command content and configures the relevant pickup area to be active or inactive according to the intent of the command, allowing for manual intervention and flexible control to adapt to special scenario requirements.
[0058] The spatial zoning status of the sound pickup area refers to the division and layout of sound pickup areas in the physical space. For example, a large conference room may be divided into multiple smaller conference rooms by movable partitions, or different discussion areas may be formed by furniture arrangement. Sensor signals refer to physical environment change signals detected by various sensors (e.g., infrared sensors, door magnetic sensors, etc.). For example, when events such as partition movement, door opening or closing, or people entering or exiting occur, sensors will generate corresponding signals. After detecting the above signals, the system will analyze the signals according to preset logic or algorithms to determine whether the spatial zoning has changed and generate spatial zoning results. The spatial zoning results are the system's understanding of the current physical space layout. For example, conference room A is divided into two independent areas, A1 and A2. Based on the above results, the system will reallocate cascading links, that is, associate the original cascading links or newly added cascading links with the new sound pickup areas and update the activity status of the corresponding areas. For example, if a large area is divided into two smaller areas, the system may configure independent cascading links for these two smaller areas and update their activity status according to their initial state or subsequent detection results.
[0059] This application enables accurate and dynamic judgment of the active state of the pickup area, allowing the audio processing device 103 to flexibly adjust its audio data processing strategy based on actual speaking conditions and environmental changes. This avoids ineffective processing of inactive areas, effectively reducing system resource consumption and improving the efficiency of audio data transmission and processing. Simultaneously, the multi-source information fusion judgment mechanism significantly improves the accuracy and robustness of active state recognition, ensuring the system can continuously provide high-quality audio services in complex and ever-changing application scenarios.
[0060] In some embodiments, the audio processing device 103 is further configured to dynamically adjust a preset strategy based on the transmission status of the cascaded link and / or the processing load of the microphone device 101; wherein the transmission status includes at least one: link bandwidth occupancy status, data transmission delay or packet loss status; the processing load includes at least one: computing resource occupancy rate of the microphone device 101, cache occupancy status or power consumption status; when the transmission status or processing load meets a fourth preset condition, the audio processing device 103 sends at least one of the first strategy or the second strategy to the corresponding cascaded link to change the way the microphone device 101 processes audio data.
[0061] Among these, transmission status refers to the performance of the cascaded link during data transmission. Link bandwidth occupancy status refers to the proportion of the network link occupied by data traffic, which can be obtained by monitoring network interface traffic statistics or packet sending / receiving rates. Data transmission latency refers to the time required for data to travel from the sender to the receiver, which can be detected by sending probe packets and measuring round-trip time (RTT) or by using Precise Time Protocol (PTP) / Network Time Protocol (NTP) for time synchronization. Packet loss status refers to the situation where data packets are lost during transmission, which can be identified by statistically analyzing packet sequence numbers or loss reports from the receiver. Processing load refers to the computing resources consumed by microphone device 101 when performing audio processing tasks.
[0062] The computing resource utilization rate of microphone device 101 can refer to the utilization rate of the central processing unit (CPU) or memory usage rate, which is usually statistically analyzed through the device's internal performance monitoring module. Cache occupancy status refers to the usage status of the buffer used to store audio data within microphone device 101, such as whether the buffer is full or nearly full. Power consumption status refers to the electrical energy consumed by microphone device 101 during operation, which can be monitored or estimated through the built-in power management unit or sensors.
[0063] Dynamically adjusting preset strategies means that the audio processing device 103 can flexibly switch or modify the audio data processing strategy executed by the microphone device 101 based on changes in the transmission status and / or processing load monitored in real time. For example, when system resources are scarce, it can switch to a strategy with lower resource consumption; when resources are plentiful, it can switch to a strategy that provides better audio quality.
[0064] The fourth preset condition refers to a specific threshold or set of rules that triggers the audio processing device 103 to adjust its strategy. This condition can be a combination of one or more transmission status parameters (e.g., bandwidth utilization exceeding a certain percentage, packet loss rate exceeding a certain threshold) or processing load parameters (e.g., CPU utilization reaching a certain upper limit, cache utilization reaching a certain critical value). For example, when the link bandwidth utilization exceeds 80% for a continuous period of time, or when the CPU utilization of the microphone device 101 is consistently higher than 90%, the fourth preset condition is met, thereby triggering a strategy adjustment.
[0065] This application can dynamically adjust the audio data processing strategy based on the transmission status of the cascaded link and the processing load of the microphone device 101, effectively solving the problems of low strategy execution efficiency or system instability that may occur when the transmission status of the cascaded link is poor or the processing load of the microphone device 101 is too high. This solution ensures that the system maintains efficient and stable audio data transmission and processing capabilities under various complex and dynamically changing operating conditions, significantly improving the robustness and reliability of the system. Simultaneously, by intelligently optimizing resource allocation, unnecessary resource waste is avoided, thereby guaranteeing the transmission quality of audio data and the user experience.
[0066] In some embodiments, the audio processing device 103 includes a first clock, and the microphone device 101 and / or the network switching device 102 includes a second clock; wherein, the first clock is used to send downlink data containing current time data to the microphone device 101 and / or the network switching device 102; the second clock is used to determine the delay data of the downlink data; the delay data is used to provide time compensation for the downstream microphone device 101; a target time data synchronized with the current time data is determined based on the delay data; the target time data is used to provide a time reference for the downstream microphone device 101.
[0067] The first clock can be the master clock located in the audio processing device 103, which generates and maintains a high-precision, globally unified time reference. The first clock can encapsulate the current time data in downlink data using a time synchronization protocol and periodically send it to the microphone device 101 and / or network switching device 102 in the system as a time reference for the entire system. The time synchronization protocol can be a Precision Time Protocol (PTP), such as IEEE 1588 PTP. The second clock can be a local clock configured in the microphone device 101 and / or network switching device 102. It receives time information from the first clock in the audio processing device 103 and adjusts its local time according to the received information to achieve synchronization with the first clock.
[0068] In some implementations, the second clock is a transparent clock used to correct for accumulated forwarding delay. The second clock can be calibrated in frequency and phase using a clock synchronization algorithm (e.g., a slave clock algorithm of PTP) to keep it in sync with the master clock. Downlink data includes current time data generated by the first clock. Delay data can be the time delay experienced between the transmission of downlink data from the audio processing device 103 and its reception.
[0069] In other implementations, the second clock can be a boundary clock used to reconstruct the local slave clock at branch points, providing a more stable reference downstream. Upon receiving downlink data containing current time data, the second clock calculates delay data by measuring the round-trip time of the data packet or using other synchronization mechanisms. This delay data can be used to provide time compensation for downstream microphone devices 101; for example, in a multi-stage cascaded system, an upstream device can pass its calculated delay data to a downstream device to help it synchronize more accurately. The target time data can be the local time synchronized with the first clock, determined by the second clock after receiving the current time data and calculating and calibrating it in conjunction with the delay data. In other words, the target time data can be the accurate time achieved by the second clock after compensation and adjustment, ensuring consistency with the first clock. The target time data is not only used for the internal operation of the second clock itself but also serves as a time reference for its connected downstream microphone devices 101, ensuring that all devices in the entire audio data transmission chain can share a unified and accurate time reference.
[0070] As an example, the network-wide audio clock is aligned using IEEE 1588 PTP. Each node can be configured as a transparent clock or a boundary clock according to deployment requirements. Adjustable jitter buffering and fixed-frame delay compensation achieve precise signal alignment at the sample level across multiple nodes. Time synchronization and audio transmission employ a fixed-frame transmission mechanism based on network-wide PTP alignment. Each node completes PTP convergence before sampling, then samples and packages audio with a fixed frame length, embedding a PTP-aligned timestamp at the transmitting end. The receiving end performs fixed-frame delay compensation based on this timestamp and the latency measured on the link, and absorbs network jitter through adjustable jitter buffering to achieve sample-level aligned playback and processing. For ease of engineering integration, the audio bearer can adopt common real-time transmission formats (e.g., packetization and timestamps based on Real-time Transport Protocol (RTP) / User Datagram Protocol (UDP)), ensuring phase consistency between cross-node synthesis and algorithm processing within the same clock domain.
[0071] For example, in applications requiring sound source localization or beamforming, precise time synchronization ensures that sound signals collected by different microphones are aligned on the time axis, thereby enabling high-precision spatial information extraction and processing; in scenarios such as echo cancellation, an accurate time reference helps to accurately identify and eliminate echoes, improving call or recording quality.
[0072] In some embodiments, the audio processing device 103 is further configured to send a reference signal to the microphone device 101 and / or the network switching device 102; the network switching device 102 is configured to send the reference signal to the downstream microphone device 101 and / or the downstream network switching device 102; and the microphone device 101 is configured to perform echo cancellation processing on the audio data based on the reference signal.
[0073] In this system, audio processing device 103, acting as the source of audio output, generates and transmits a known and controllable reference signal. Upon receiving this reference signal, network switching device 102 forwards it to downstream microphone device 101 and / or downstream network switching device 102 according to its internal configuration or routing rules. Upon receiving the reference signal, microphone device 101 uses its internal digital signal processor or microcontroller to execute an echo cancellation algorithm, obtaining clean audio data without echo. As an example, the remote reference signal is generated by the central audio processing device 103 and synchronously distributed to each node along the link, supporting node-local AEC or centralized AEC strategies.
[0074] like Figure 8 As shown, Figure 8 This application provides a flowchart illustrating a data transmission method according to an embodiment of the present application. The present application also provides a data transmission method applied to an audio data transmission system; the method includes: Step 801: Collect local audio data using a microphone device and transmit the audio data upstream in the cascaded link.
[0075] Step 802: Receive audio data streams from one or more downstream devices through a network switching device, and perform path management and forwarding control on the audio data streams to transmit the audio data streams upstream of the cascaded link.
[0076] Step 803: Connect and process the audio data aggregated from the cascaded link through the audio processing device.
[0077] In this application, the microphone device not only serves as a data source but also possesses preliminary data forwarding capabilities, while the network switching device acts as a crucial aggregation and routing node, effectively managing and scheduling audio data streams within the network. This significantly reduces the complexity of traditional cabling and improves the system's deployment flexibility and scalability in large or complex environments. The audio processing device can receive data from different nodes in the network, further enhancing the system's robustness and adaptability, ensuring that audio data can be stably and timely transmitted to the processing center, providing a reliable data foundation for subsequent audio analysis, processing, or applications. In some embodiments, the method further includes: sending a preset strategy to the cascaded link via an audio processing device based on system configuration or environmental status; and processing the first audio data according to the preset strategy to obtain second audio data upon receiving first audio data sent by a downstream microphone device, and transmitting the second audio data upstream of the cascaded link.
[0078] After receiving the first audio data sent by the downstream microphone device, the microphone device of this application can process it according to a preset strategy to generate the second audio data and send it upstream, effectively reducing the amount of raw data that needs to be transmitted and reducing the network bandwidth usage.
[0079] In some embodiments, the method further includes: the microphone device receiving Ethernet power and data from an upstream source through its network port; controlling the microphone device to use Ethernet power when the Ethernet power meets a first preset condition; and controlling the microphone device to use an external DC power supply when the Ethernet power does not meet the first preset condition; wherein the first preset condition represents a condition in which the Ethernet power is sufficient to supply power to the microphone device itself and its downstream devices.
[0080] In this application, when the Ethernet power supply provided by the upstream microphone device is sufficient to meet the power supply needs of downstream devices, including itself, the microphone device prioritizes using Ethernet power supply, thereby simplifying wiring and reducing installation complexity. When Ethernet power supply is insufficient to support the power supply needs of the entire link, the system can promptly identify and automatically switch to external DC power supply.
[0081] In some embodiments, the method further includes: sending downlink data containing current time data to a microphone device and / or a network switching device via a first clock of the audio processing device; determining delay data of the downlink data via a second clock of the microphone device and / or the network switching device; the delay data being used to provide time compensation for downstream microphone devices; determining target time data synchronized with the current time data based on the delay data via the second clock; the target time data being used to provide a time reference for downstream microphone devices.
[0082] This application provides a precise time synchronization mechanism. The first clock of the audio processing device serves as the master clock source, sending downlink data containing current time information to downstream devices, providing a unified time reference for the entire system. Upon receiving this data, the second clock of the microphone device and / or network switching device accurately calculates the data transmission delay and determines the target time data to be synchronized with the master clock based on this delay data. This target time data is not only used to synchronize the internal clock of the current device but also serves as a time base provided to the downstream microphone device, thereby ensuring precise time synchronization of all devices in the distributed audio data transmission system.
[0083] The following describes the data transmission system and method provided in the embodiments of this application. This application belongs to the field of distributed audio acquisition and network transmission technology, and relates to a scheme for the infinite expansion of distributed matrix microphones, which may include a distributed microphone system and method that can be linearly cascaded and can be expanded approximately infinitely in engineering.
[0084] The system workflow is as follows: During the deployment phase, the trunk and branch locations are determined based on coverage requirements, and the physical connections of microphone nodes are completed. After power-on, each microphone node determines its current power path and establishes a steady-state power supply via a hardware switch. Subsequently, the uplink port completes link negotiation and PTP network access. Once PTP is locked and stable, the microphone nodes begin sampling, packaging, and transmitting at fixed frame lengths. Upstream microphone nodes synchronously receive the lower-level audio and perform local aggregation before outputting the aggregated stream upstream. The central audio processing equipment performs necessary global processing and remote reference generation at the same time, and distributes the references to each node in groups to support AEC. During operation, if the upstream power supply or link status changes, the nodes smoothly switch between PoE and DC according to predetermined priorities, maintaining audio continuity without disrupting PTP alignment.
[0085] Specific implementation: In a long, narrow conference room, several microphone nodes can be cascaded along the edge of the table. The first microphone node receives power and supplies power to subsequent stages. The end stage can be connected to a DC power supply as needed to compensate for power and voltage margins. In a lecture hall or multi-section space, branch nodes can be set up in the middle of the trunk line to perform local mixing for each section before uploading the aggregated stream. This significantly reduces uplink bandwidth and simplifies the computational load of central processing while ensuring sample-level synchronization. All the above implementations follow a unified design where port A is for PD uplink, port B is for PSE downlink, and the DC input is optional redundant or main power supply. Combined with PTP synchronization and fixed frame delay compensation, this enables replicable deployment and expansion in engineering applications.
[0086] like Figure 9 and Figure 10 As shown, Figure 9 This is a schematic diagram of the visual management interface of the data transmission system provided in the embodiments of this application. Figure 10 This is another schematic diagram of the data transmission system provided in an embodiment of this application. In a large training room accommodating 100 people, a main podium is located at the front, and the middle is divided into four student discussion areas (front left, rear left, front right, and rear right). Each discussion area is equipped with a round table for 6-8 people, and there are tiered seats at the back. A high-performance ceiling microphone (node A1) is installed above the main podium. One ceiling microphone (nodes B1-B4) is installed above the center of each of the four discussion areas (front left, rear left, front right, and rear right). Four desktop microphones (nodes C1-C16) are set up in each discussion area. The tiered seating area can also be divided into sound pickup zones and ceiling microphones can be deployed. Figure 11 As shown, Figure 11 This is a schematic diagram of a microphone device visual management interface provided in an embodiment of this application, including status guidance when the microphone device is paired or physically located, to help users quickly match the virtual list with the actual microphone device.
[0087] The microphones described above employ a star-chain hybrid topology. Specifically, the central processing unit (CPU) connects to node A1, speakers / cameras / control device. Port B of node A1 connects to a network switch (e.g., a network switch). The network switch acts as a branch point, connecting to ports A of nodes B1-B4. Port B of B1 connects to the network switch, which in turn connects to four desktop microphones (C1-C4). Similarly, B2-B4 each connect to their respective desktop microphones and several speakers / cameras within their discussion area via the network switch. All connections use standard network cables and support PoE power supply. The CPU connects to the PoE network switch / DC input to power node A1. Node A1 acts as a power source (PSE), providing power to the network switch, which in turn provides power to nodes B1-B4. Nodes B1-B4 distribute power to their local network switch and the four desktop microphones. If power is insufficient, the DC input can be used as optional redundancy or as the primary power source. The entire system uses the PTP protocol for clock synchronization, achieving microsecond-level accuracy. The CPU can be connected to a PC or host device for management via an interface.
[0088] Dynamic audio pickup strategies in large training room scenarios: For the main speaker mode, the first strategy can be used, where node A1 beamforming focuses on the main speaker area, suppressing noise in other areas. Nodes B1-B4 in the discussion area only perform basic noise reduction, outputting mute or low gain. For the discussion mode, a hybrid strategy can be used, switching to the second strategy when the main speaker is no longer focused, while the discussion area switches to the first strategy to extract the speaker's voice from the desktop. The branch junctions of nodes B1-B4 selectively aggregate the four signals, and the central processor synthesizes five audio streams (1 main speaker stream + 4 discussion area streams). For the free speaking mode, the second strategy can be used, optionally combining sound source localization to identify the speaker's location and dynamically adjusting the gain of each area. Different modes can be automatically identified or manually switched in the management interface.
[0089] In other embodiments, such as Figure 12 As shown, Figure 12This is another schematic diagram of the data transmission system provided in the embodiments of this application. It can segment a conference room scene, adding sensors to the physical partition doors for detection. When the door is open, it operates in a merged mode; when the door is closed, it operates in a segmented mode. The system performs partitioned acquisition, processing, and management of the audio pickup area. The number of microphone nodes and the cascading method are adjusted according to the size of the conference room. A network switching device acts as a data hub, connecting the central processing unit to three partitions: A, B, and C. Microphone nodes within each partition (e.g., nodes 1-4 in partition A) are cascaded serially. For microphones near the partitions in the space, in addition to normal audio pickup, optional audio pickup outside the pickup area (on the other side of the partition) is used for algorithmic cancellation.
[0090] To implement the method of the embodiments of this application, such as Figure 13 As shown, Figure 13 This is a schematic diagram of a hardware structure of an electronic device provided in an embodiment of this application. This application also provides an electronic device 130 that may include: a memory 1301 for storing a computer program; and a processor 1302 for implementing the method described above when executing the computer program. The processor 1302 can implement the steps of any of the methods described above, which will not be elaborated further here.
[0091] Of course, in practical applications, such as Figure 13 As shown, the electronic device 130 may further include at least one network interface 1303. Various components in the electronic device are coupled together via a bus system 1304. It is understood that the bus system 1304 is used to implement communication between these components. In addition to a data bus, the bus system 1304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 13Various buses are labeled as bus system 1304. The number of processors 1302 can be at least one. Network interface 1303 is used for wired or wireless communication between electronic devices and other devices. Memory 1301 in this embodiment is used to store various types of data to support the operation of the electronic device. The methods disclosed in the above embodiments can be applied to processor 1302, or implemented by processor 1302. Processor 1302 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 1302 or by instructions in software form. The processor 1302 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 1302 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected in the combined execution of hardware and software modules in a microcontroller. The software module can reside in a storage medium located in memory 1301. Processor 1302 reads information from memory 1301 and, in conjunction with its hardware, completes the steps of the aforementioned method. In an exemplary embodiment, electronic device 130 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.
[0092] Specifically, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, such as a memory 1301 storing the computer program, which can be executed by a processor 1302 to complete the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0093] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. A data transmission system, characterized in that, The system includes: At least one microphone device, at least one network switching device, and an audio processing device; The at least one microphone device is connected to the network switching device in a cascaded manner to form at least one cascaded link; The microphone device is used to collect local audio data and transmit the audio data upstream in the cascaded link; The network switching device is used to receive audio data streams from one or more downstream devices and transmit the audio data streams upstream of the cascaded link. The network switching device is used as a branch node in the cascaded link to perform path management and forwarding control of audio data streams from different downstream devices. The audio processing device is connected to the upstream end of the at least one cascaded link and is used to receive and process the converged audio data.
2. The system according to claim 1, characterized in that, include: The microphone device is also used to receive first audio data sent by a downstream microphone device, and process the first audio data according to a preset strategy to obtain second audio data; The second audio data is transmitted upstream in the cascaded link; The network switching device is also used to receive a second audio data stream from one or more downstream devices; The second audio data stream is transmitted upstream in the cascaded link; The audio processing device is also used to send the preset strategy to the cascaded link based on system configuration or environmental status.
3. The system according to claim 1, characterized in that, The microphone device also includes a DC input interface and a power supply control component; wherein... The microphone device is configured to receive Ethernet power and data from upstream via its network port; The DC input interface is used to connect an external DC power supply; The power supply control component is configured to control the microphone device to use the Ethernet power supply when the Ethernet power supply meets a first preset condition; and to control the microphone device to receive power from the external DC power supply using a DC input interface when the Ethernet power supply does not meet the first preset condition. The first preset condition represents the condition that the Ethernet power supply is sufficient to power the microphone device itself and its downstream devices.
4. The system according to claim 2, characterized in that, The preset strategy includes a first strategy of merging the initial audio data and a second strategy of not merging the initial audio data. The merging process includes aggregation processing or weighted mixing processing; wherein... The microphone device is further configured to, based on the first strategy, filter the first audio data that meets the second preset conditions, aggregate the first audio data that meets the second preset conditions to obtain the second audio data; or, perform weighted mixing processing on multiple first audio data to obtain the second audio data. The microphone device is further configured to determine all the first audio data as the second audio data based on the second strategy; The preset strategy is configured by the audio processing device or loaded by the microphone device based on a locally stored scene template.
5. The system according to claim 4, characterized in that, The system is applied in application scenarios that include multiple pickup areas, wherein each of the cascaded links is assigned to at least one pickup area; When different pickup areas are in different active states, the audio processing device issues different preset strategies to the corresponding cascaded links, including at least one of the following: When the pickup area is active, the microphone device on the corresponding cascaded link executes the first strategy to perform weighted mixing or aggregation processing on multiple audio data. When the pickup area is inactive, the microphone device on the corresponding cascaded link executes the second strategy, either transmitting the original audio data or outputting a suppression signal. The allocation of the cascaded links is dynamically adjusted by the audio processing device.
6. The system according to claim 5, characterized in that, When the system is configured in speaker mode, the speaker area is set to active state, and other sound pickup areas are set to inactive state. When the system is configured in discussion mode, the discussion area is set to active and the main speaker area is set to inactive. When the system is configured in free discussion mode, each pickup area is dynamically configured to be active or inactive based on the real-time activity status.
7. The system according to claim 5, characterized in that, The active or inactive state of the pickup area is determined based on at least one of the following methods: When the voice energy collected by the microphone device continuously exceeds the third preset condition, the corresponding pickup area is configured to be in an active state. When the camera detects human activity in the sound pickup area and identifies it as a speaker, the corresponding sound pickup area is configured to be active. When an external command is received from the user, the relevant pickup area is configured to be active or inactive according to the command content. When the system detects a sensor signal indicating a change in the spatial partitioning state of the pickup area, it reallocates the cascaded links and updates the active state of the corresponding area based on the spatial partitioning results.
8. The system according to claim 5, characterized in that, The audio processing device is also used to dynamically adjust the preset strategy based on the transmission status of the cascaded link and / or the processing load of the microphone device; The transmission status includes at least one of the following: link bandwidth occupancy status, data transmission delay, or packet loss status. The processing load includes at least one of the following: the microphone device's computing resource utilization rate, cache utilization status, or power consumption status. When the transmission state or the processing load meets the fourth preset condition, the audio processing device sends at least one of the first strategy or the second strategy to the corresponding cascaded link to change the way the microphone device processes audio data.
9. The system according to claim 1, characterized in that, The audio processing device includes a first clock, and the microphone device and / or the network switching device includes a second clock; wherein... The first clock is used to send downlink data containing current time data to the microphone device and / or the network switching device; The second clock is used to determine the delay data of the downlink data; the delay data is used to provide time compensation for downstream microphone devices; a target time data synchronized with the current time data is determined based on the delay data; the target time data is used to provide a time reference for the downstream microphone devices.
10. The system according to claim 1, characterized in that, The audio processing device is also used to send reference signals to the microphone device and / or the network switching device; The network switching device is used to transmit the reference signal to a downstream microphone device and / or a downstream network switching device; The microphone device is used to perform echo cancellation processing on the audio data based on the reference signal.
11. An audio data transmission method, characterized in that, Applied to the audio data transmission system as described in any one of claims 1-10; the method includes: Local audio data is collected using a microphone device and transmitted upstream in the cascaded link. The network switching device receives audio data streams from one or more downstream devices and performs path management and forwarding control on the audio data streams to transmit the audio data streams upstream of the cascaded link. The audio data aggregated from the cascaded link is processed by an audio processing device.
12. The method according to claim 11, characterized in that, The method further includes: The audio processing device sends a preset strategy to the cascaded link based on system configuration or environmental status. The microphone device, upon receiving first audio data sent by a downstream microphone device, processes the first audio data according to the preset strategy to obtain second audio data, and then transmits the second audio data upstream in the cascaded link.
13. The method according to claim 11, characterized in that, The method further includes: The microphone device receives Ethernet power and data from upstream via its network port; When the Ethernet power supply meets the first preset condition, the microphone device is controlled to use the Ethernet power supply. If the Ethernet power supply does not meet the first preset condition, the microphone device is controlled to use an external DC power supply. The first preset condition represents the condition that the Ethernet power supply is sufficient to power the microphone device itself and its downstream devices.
14. The method according to claim 11, characterized in that, The method further includes: The audio processing device sends downlink data containing the current time data to the microphone device and / or the network switching device via the first clock of the audio processing device. The delay data of the downlink data is determined by a second clock of the microphone device and / or the network switching device; the delay data is used to provide time compensation for the downstream microphone device; The second clock determines a target time data that is synchronized with the current time data based on the delay data; the target time data is used to provide a time reference for the downstream microphone device.
15. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 11-14.