Audio tag management-based multi-party intercom method and system
By optimizing the transmission path through audio tag management and tree diagrams, the problems of limited connectivity, complex mixing, and poor user experience in multi-party intercom were solved, achieving efficient and reliable multi-party intercom connectivity and dynamic device management.
Patent Information
- Application Number
- PCT/CN2025/090426
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2025-04-22
- Publication Date
- 2026-03-05
AI Technical Summary
In existing multi-party intercom technologies, direct connection for transmitting audio data is limited by the small range and cumbersome nature of wired connections, while wireless connections pose security risks. Server-side audio mixing is complex and inefficient, making it difficult to handle large-scale multi-person intercom scenarios. Furthermore, the status and flow of audio data are not intuitive.
By managing audio tags, we can identify forwarding servers and intercom servers, set audio packet transmission formats, construct a tree diagram to optimize transmission paths, and use audio tags to filter and mix audio packets to avoid echoes and noise, thereby improving server processing efficiency and user experience.
It enables fast and reliable multi-party communication, reduces voice latency, improves user experience and server processing efficiency, and supports multi-tasking and smooth switching when devices join or leave.
Smart Images

Figure CN2025090426_05032026_PF_FP_ABST
Abstract
Description
A multi-party intercom method and system based on audio tag management Technical Field
[0001] This application relates to the field of multi-party wireless intercom technology, and in particular to a multi-party intercom method and system based on audio tag management. Background Technology
[0002] Existing intercom solutions mainly follow two directions: one is to transmit audio data through direct connection between intercom devices, and the other is for intercom devices to transmit audio data to a server for mixing and then distribute it to each intercom device.
[0003] Intercom solutions that directly transmit audio data include wired and wireless connections. Wired connections may use RS485, RS232, or TCP / IP protocols. However, due to the limitations of the data cable, the effective range of the intercom is limited, and the wiring is cumbersome; therefore, wired connections are generally not chosen. Wireless connections may use Bluetooth, Wi-Fi, or other wireless communication protocols. Intercom devices need to have audio encoding and decoding capabilities to ensure efficient audio data transmission. Common audio encoding formats include, but are not limited to, PCM (Pulse Code Modulation), MP3, AAC, and Opus. Audio data is usually transmitted using the UDP protocol. Although UDP has lower latency, it does not guarantee reliable data packet transmission, posing certain security risks. Furthermore, in direct audio data transmission, the data flow is difficult to capture during direct transmission between intercom devices, making troubleshooting challenging; obtaining and updating the addresses of other intercom devices during direct audio data transmission is also difficult.
[0004] In server-connected scenarios, intercom devices transmit audio to the server via network protocols (such as UDP or TCP), and the server receives audio data packets from each intercom device. The audio streams are then mixed using common mixing libraries (such as WebRTC, FFmpeg, and PortAudio) and distributed to the intercom devices to enable multi-party communication. However, this approach requires the server-level mixing of audio streams to generate corresponding mixed audio streams for different intercom devices. Each mixed audio stream needs to exclude the target intercom's own audio data; otherwise, echoes will occur. This process is complex, inefficient, and unsuitable for large-scale multi-person intercom scenarios. Furthermore, the audio data status and flow are not intuitive. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a multi-party intercom method and system based on audio tag management, which improves the server's efficiency in processing audio data and enhances the user experience.
[0006] Firstly, this application provides a multi-party intercom method based on audio tag management, including:
[0007] Obtain an intercom task, wherein the intercom task includes a task identifier, each intercom device participating in the intercom, and a first audio tag corresponding to each intercom device;
[0008] The forwarding server, several intercom servers, and data transmission path are determined according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices;
[0009] According to the data transmission path, the intercom task is sent to each of the intercom servers through the forwarding server, so that each of the intercom servers registers its corresponding intercom device according to the intercom task and sets the audio packet transmission format of its corresponding intercom device according to each first audio tag.
[0010] The multi-party intercom is conducted by transmitting audio packets through each of the intercom devices, each of the intercom servers, and the forwarding server according to the data transmission path.
[0011] This application provides a multi-party intercom method based on audio tag management. It determines several intercom servers participating in the intercom and a forwarding server responsible for forwarding audio packets through intercom tasks. The forwarding server distributes the intercom task to each intercom server, and the intercom devices involved in the intercom task are registered to establish a multi-party intercom connection. Simultaneously, each intercom server sets its own audio packet upload and reception formats based on the first audio tag in the intercom task, realizing multi-party intercom based on audio tag management. This ensures that during the intercom process, each audio packet is transmitted in a format preset by the intercom server. Therefore, each intercom server can quickly parse the audio packets, determine the intercom task to which the audio packet belongs and the data transmission path based on the audio tag, and achieve rapid forwarding of the audio packets. Because audio tags are set, the forwarding server does not need to generate corresponding mixed audio streams for different intercom devices. Instead, it can directly mix audio packets with the same task. After receiving the mixed audio packets, each intercom device can compare the received audio packet format with the audio tags carried by the mixed audio packets to filter the mixed audio packets, avoid echoes and other noises, improve the server's processing efficiency of audio data and the user experience.
[0012] Furthermore, the step of determining the forwarding server, several intercom servers, and data transmission path based on the intercom task includes:
[0013] Based on the topological connections between all servers, a tree diagram is constructed, where each server is a node in the tree diagram;
[0014] Based on the server ownership information of each intercom device participating in the intercom task, several intercom servers are determined.
[0015] Based on the position of each intercom server in the tree diagram, determine the least common ancestor node of each intercom server;
[0016] The server corresponding to the least common ancestor node is used as the forwarding server;
[0017] Based on the tree diagram, the shortest path from the forwarding server to each of the intercom servers is determined, and the shortest path is used as the data transmission path.
[0018] This application provides a method for confirming transmission paths. Since the servers in this embodiment are connected in a hierarchical tree structure, direct communication between them may be impossible. Therefore, a common ancestor server is needed to distribute and acquire audio packets, i.e., the forwarding server. Thus, after acquiring an intercom task, it is first necessary to determine each intercom server, the forwarding server, and the data transmission path. This application sets the least common ancestor node of each intercom server as the forwarding server and uses the shortest path from the forwarding server to each intercom server as the data transmission path, minimizing data transmission distance, reducing voice intercom latency, and improving the user experience.
[0019] In one possible implementation, when any first intercom server among the intercom servers sets the audio packet upload format and audio packet reception format of the corresponding first intercom device according to the respective first audio tags, each of the intercom servers sets the audio packet transmission format of its respective corresponding intercom device according to the respective first audio tags, including:
[0020] A second audio tag set corresponding to the first intercom device is determined based on each of the first audio tags, wherein the second audio tag set is a tag set composed of each of the first audio tags corresponding to each non-first intercom device;
[0021] A control input command is generated based on the first audio tag and audio packet sending address corresponding to the first intercom device, and the control input command is sent to the first intercom device to set the audio packet upload format of the first intercom device.
[0022] A control output command is generated based on the second audio tag set and audio packet receiving address corresponding to the first intercom device, and the control output command is sent to the first intercom device to set the audio packet receiving format of the first intercom device.
[0023] This application embodiment takes a first intercom server as an example and provides a method for setting the audio packet upload format and audio packet reception format of an intercom device. By setting a second audio tag set, the audio tags of the first intercom device itself are distinguished from the audio tags of other intercom devices. By sending control input commands to the first intercom device, the audio packet sending address and the first audio tag carried by the intercom device when uploading audio packets are set. By sending control output commands to the first intercom device, the intercom device is set to only receive audio packets with audio tags from the second audio tag set and whose audio packet reception address is used, thus filtering out other audio packets and avoiding echoes and other noise. At the same time, the server does not need to perform targeted processing on mixed audio packets, but can directly use the intercom device for audio filtering, improving the server's processing efficiency of audio data and the user experience.
[0024] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0025] Intercom devices acquire user voice input;
[0026] The intercom device converts the voice input into a corresponding input audio packet according to the corresponding audio packet upload format;
[0027] The intercom device uploads the corresponding input audio packet to the corresponding intercom server, so that each intercom server uploads the input audio packet to the forwarding server;
[0028] The intercom device receives a mixed audio packet sent by a corresponding intercom server. The mixed audio packet is obtained by the forwarding server merging the various input audio packets and then sending it to each of the intercom servers through the forwarding server.
[0029] The intercom device converts the mixed audio packet into several corresponding output audio packets according to the corresponding audio packet receiving format;
[0030] The intercom device outputs voice based on a number of corresponding output audio packets.
[0031] This application describes the actions performed by the intercom device during the intercom process. During voice input, the intercom device packages the user's voice input into a corresponding input audio packet according to the audio packet upload format, and uploads it to the forwarding server through the intercom server. During voice output, the intercom device parses the obtained mixed audio packet according to the audio packet receiving format, obtains several output audio packets, and performs voice output based on the several output audio packets, thereby realizing multi-party intercom based on audio tag management.
[0032] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0033] The intercom server continuously acquires input audio packets and mixed audio packets;
[0034] After the intercom server obtains the input audio packet, it determines the data transmission path corresponding to the input audio packet based on the first audio tag of the input audio packet, and then uploads the input audio packet to the corresponding forwarding server or the next-level server according to the data transmission path.
[0035] After the intercom server obtains the mixed audio packet, it determines the data transmission path corresponding to the obtained mixed audio packet based on any first audio tag in the obtained mixed audio packet, and then sends the mixed audio packet to the corresponding intercom device or the next-level server according to the data transmission path.
[0036] This application embodiment describes the actions performed by the intercom server during the intercom process. The intercom server can continuously acquire input audio packets and mixed audio packets, and identify the task and transmission path corresponding to the audio packets according to the audio tags in the audio packets, and then correctly send each audio packet to the corresponding address. Therefore, in this application embodiment, the intercom server can support multiple intercom tasks to conduct intercom at the same time, which improves the utilization efficiency of the server and the user experience.
[0037] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0038] Within a preset time interval, the forwarding server acquires several input audio packets;
[0039] Based on the first audio tag of each input audio packet, determine the intercom task to which each input audio packet belongs and the corresponding data transmission path;
[0040] Several input audio packets for the same intercom task are merged into a mixed audio packet;
[0041] Each of the mixed audio packets is sent to the corresponding intercom device or the next-level server according to its corresponding data transmission path.
[0042] This application describes the actions performed by the forwarding server during intercom communication. The forwarding server mixes several input audio packets acquired within a preset time interval to generate mixed audio packets corresponding to each intercom task, and sends each mixed audio packet down based on the corresponding transmission path. This setting enables the forwarding server to support multiple intercom tasks simultaneously, improving server utilization efficiency and user experience.
[0043] Furthermore, when a certain intercom device has multiple intercom tasks, a task queue is created based on the task priority of each intercom task, and the intercom tasks are performed sequentially according to the task queue.
[0044] This application embodiment further considers the multi-tasking of the intercom device. Since there is a task priority setting in each intercom task, the intercom device can create a task queue based on the task priority in each intercom task, sort the intercom tasks, and then perform intercom tasks in sequence according to the task queue, so as to avoid missing intercom tasks or long-term occupation of low-priority intercom tasks, and further improve the user experience.
[0045] Furthermore, during the intercom communication process, if a new intercom device is added or an intercom device exits, the intercom task is updated through the forwarding server and distributed to each intercom server, so that each intercom server updates its intercom task and updates the audio packet receiving format of the corresponding intercom device according to the updated intercom task.
[0046] In this embodiment, thanks to audio tag management, new intercom devices can be added or removed at any time during an intercom task without interrupting the current intercom connection. Each intercom server updates the intercom task and audio packet reception format in real time, ensuring that newly added intercom devices can participate in the intercom promptly, further improving the user experience.
[0047] Furthermore, the first audio tag corresponding to each of the intercom devices is the MAC address of each of the intercom devices.
[0048] In this embodiment, the MAC address of the device is used as the first audio tag corresponding to the intercom device, which facilitates quick identification of the device to which the audio stream belongs by packet capture during troubleshooting, making the status and flow of audio data more intuitive during intercom.
[0049] Secondly, correspondingly, this application provides a multi-party intercom system based on audio tag management, including an acquisition module, a server determination module, a task distribution module, and an intercom module;
[0050] The acquisition module is used to acquire intercom tasks, which include task identifiers, each intercom device participating in the intercom, a first audio tag corresponding to each intercom device, and task priority.
[0051] The server determination module is used to determine a forwarding server, several intercom servers and a data transmission path according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices.
[0052] The task distribution module is used to send the intercom task to each intercom server through the forwarding server according to the data transmission path, so that each intercom server registers its corresponding intercom device according to the intercom task, and sets the audio packet upload format and audio packet reception format of its corresponding intercom device according to each first audio tag.
[0053] The intercom module is used to transmit audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers, and the forwarding server to conduct the multi-party intercom. Attached Figure Description
[0054] Figure 1: A flowchart illustrating a multi-party intercom method based on audio tag management provided in an embodiment of this application.
[0055] Figure 2: A schematic diagram showing the distribution of the server and intercom devices in a multi-party intercom method based on audio tag management provided in an embodiment of this application.
[0056] Figure 3: A schematic diagram of the data structure of a multi-party intercom method based on audio tag management provided in an embodiment of this application.
[0057] Figure 4: A schematic diagram showing the change of control output commands when adding or removing an intercom device in a multi-party intercom method based on audio tag management provided in an embodiment of this application.
[0058] Figure 5: A schematic diagram of the structure of a multi-party intercom system based on audio tag management provided in an embodiment of this application. Detailed Implementation
[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0060] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0061] Example 1:
[0062] As shown in Figure 1, Embodiment 1 provides a multi-party intercom method based on audio tag management, including steps S1-S4:
[0063] Step S1: Obtain the intercom task, wherein the intercom task includes a task identifier, each intercom device participating in the intercom, and a first audio tag corresponding to each intercom device;
[0064] Step S2: Determine a forwarding server, several intercom servers, and a data transmission path according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices;
[0065] Step S3: According to the data transmission path, the intercom task is sent to each of the intercom servers through the forwarding server, so that each of the intercom servers registers its corresponding intercom device according to the intercom task and sets the audio packet transmission format of its corresponding intercom device according to each first audio tag.
[0066] Step S4: Through each of the intercom devices, each of the intercom servers and the forwarding server, audio packets are transmitted according to the data transmission path to conduct the multi-party intercom.
[0067] This application provides a multi-party intercom method based on audio tag management. It determines several intercom servers participating in the intercom and a forwarding server responsible for forwarding audio packets through intercom tasks. The forwarding server distributes the intercom task to each intercom server, and the intercom devices involved in the intercom task are registered to establish a multi-party intercom connection. Simultaneously, each intercom server sets its own audio packet upload and reception formats based on the first audio tag in the intercom task, realizing multi-party intercom based on audio tag management. This ensures that during the intercom process, each audio packet is transmitted in a format preset by the intercom server. Therefore, each intercom server can quickly parse the audio packets, determine the intercom task to which the audio packet belongs and the data transmission path based on the audio tag, and achieve rapid forwarding of the audio packets. Because audio tags are set, the forwarding server does not need to generate corresponding mixed audio streams for different intercom devices. Instead, it can directly mix audio packets with the same task. After receiving the mixed audio packets, each intercom device can compare the received audio packet format with the audio tags carried by the mixed audio packets to filter the mixed audio packets, avoid echoes and other noises, improve the server's processing efficiency of audio data and the user experience.
[0068] In this embodiment, the intercom devices are configured under a server. A single intercom system can contain multiple servers at the same or different levels (upper and lower levels). Each server is configured with a site number (001, 002, etc.). Site information is synchronized from lower to upper levels. If there are peer-to-peer platform information at the top-level server, they are synchronized and updated together. Since intercom device data is only synchronized unidirectionally from lower-level servers to upper-level servers, only the top-level server can directly initiate intercoms with intercom devices at any site. Other servers can only initiate intercoms with intercom devices on that server and its subordinate servers. The preemption information is distributed by the initiating upper-level server (i.e., the forwarding server). This preemption information typically includes a unique task identifier (used for subsequent operations such as adjusting volume, adding or deleting intercom users), audio tags, task priority, and information about the included intercom devices. This preemption information is the intercom task. If it is necessary for each site to freely initiate intercoms, data from each site is uploaded according to the site number to avoid duplicate data synchronization, and the upper-level server needs to synchronize device information with the lower-level devices. When a lower-level server initiates a cross-site intercom, the information is pushed layer by layer through the upload to the upper-level server (i.e., the forwarding server). Therefore, whether the intercom is initiated by an upper-level server or a lower-level server, it is necessary to determine the forwarding server and the intercom server, and the forwarding server issues the intercom task. It should be noted that in some cases, the forwarding server itself can also be an intercom server. For example, among the servers participating in the intercom, if there is an intercom server that is the upper-level server of other intercom servers, then that intercom server can be directly used as a forwarding server. When an intercom task is established, the task is decomposed starting from the uppermost server, i.e., the forwarding server. It registers audio tags for itself and its subordinate servers (registering the audio streams to be received and being responsible for receiving and forwarding them to the lower-level servers and the intercom devices on this site). After receiving the registration information, the lower-level server decomposes the task again until all intercom devices are registered.
[0069] Furthermore, in step S1, determining the forwarding server, several intercom servers, and data transmission path based on the intercom task includes:
[0070] Based on the topological connections between all servers, a tree diagram is constructed, where each server is a node in the tree diagram;
[0071] Based on the server ownership information of each intercom device participating in the intercom task, several intercom servers are determined.
[0072] Based on the position of each intercom server in the tree diagram, determine the least common ancestor node of each intercom server;
[0073] The server corresponding to the least common ancestor node is used as the forwarding server;
[0074] Based on the tree diagram, the shortest path from the forwarding server to each of the intercom servers is determined, and the shortest path is used as the data transmission path.
[0075] To more quickly determine the transmission path of preemption information in cross-level intercom, this application embodiment constructs a tree (if there are multiple top-level servers, all top-level servers are grouped under one system node) and confirms the path through the nearest common ancestor or depth-first search.
[0076] This application provides a method for confirming transmission paths. Since the servers in this embodiment are connected in a hierarchical tree structure, direct communication between them may be impossible. Therefore, a common ancestor server is needed to distribute and acquire audio packets, i.e., the forwarding server. Thus, after acquiring an intercom task, it is first necessary to determine each intercom server, the forwarding server, and the data transmission path. This application sets the least common ancestor node of each intercom server as the forwarding server and uses the shortest path from the forwarding server to each intercom server as the data transmission path, minimizing data transmission distance, reducing voice intercom latency, and improving the user experience.
[0077] In one possible implementation, when any first intercom server among the intercom servers sets the audio packet upload format and audio packet reception format of the corresponding first intercom device according to the respective first audio tags, each of the intercom servers sets the audio packet transmission format of its respective corresponding intercom device according to the respective first audio tags, including:
[0078] A second audio tag set corresponding to the first intercom device is determined based on each of the first audio tags, wherein the second audio tag set is a tag set composed of each of the first audio tags corresponding to each non-first intercom device;
[0079] A control input command is generated based on the first audio tag and audio packet sending address corresponding to the first intercom device, and the control input command is sent to the first intercom device to set the audio packet upload format of the first intercom device.
[0080] A control output command is generated based on the second audio tag set and audio packet receiving address corresponding to the first intercom device, and the control output command is sent to the first intercom device to set the audio packet receiving format of the first intercom device.
[0081] This application embodiment takes a first intercom server as an example and provides a method for setting the audio packet upload format and audio packet reception format of an intercom device. By setting a second audio tag set, the audio tags of the first intercom device itself are distinguished from the audio tags of other intercom devices. By sending control input commands to the first intercom device, the audio packet sending address and the first audio tag carried by the intercom device when uploading audio packets are set. By sending control output commands to the first intercom device, the intercom device is set to only receive audio packets with audio tags from the second audio tag set and whose audio packet reception address is used, thus filtering out other audio packets and avoiding echoes and other noise. At the same time, the server does not need to perform targeted processing on mixed audio packets, but can directly use the intercom device for audio filtering, improving the server's processing efficiency of audio data and the user experience.
[0082] In a preferred embodiment, as shown in Figure 2, taking the upper-level server 001 and its two lower-level servers 002 and 003 as an example (the upper-level server also functions as a forwarding server), when a multi-party intercom is initiated, the upper-level server 001 is responsible for decomposing the intercom task. It directly prioritizes and preempts intercom device A within the same station. After successful preemption, it sends control input commands to intercom device A, thereby setting the audio packet upload format for intercom device A. Intercom device A then uploads the audio stream to its parent server 001 according to the audio packet upload format. For intercom devices (B, C, D) belonging to lower-level servers (002, 003), the upper-level server 001 notifies the lower-level servers (002, 003) to preempt intercom devices (B, C, D). After successful contention, the intercom devices (B, C, D) on the lower-level servers (002, 003) send control input commands to each other. Each device uploads its audio stream to its respective lower-level server (002, 003) according to its audio packet upload format. Upon receiving the audio data, the lower-level servers (002, 003) forward it to the higher-level server (001). The higher-level server (001) is responsible for merging and distributing the audio streams. When an invited intercom device accepts the invitation, the server sends a control output command to that device. The audio identifier in the command is the input audio identifier of the other devices in the intercom (multiple identifiers exist for three-way or more intercoms). This completes the intercom control process where the control input command determines the audio identifier of the uploaded audio packet, and the control output command determines the audio packet received by the intercom device. The data structure of the control input command, control output command, and audio packet is shown in Figure 3.
[0083] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0084] Intercom devices acquire user voice input;
[0085] The intercom device converts the voice input into a corresponding input audio packet according to the corresponding audio packet upload format;
[0086] The intercom device uploads the corresponding input audio packet to the corresponding intercom server, so that each intercom server uploads the input audio packet to the forwarding server;
[0087] The intercom device receives a mixed audio packet sent by a corresponding intercom server. The mixed audio packet is obtained by the forwarding server merging the various input audio packets and then sending it to each of the intercom servers through the forwarding server.
[0088] The intercom device converts the mixed audio packet into several corresponding output audio packets according to the corresponding audio packet receiving format;
[0089] The intercom device outputs voice based on a number of corresponding output audio packets.
[0090] This application describes the actions performed by the intercom device during the intercom process. During voice input, the intercom device packages the user's voice input into a corresponding input audio packet according to the audio packet upload format, and uploads it to the forwarding server through the intercom server. During voice output, the intercom device parses the obtained mixed audio packet according to the audio packet receiving format, obtains several output audio packets, and performs voice output based on the several output audio packets, thereby realizing multi-party intercom based on audio tag management.
[0091] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0092] The intercom server continuously acquires input audio packets and mixed audio packets;
[0093] After the intercom server obtains the input audio packet, it determines the data transmission path corresponding to the input audio packet based on the first audio tag of the input audio packet, and then uploads the input audio packet to the corresponding forwarding server or the next-level server according to the data transmission path.
[0094] After the intercom server obtains the mixed audio packet, it determines the data transmission path corresponding to the obtained mixed audio packet based on any first audio tag in the obtained mixed audio packet, and then sends the mixed audio packet to the corresponding intercom device or the next-level server according to the data transmission path.
[0095] This application embodiment describes the actions performed by the intercom server during the intercom process. The intercom server can continuously acquire input audio packets and mixed audio packets, and identify the task and transmission path corresponding to the audio packets according to the audio tags in the audio packets, and then correctly send each audio packet to the corresponding address. Therefore, in this application embodiment, the intercom server can support multiple intercom tasks to conduct intercom at the same time, which improves the utilization efficiency of the server and the user experience.
[0096] In one possible implementation, the multi-party intercom transmission via the respective intercom devices, intercom servers, and forwarding servers, according to the data transmission path, includes:
[0097] Within a preset time interval, the forwarding server acquires several input audio packets;
[0098] Based on the first audio tag of each input audio packet, determine the intercom task to which each input audio packet belongs and the corresponding data transmission path;
[0099] Several input audio packets for the same intercom task are merged into a mixed audio packet;
[0100] Each of the mixed audio packets is sent to the corresponding intercom device or the next-level server according to its corresponding data transmission path.
[0101] This application describes the actions performed by the forwarding server during intercom communication. The forwarding server mixes several input audio packets acquired within a preset time interval to generate mixed audio packets corresponding to each intercom task, and sends each mixed audio packet down based on the corresponding transmission path. This setting enables the forwarding server to support multiple intercom tasks simultaneously, improving server utilization efficiency and user experience.
[0102] In a preferred embodiment, when an intercom task is established, each intercom server records the IP addresses (servers or local devices) to be sent by its own site and all subordinate audio tags. After receiving the audio packet, the server forwards the audio packet to the corresponding address according to the recorded audio tag and the corresponding IP address.
[0103] In a preferred embodiment, the intercom process can also include a ringing prompt stage. When ringing prompts are enabled, during the preemption stage, a control output command is first sent to the invited intercom device, and the server sends the corresponding ringing prompt audio data to the invited device. When the invited device joins the intercom, control input commands and control output commands are sent to that intercom device again, and control output commands are sent to other devices in the intercom task to update the audio identifier (adding the audio tag of the newly joined intercom device to the command).
[0104] In a preferred embodiment, since each intercom device carries its own audio tag in its audio stream, other devices can process different audio streams according to the reception and individual needs. For example, filtering a specified audio tag can achieve targeted mute, recording, and volume adjustment effects. Specifically: Mute: Filtering the audio stream with the specified audio tag achieves a mute effect. Recording: Directly saving the audio stream with the specified audio tag achieves a recording effect. Volume Adjustment: By modifying the audio volume data in the audio information of the audio stream with the specified audio tag, the intercom device can adjust the output volume according to the audio volume in the audio information during playback. Similarly, each intercom device can also process its own audio stream to achieve mute, recording, and volume adjustment effects. Similarly, the server can process the received audio stream to achieve mute, recording, volume adjustment, and sound effect changes. It should be noted that after audio processing, the audio tag in the audio header should not be changed (or the audio header should be truncated before processing and inserted before sending).
[0105] In a preferred embodiment, the server can control the source volume and output volume of the device via control input commands and control output commands, respectively. The control input commands can change the source volume of the audio information in the audio stream uploaded by the device; other devices receiving the audio stream can obtain the source volume when parsing the audio information. The control output commands can control the output volume of the device, and the final output volume is the product of the input volume and the output volume.
[0106] If external devices need to be connected, the audio stream is mixed by the server before being sent to the external device via a common format (MP3, WAV). An audio header needs to be added to the received audio. If the external device needs to change the output audio quality, the audio needs to be processed and mixed by the server before output.
[0107] Furthermore, when a certain intercom device has multiple intercom tasks, a task queue is created based on the task priority of each intercom task, and the intercom tasks are performed sequentially according to the task queue.
[0108] A single system can support multiple intercom tasks simultaneously, with task processing done on demand. Common multi-tasking solutions are as follows:
[0109] I. Prioritize tasks according to their priority.
[0110] Each intercom task has a task priority setting (whether tasks with the same priority can preempt each other can be set as needed). Higher priority tasks can preempt lower priority tasks. A preempted device leaves the lower priority task and joins the higher priority task.
[0111] 2. The user decides whether to enter a new intercom task.
[0112] To facilitate confirmation of new intercom information, the intercom task name, initiator, and other information can be generated into audio using text-to-speech (TTS) software. This audio can then be sent to the device as a notification tone via control output commands. If a separate notification tone for a new task is desired, the control output command should only include the audio tag for that new task notification tone. If the output of the original intercom task is to be unaffected, the audio tag for the new intercom task notification tone can be added to the control output command. The output effect of the new intercom notification tone can be altered by changing the audio information in the audio stream. For devices with a display screen, the server can also send new intercom information to the device for display.
[0113] Furthermore, if it's necessary to automatically return to the previous intercom task after it ends, the server needs to add a task queue for each intercom device on its server. When an intercom task ends, it should return to the specified intercom task according to priority or by user selection. When an intercom device returns, control output commands need to be updated for other devices, and the audio stream from the new intercom device needs to be forwarded to other devices. Whether returning to intercom requires initiator confirmation can be configured through server settings or task settings (set as needed).
[0114] This application embodiment further considers the multi-tasking of the intercom device. Since there is a task priority setting in each intercom task, the intercom device can create a task queue based on the task priority in each intercom task, sort the intercom tasks, and then perform intercom tasks in sequence according to the task queue, so as to avoid missing intercom tasks or long-term occupation of low-priority intercom tasks, and further improve the user experience.
[0115] Furthermore, during the intercom communication process, if a new intercom device is added or an intercom device exits, the intercom task is updated through the forwarding server and distributed to each intercom server, so that each intercom server updates its intercom task and updates the audio packet receiving format of the corresponding intercom device according to the updated intercom task.
[0116] In a preferred embodiment, the intercom task ends when one party leaves a two-person intercom. When there are three or more parties in an intercom, the task does not end when one party leaves; the remaining intercom devices update their control output commands (removing the audio tag of the departing party). The intercom task ends when only one party remains. Intercom devices can terminate the intercom task by notifying the server via a protocol or directly through the server's web or client interface. The permission to terminate the intercom task is allocated as needed. During an intercom task, devices already in the intercom can invite other devices to join. After a new device joins, the other devices need to update their control output commands and add the audio tag of the newly joined device. For scenarios where intercom devices join or leave an intercom task, taking intercom device B as an example, the changes in the control output commands of intercom device B are shown in Figure 4. In addition to the control output commands, the corresponding audio stream should also be sent to intercom device B simultaneously with the command to prevent the intercom device from missing audio packets from other intercom devices.
[0117] In this embodiment, thanks to audio tag management, new intercom devices can be added or removed at any time during an intercom task without interrupting the current intercom connection. Each intercom server updates the intercom task and audio packet reception format in real time, ensuring that newly added intercom devices can participate in the intercom promptly, further improving the user experience.
[0118] Furthermore, the first audio tag corresponding to each of the intercom devices is the MAC address of each of the intercom devices.
[0119] In this embodiment, the MAC address of the device is used as the first audio tag corresponding to the intercom device, which facilitates quick identification of the device to which the audio stream belongs by packet capture during troubleshooting, making the status and flow of audio data more intuitive during intercom.
[0120] Example 2:
[0121] As shown in Figure 5, correspondingly, Embodiment 2 provides a multi-party intercom system based on audio tag management, including an acquisition module 10, a server determination module 20, a task distribution module 30, and an intercom module 40.
[0122] The acquisition module 10 is used to acquire intercom tasks, which include task identifiers, each intercom device participating in the intercom, a first audio tag corresponding to each intercom device, and task priority.
[0123] The server determination module 20 is used to determine a forwarding server, several intercom servers and a data transmission path according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices.
[0124] The task distribution module 30 is used to send the intercom task to each of the intercom servers through the forwarding server according to the data transmission path, so that each of the intercom servers registers its corresponding intercom device according to the intercom task, and sets the audio packet upload format and audio packet reception format of its corresponding intercom device according to each first audio tag.
[0125] The intercom module 40 is used to transmit audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers and the forwarding server to conduct the multi-party intercom.
[0126] This application provides a multi-party intercom system based on audio tag management. It determines several intercom servers participating in the intercom and a forwarding server responsible for forwarding audio packets through intercom tasks. The forwarding server distributes the intercom task to each intercom server, and the intercom devices involved in the intercom task are registered to establish a multi-party intercom connection. Simultaneously, each intercom server sets its own audio packet upload and reception formats based on the first audio tag in the intercom task, realizing multi-party intercom based on audio tag management. This ensures that during the intercom process, each audio packet is transmitted in a format preset by the intercom server. Therefore, each intercom server can quickly parse the audio packet, determine the intercom task to which the audio packet belongs and the data transmission path based on the audio tag, and achieve rapid forwarding of the audio packet. Because audio tags are set, the forwarding server does not need to generate corresponding mixed audio streams for different intercom devices. Instead, it can directly mix audio packets with the same task. After receiving the mixed audio packets, each intercom device can compare the received audio packet format with the audio tags carried by the mixed audio packets to filter the mixed audio packets, avoid echoes and other noises, improve the server's processing efficiency of audio data and the user experience.
[0127] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.
[0128] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.
Claims
1. A multi-party intercom method based on audio tag management, characterized in that, include: Obtain an intercom task, wherein the intercom task includes a task identifier, each intercom device participating in the intercom, and a first audio tag corresponding to each intercom device; The forwarding server, several intercom servers, and data transmission path are determined according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices; According to the data transmission path, the intercom task is sent to each of the intercom servers through the forwarding server, so that each of the intercom servers registers its corresponding intercom device according to the intercom task and sets the audio packet transmission format of its corresponding intercom device according to each first audio tag. The multi-party intercom is conducted by transmitting audio packets through each of the intercom devices, each of the intercom servers, and the forwarding server according to the data transmission path.
2. The multi-party intercom method based on audio tag management as described in claim 1, characterized in that, The step of determining the forwarding server, several intercom servers, and data transmission path based on the intercom task includes: Based on the topological connections between all servers, a tree diagram is constructed, where each server is a node in the tree diagram; Based on the server ownership information of each intercom device participating in the intercom task, several intercom servers are determined. Based on the position of each intercom server in the tree diagram, determine the least common ancestor node of each intercom server; The server corresponding to the least common ancestor node is used as the forwarding server; Based on the tree diagram, the shortest path from the forwarding server to each of the intercom servers is determined, and the shortest path is used as the data transmission path.
3. The multi-party intercom method based on audio tag management as described in claim 1, characterized in that, When any first intercom server among the intercom servers sets the audio packet upload format and audio packet reception format of the corresponding first intercom device according to the respective first audio tags, each of the intercom servers sets the audio packet transmission format of its respective intercom device according to the respective first audio tags, including: A second audio tag set corresponding to the first intercom device is determined based on each of the first audio tags, wherein the second audio tag set is a tag set composed of each of the first audio tags corresponding to each non-first intercom device; A control input command is generated based on the first audio tag and audio packet sending address corresponding to the first intercom device, and the control input command is sent to the first intercom device to set the audio packet upload format of the first intercom device. A control output command is generated based on the second audio tag set and audio packet receiving address corresponding to the first intercom device, and the control output command is sent to the first intercom device to set the audio packet receiving format of the first intercom device.
4. The multi-party intercom method based on audio tag management as described in claim 1, characterized in that, The multi-party intercom, conducted by transmitting audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers, and the forwarding server, includes: Intercom devices acquire user voice input; The intercom device converts the voice input into a corresponding input audio packet according to the corresponding audio packet upload format; The intercom device uploads the corresponding input audio packet to the corresponding intercom server, so that each intercom server uploads the input audio packet to the forwarding server; The intercom device receives a mixed audio packet sent by a corresponding intercom server. The mixed audio packet is obtained by the forwarding server merging the various input audio packets and then sending it to each of the intercom servers through the forwarding server. The intercom device converts the mixed audio packet into several corresponding output audio packets according to the corresponding audio packet receiving format; The intercom device outputs voice based on a number of corresponding output audio packets.
5. A multi-party intercom method based on audio tag management as described in claim 1, characterized in that, The multi-party intercom, conducted by transmitting audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers, and the forwarding server, includes: The intercom server continuously acquires input audio packets and mixed audio packets; After the intercom server obtains the input audio packet, it determines the data transmission path corresponding to the input audio packet based on the first audio tag of the input audio packet, and then uploads the input audio packet to the corresponding forwarding server or the next-level server according to the data transmission path. After the intercom server obtains the mixed audio packet, it determines the data transmission path corresponding to the obtained mixed audio packet based on any first audio tag in the obtained mixed audio packet, and then sends the mixed audio packet to the corresponding intercom device or the next-level server according to the data transmission path.
6. The multi-party intercom method based on audio tag management as described in claim 1, characterized in that, The multi-party intercom, conducted by transmitting audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers, and the forwarding server, includes: Within a preset time interval, the forwarding server acquires several input audio packets; Based on the first audio tag of each input audio packet, determine the intercom task to which each input audio packet belongs and the corresponding data transmission path; Several input audio packets for the same intercom task are merged into a mixed audio packet; Each of the mixed audio packets is sent to the corresponding intercom device or the next-level server according to its corresponding data transmission path.
7. A multi-party intercom method based on audio tag management as described in any one of claims 1-6, characterized in that, When a certain intercom device has multiple intercom tasks, a task queue is created based on the task priority of each intercom task, and the intercom tasks are performed sequentially according to the task queue.
8. A multi-party intercom method based on audio tag management as described in any one of claims 1-6, characterized in that, During the intercom communication process, if a new intercom device is added or an intercom device exits, the intercom task is updated through the forwarding server and sent to each intercom server so that each intercom server updates its intercom task and updates the audio packet receiving format of the corresponding intercom device according to the updated intercom task.
9. A multi-party intercom method based on audio tag management as described in any one of claims 1-6, characterized in that, The first audio tag corresponding to each of the intercom devices is the MAC address of each of the intercom devices.
10. A multi-party intercom system based on audio tag management, characterized in that, It includes an acquisition module, a server determination module, a task distribution module, and an intercom module; The acquisition module is used to acquire intercom tasks, which include task identifiers, each intercom device participating in the intercom, a first audio tag corresponding to each intercom device, and task priority. The server determination module is used to determine a forwarding server, several intercom servers and a data transmission path according to the intercom task, wherein the intercom server is a preset server for each of the intercom devices. The task distribution module is used to send the intercom task to each intercom server through the forwarding server according to the data transmission path, so that each intercom server registers its corresponding intercom device according to the intercom task, and sets the audio packet upload format and audio packet reception format of its corresponding intercom device according to each first audio tag. The intercom module is used to transmit audio packets according to the data transmission path through each of the intercom devices, each of the intercom servers, and the forwarding server to conduct the multi-party intercom.
Citation Information
Patent Citations
One-to-many talkback method and device
CN109963108A
Multi-party intercom call method, device and system and electronic equipment
CN111327667A
Voice intercom service implementation method and device, and storage medium
CN114374729A
Message transmission method and device, electronic equipment and computer readable medium
CN117176649A
Multi-party talkback method and system based on audio tag management
CN119071274A