Control method for transmitting audio data across regions based on a wide area network and related apparatus

CN122293650BActive Publication Date: 2026-08-21SOYO TECH DEV CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610746225.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-21
Estimated Expiration
2046-05-28

AI Technical Summary

Technical Problem

[0002]当前的物联设备在广域网双向音频传输场景中通常采用基于单一端口复用多路复用进行音频传输,或者是基于负载均衡策略选择转发服务器节点,可能存在多设备组并发易引发数据冲突的问题

Benefits of technology

本申请提供了一种基于广域网的跨地域传输音频数据的控制方法及相关装置,应用于音频传输控制系统的调度服务器,其中,音频传输控制系统还包括设置在多个区域的多个转发服务器,与多个转发服务器通信连接的终端,多个转发服务器与调度服务器通信连接,该方法包括:获取终端的设备信息和网络地址信息,以及多个转发服务器的多个第一负载参数;根据设备信息和网络地址信息生成用于传输音频数据的虚拟设备标识和传输标识;基于传输标识和虚拟设备标识确定传输音频数据的音频传输的链路信息,得到多个音频传输链路信息;根据多个第一负载参数和多个音频传输链路信息构建多个转发服务器传输音频数据的多个音频传输通道;获取多个音频传输通道的网络状态参数;根据网络状态参数对预设的稳定传输阈值进行稳定性判断,得到稳定性状态结果;若稳定性状态结果为不稳定状态时,根据音频数据生成音频编码指令,和/或,根据网络状态参数确定多个音频传输通道中音频数据的音频调整发送参数;若稳定性状态结果为稳定状态时,控制多个转发服务器通过多个音频传输通道将音频数据转发至所述终端。如此,一方面,通过调度服务器并发隔离与任务分配,通过构建音频传输通道传输音频数据,从而避免在多设备音频传输的时候出现数据冲突;另一方面,转发服务器根据网络状态参数(传输带宽、丢包率等),动态的将音频信号进行码率转换并传输,以及控制转发服务器的发送任务,从而减少了交互的音频卡顿问题,以提升了传输的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293650B_ABST
    Figure CN122293650B_ABST
Patent Text Reader

Abstract

The application provides a control method for transmitting audio data across regions based on a wide area network and related devices, the method comprising: obtaining device information, network address information of a terminal, and load parameters of a plurality of forwarding servers, determining link information of audio transmission according to the device information and the network address information, and constructing a plurality of audio transmission channels for the plurality of forwarding servers to transmit audio data according to the plurality of first load parameters and the plurality of audio transmission link information; and obtaining network state parameters of the plurality of audio transmission channels; performing stability judgment on a preset stable transmission threshold according to the network state parameters, if the state is unstable, generating an audio encoding instruction according to the audio data, and / or determining audio adjustment sending parameters of the audio data according to the network state parameters; if the state is stable, controlling the plurality of forwarding servers to forward the audio data to the terminal through the plurality of audio transmission channels. In this way, the stability of audio transmission is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio transmission technology, and in particular to a control method and related apparatus for cross-regional transmission of audio data based on a wide area network. Background Technology

[0002] Current IoT devices typically employ single-port multiplexing for audio transmission in wide area network (WAN) bidirectional audio transmission scenarios, or use load balancing strategies to select forwarding server nodes. This can lead to data conflicts due to multiple devices operating concurrently. Furthermore, bandwidth fluctuations in WANs, such as unstable 4G / 5G signals and public network congestion, can cause audio stuttering and packet loss, affecting audio transmission from IoT devices. Current WAN-based bidirectional audio transmission methods lack efficient concurrency isolation, bandwidth adaptation, and fault tolerance mechanisms, making it difficult to guarantee transmission stability and consistent audio quality in high-concurrency scenarios. They also cannot adapt to cross-carrier and cross-regional application requirements, severely limiting the large-scale deployment of traditional IoT devices for WAN bidirectional interaction.

[0003] Therefore, in cross-regional audio transmission, it is urgent to solve the problems of improving the stability of audio transmission and reducing data conflicts. Summary of the Invention

[0004] This application provides a control method and related apparatus for cross-regional audio data transmission based on a wide area network. By adopting a distributed architecture of a scheduling server, a storage server, and a forwarding server, the scheduling server provides concurrent isolation and task allocation. By constructing audio transmission channels and distributing audio data to multiple audio transmission channels for transmission, data conflicts are avoided when audio is transmitted from multiple devices. In addition, the forwarding server dynamically converts the bitrate of the audio signal and transmits it based on network status parameters (transmission bandwidth, packet loss rate, etc.), and controls the sending tasks of the forwarding server, thereby reducing audio stuttering issues during interaction and improving transmission stability.

[0005] In a first aspect, this application provides a control method for cross-regional transmission of audio data based on a wide area network, applied to a scheduling server of an audio transmission control system. The audio transmission control system further includes multiple forwarding servers located in multiple regions, and a terminal communicatively connected to the multiple forwarding servers. The multiple forwarding servers are communicatively connected to the scheduling server. The method includes: Obtain the device information and network address information of the terminal; and the multiple first load parameters of the multiple forwarding servers; A virtual device identifier and a transmission identifier for transmitting audio data are generated based on the device information and the network address information; Based on the transmission identifier and the virtual device identifier, the audio transmission link information for transmitting the audio data is determined, resulting in multiple audio transmission link information. Based on the multiple first load parameters and the multiple audio transmission link information, multiple audio transmission channels for transmitting audio data are constructed by the multiple forwarding servers, resulting in multiple audio transmission channels; and network status parameters of the multiple audio transmission channels are obtained. The stability of the preset stable transmission threshold is determined based on the network state parameters to obtain the stability state result. If the stability state result is unstable, an audio encoding instruction is generated based on the audio data, and / or, the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels are determined based on the network state parameters; If the stability status result is stable, control the multiple forwarding servers to forward the audio data to the terminal through the multiple audio transmission channels.

[0006] Secondly, this application provides a control device for cross-regional transmission of audio data based on a wide area network, applied to a scheduling server of an audio transmission control system. The audio transmission control system further includes multiple forwarding servers located in multiple areas, and a terminal communicatively connected to the multiple forwarding servers. The multiple forwarding servers are communicatively connected to the scheduling server. The device includes: The acquisition unit is used to acquire the device information and network address information of the terminal; and multiple first load parameters of the multiple forwarding servers; The determining unit is configured to generate a virtual device identifier and a transmission identifier for transmitting audio data based on the device information and the network address information; and determine the audio transmission link information for transmitting the audio data based on the transmission identifier and the device identifier, thereby obtaining multiple audio transmission link information. The processing unit is configured to construct multiple audio transmission channels for transmitting audio data from the multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information; and to obtain network status parameters of the multiple audio transmission channels; The control unit is configured to determine the stability of a preset stable transmission threshold based on the network status parameters, and obtain a stability status result; if the stability status result is unstable, generate an audio encoding instruction based on the audio data, and / or determine the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters; if the stability status result is stable, control the plurality of forwarding servers to forward the audio data to the terminal through the plurality of audio transmission channels.

[0007] Thirdly, this application provides a server including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any of the methods of the first aspect of the embodiments of this application.

[0008] Fourthly, this application provides a control for cross-regional transmission of audio data based on a wide area network, wherein the control system for cross-regional transmission of audio data based on a wide area network can perform some or all of the steps described in any of the methods in the first aspect of the embodiments of this application.

[0009] Fifthly, this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of the embodiments of this application. The computer program product may be a software installation package.

[0010] By implementing this application, the following beneficial effects can be achieved: This application provides a control method and related apparatus for cross-regional audio data transmission based on a wide area network (WAN), applied to a scheduling server of an audio transmission control system. The audio transmission control system further includes multiple forwarding servers located in multiple regions, terminals communicatively connected to the multiple forwarding servers, and the multiple forwarding servers communicatively connected to the scheduling server. The method includes: acquiring device information and network address information of the terminals, and multiple first load parameters of the multiple forwarding servers; generating virtual device identifiers and transmission identifiers for transmitting audio data based on the device information and network address information; determining audio transmission link information for transmitting audio data based on the transmission identifiers and virtual device identifiers, obtaining multiple audio transmission link information; constructing multiple audio transmission channels for the multiple forwarding servers to transmit audio data based on the multiple first load parameters and the multiple audio transmission link information; acquiring network status parameters of the multiple audio transmission channels; performing a stability judgment on a preset stable transmission threshold based on the network status parameters, obtaining a stability status result; if the stability status result is unstable, generating audio encoding instructions based on the audio data, and / or determining audio adjustment transmission parameters for the audio data in the multiple audio transmission channels based on the network status parameters; if the stability status result is stable, controlling the multiple forwarding servers to forward the audio data to the terminals through the multiple audio transmission channels. Thus, on the one hand, by scheduling the server to isolate concurrency and allocate tasks, and by building an audio transmission channel to transmit audio data, data conflicts can be avoided when multiple devices transmit audio. On the other hand, the forwarding server dynamically converts the bitrate of the audio signal and transmits it based on network status parameters (transmission bandwidth, packet loss rate, etc.), and controls the sending tasks of the forwarding server, thereby reducing audio stuttering issues during interaction and improving transmission stability. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an architecture diagram of an audio transmission control system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a server provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a control method for cross-regional audio data transmission based on a wide area network, as provided in an embodiment of this application. Figure 4 This is a flowchart illustrating an audio transmission control method for transmission failures provided in an embodiment of this application. Figure 5 This is a flowchart illustrating another control method for cross-regional audio data transmission based on a wide area network, provided in an embodiment of this application. Figure 6 This is a schematic diagram of the structure of a control method for cross-regional transmission of audio data based on a wide area network, provided in an embodiment of this application. Figure 7 This is a block diagram of the functional modules of a control device for cross-regional transmission of audio data based on a wide area network, provided in an embodiment of this application. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0014] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0015] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.

[0016] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.

[0017] In this application embodiment, "connection" refers to various connection methods such as direct connection or indirect connection to realize communication between devices. This application embodiment does not limit this in any way.

[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] In wide area network (WAN) scenarios, most IoT devices (such as speakers and smart panels) engage in two-way audio interaction via the public network or use a server-forwarding model for audio data transmission. Few designs specifically address the characteristics of WANs and high-concurrency scenarios. However, using the public network for two-way audio interaction or the conventional server-forwarding model for audio data transmission suffers from several drawbacks. The conventional server-forwarding model lacks concurrency isolation mechanisms, leading to data conflicts during audio transmission from multiple devices. Furthermore, using the public network for two-way audio interaction is susceptible to network congestion and fluctuations, causing audio transmission stuttering and other issues.

[0020] To address the aforementioned issues, this application provides a control method and related apparatus for cross-regional audio data transmission based on a wide area network. By employing a distributed architecture of a scheduling server, a storage server, and a forwarding server, and with the scheduling server implementing concurrent isolation and task allocation, audio data is transmitted through an audio transmission channel, thus avoiding data conflicts during multi-device audio transmission. Furthermore, the forwarding server dynamically converts the bitrate of the audio signal and transmits it based on network status parameters (transmission bandwidth, packet loss rate, etc.), and controls the forwarding server's sending tasks, thereby reducing audio stuttering issues during interaction and improving transmission stability.

[0021] The following is combined with Figure 1 The system architecture of a control method for cross-regional audio data transmission based on a wide area network, as described in an embodiment of this application, is explained. Figure 1 This is an architecture diagram of an audio transmission control system provided in an embodiment of this application. The audio transmission control system 100 includes an Internet of Things terminal device 110, a forwarding server cluster 120, a storage server 130, and a scheduling server 140.

[0022] The IoT terminal device 110 is used to collect, generate, and transmit audio data, and establish a communication connection with the forwarding server cluster 120 via a network. The IoT terminal device 110 can be a terminal device with audio acquisition or playback functions, such as a smart intercom device, network camera device, mobile terminal, or smart voice terminal, etc., and is not limited thereto. The IoT terminal device 110 can collect audio signals locally, encode the audio signals into digital audio data, and send them to the forwarding server cluster 120 via a wide area network, thereby initiating a cross-regional audio communication request. In one possible embodiment, the IoT terminal device 110 includes an audio acquisition module, which is used to collect sound signals in the environment and convert the collected analog audio signals into digital audio data. The audio acquisition module may include a microphone unit and an audio encoding unit, which encodes and compresses the collected audio signals to form an audio data stream suitable for network transmission, and sends the audio data stream to the forwarding server cluster 120 for further processing. In addition, the IoT terminal device 110 may also include an audio playback module for receiving audio data from the network, decoding and playing the received audio data to achieve functions such as remote voice communication or audio broadcasting. Thus, the IoT terminal device 110 can realize the acquisition, transmission, reception, and playback of audio data, thereby constituting a terminal node in the audio data transmission system.

[0023] The forwarding server cluster 120 is used to receive, process, and forward audio data from the IoT terminal device 110. The forwarding server cluster 120 consists of multiple forwarding server nodes, which can be distributed across data centers in different regions, forming a cross-regional server network structure. When the IoT terminal device 110 sends audio data, the forwarding server cluster 120 can select a suitable forwarding server node to process the audio data according to the scheduling strategy of the scheduling server 140, and forward the audio data to the target terminal device or target server node. By adopting a server cluster, the processing capacity and concurrent processing capacity of the audio data transmission system can be effectively improved, thereby meeting the needs of large-scale terminal device access. In one possible embodiment, each forwarding server node in the forwarding server cluster 120 can also perform data fragmentation, timestamp marking, and integrity verification on the audio data to improve the stability and reliability of audio data transmission in a wide area network environment. Simultaneously, when a server node is detected to be overloaded or experiencing network anomalies, the corresponding audio transmission task can be migrated to other server nodes, thereby achieving node-level fault tolerance.

[0024] The storage server 130 is used to store and manage relevant data generated during audio transmission. It can store audio data files, audio transmission logs, and system operating status information. For example, when the audio transmission control system 100 needs to record or replay the audio communication process, it can store the corresponding audio data stream or audio data segments in the storage server 130 for subsequent querying or analysis. Simultaneously, the storage server 130 also stores statistical data generated during audio transmission, such as network latency information, packet loss rate information, and server load information, providing data support for system scheduling and optimization.

[0025] The scheduling server 140 is used for unified scheduling and management of the entire audio transmission control system 100. The scheduling server 140 establishes communication connections with the forwarding server cluster 120 and the IoT terminal device 110 to obtain real-time network status information, server operating status information, and audio transmission status information. In one possible embodiment, the scheduling server 140 can also construct a global network status view of the system, and dynamically adjust the audio transmission strategy by comprehensively analyzing information such as network bandwidth usage, data packet loss rate, and processing load reported by each forwarding server node. For example, when a server node is detected to have a high load, the scheduling server 140 can reallocate some audio transmission tasks to server nodes with lower loads, thereby achieving balanced distribution of system load.

[0026] As can be seen, through the system architecture of the above-described control method for cross-regional audio data transmission based on wide area networks, the audio transmission control system 100 can achieve efficient communication between IoT terminal devices 110 and forwarding server cluster 120 in a wide area network environment. It can also uniformly schedule each server node through scheduling server 140 and store and manage relevant data using storage server 130, thereby forming a distributed audio transmission system that supports cross-regional audio data transmission. This system can effectively support cross-regional audio data transmission and control based on wide area networks, and realize efficient collaboration and stable interaction of audio data between nodes in different regions.

[0027] The following is combined with Figure 2 The server in the embodiments of this application will be described. Figure 2 This is a schematic diagram of the structure of a server provided in an embodiment of this application, such as... Figure 2 As shown, the server 200 includes a processor 210, a memory 220, a communication interface 230, and one or more programs 221. The processor 210 is communicatively connected to the memory 220 and the communication interface 230 via an internal communication bus.

[0028] The one or more programs 221 are stored in the memory 220 and configured to be executed by the processor 210. The one or more programs 221 include instructions for performing any step in the above method embodiments.

[0029] The processor 210 can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, units, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication unit can be a communication interface, transceiver, transceiver circuit, etc., and the storage unit can be a memory.

[0030] The memory 220 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0031] It is understood that server 200 may include more or fewer structural elements than those shown in the block diagram above, such as a power module, physical buttons, a Wi-Fi module, a speaker, a Bluetooth module, sensors, a display module, etc., without limitation. It is understood that server 200 may be equipped with... Figure 1 The architecture of an audio transmission control system is described above.

[0032] After understanding the software and hardware architecture of this application, the following will be combined with... Figure 3 This application describes a control method for cross-regional audio data transmission based on a wide area network (WAN). Figure 3 This is a flowchart illustrating a control method for cross-regional audio data transmission based on a wide area network, provided in an embodiment of this application. The method is applied to a scheduling server of an audio transmission control system. The audio transmission control system also includes multiple forwarding servers located in multiple regions, and terminals communicatively connected to the multiple forwarding servers. The multiple forwarding servers are communicatively connected to the scheduling server. The method specifically includes the following steps: Step S310: Obtain the device information and network address information of the terminal; and the multiple first load parameters of the multiple forwarding servers.

[0033] The term "terminal" refers to an IoT terminal device connected to a wide area network (WAN), which may include devices with two-way audio interaction capabilities such as speakers and smart panels, without further limitation. Device information refers to the terminal's inherent identification information, while network address information refers to the terminal's network location and access attribute information after connecting to the WAN. The first load parameter is a real-time operating load-related indicator of the forwarding server.

[0034] Specifically, the scheduling server first initiates an information collection request to all IoT terminals connected to the wide area network. Upon receiving the request, each terminal returns its own device information and network address information. The device information includes the terminal's factory MAC address, device number, device model, and other identifying information unrelated to network access status; this information is used for terminal identification and grouping. The network address information includes the terminal's public IP address, network access type (Wi-Fi / 4G / 5G), signal strength, and access region, among other dynamic network access information. Simultaneously, the scheduling server establishes real-time communication connections with the forwarding server clusters distributed across various regions. It collects first load parameters from each forwarding server according to a preset period. These first load parameters specifically include the forwarding server's CPU utilization, bandwidth utilization, number of currently accessing audio transmission tasks, and network connectivity status. Each forwarding server transmits these load parameters back to the scheduling server in real time, and the scheduling server then summarizes and processes the load status of all forwarding servers.

[0035] It should be noted that the collection of device information and network address information is an initial collection and real-time update mode after the terminal connects to the wide area network. After initial access, the terminal reports basic information to the scheduling server. If the terminal undergoes network switching (e.g., switching from Wi-Fi to 4G) or geographical relocation, it will update the network address information to the scheduling server in real time to ensure information timeliness. The first load parameter of the forwarding server adopts a high-frequency periodic collection mode, with the collection cycle matching the response requirements of subsequent load scheduling. This reflects the dynamic changes in server load and avoids scheduling errors caused by lagging load information. Furthermore, all collected information is uniformly stored and categorized by the scheduling server, and a lightweight communication protocol is used during information transmission to avoid excessive network resource consumption during information collection, which could affect the normal transmission of audio data.

[0036] Step S320: Generate a virtual device identifier and a transmission identifier for transmitting audio data based on the device information and the network address information.

[0037] The virtual device identifier is a unique virtual identifier assigned to the terminal by the scheduling server. It is used to virtualize the terminal's identity during transmission and forms a unique mapping relationship with the terminal's physical device information. The transmission identifier is an identifier adapted for audio data transmission and is used to identify audio data packets during data transmission.

[0038] Specifically, the scheduling server first performs structured parsing of the collected terminal device information and network address information. Based on the uniqueness of the device information, it completes terminal identity verification and grouping, dividing terminals with audio interaction needs into several device groups. Then, it performs IP address mapping processing based on the public IP addresses in the network address information, establishing a correspondence between the terminal's physical identity, network address, and virtual identity. For each individual terminal or terminal device group after division, the scheduling server assigns a unique virtual device identifier according to the identifier generation rules. This identifier uses a fixed-format encoding form, possessing non-repeatability and identifiability. Simultaneously, a corresponding transmission identifier is generated based on the virtual device identifier. The transmission identifier uses a lightweight encoding design, adapting to the field embedding requirements of the audio data packet header. After completing the identifier generation, the scheduling server associates and binds the virtual device identifier and the transmission identifier, forming a mapping relationship of "device information - network address information - virtual device identifier - transmission identifier," and stores this mapping relationship uniformly to ensure the accuracy and efficiency of identifier parsing during subsequent audio data transmission.

[0039] As can be seen, by generating virtual device identifiers and transmission identifiers based on terminal device information and network address information, an audio data virtual identifier system adapted to high-concurrency wide area network scenarios has been constructed. The virtual device identifier separates the physical identity of the terminal from the virtualized transmission identity, facilitating subsequent dynamic access and scheduling management of terminals. The generation and binding of the transmission identifier gives each audio data stream an identity tag, laying the foundation for accurate identification and forwarding of subsequent audio data packets. At the same time, the identifier allocation mode based on device groups realizes the logical differentiation of audio data from multiple terminal groups, avoiding the conflict problem of audio data from different terminals / terminal groups in high-concurrency scenarios from the data source, and providing an identifier basis for the subsequent isolation construction of transmission links.

[0040] Step S330: Based on the transmission identifier and the virtual device identifier, determine the audio transmission link information for transmitting the audio data, and obtain multiple audio transmission link information.

[0041] The audio transmission link information is a set of logical link-related parameters planned by the scheduling server for audio data transmission between the terminal and the forwarding server. This includes: identification association information, data forwarding direction information, and matching relationship information between the terminal and the forwarding server. Each link information is bound to a transmission identifier and a virtual device identifier to ensure the uniqueness and isolation of the link.

[0042] Specifically, the scheduling server first uses the mapping relationship of "device information - network address information - virtual device identifier - transmission identifier" to index the transmission identifier and virtual device identifier corresponding to each terminal or device group as the link information. For a single indexed terminal or device group, the scheduling server combines the access region, network type, and other attributes in its network address information to plan the basic data transmission path from the terminal to the corresponding regional forwarding server, clarifying the topology structure where the starting node of the link is the audio data sending terminal, the intermediate node is the corresponding forwarding server, and the ending node is the audio data receiving terminal. Then, the scheduling server embeds the transmission identifier and virtual device identifier into the planned basic path parameters, establishing the correspondence between identifiers and paths, and configuring basic data forwarding rules for the link, forming complete audio transmission link information containing identifier information, node matching information, path topology information, and forwarding rule information. For high-concurrency transmission scenarios with multiple terminals or device groups, the scheduling server independently plans the link information for each terminal or device group according to the same process. Finally, it generates multiple independent audio transmission link information and classifies, stores, and numbers all link information to ensure the accuracy of subsequent calls.

[0043] As can be seen, by identifying multiple independent audio transmission link information based on transmission identifiers and virtual device identifiers, the limitations of traditional audio transmission link planning—such as the lack of dedicated identifier associations and the ease of confusion among multiple links—are overcome, thus achieving the binding of logical links with device groups. Each link information and identifier combination is matched, allowing the forwarding server to quickly match the corresponding link information by parsing the identifiers in the data during subsequent audio data transmission, improving the accuracy and efficiency of data forwarding. Furthermore, link planning combined with terminal network address information ensures that the links are adapted to the actual network access attributes of the terminals, effectively shortening the audio data transmission path and reducing the basic latency of cross-regional transmission. The structured and standardized set of link information improves the link planning and management efficiency of the entire transmission control system, ensuring the orderliness and stability of cross-regional audio data transmission over wide area networks.

[0044] Step S340: Construct multiple audio transmission channels for the multiple forwarding servers to transmit audio data based on the multiple first load parameters and the multiple audio transmission link information; and obtain the network status parameters of the multiple audio transmission channels.

[0045] The audio transmission channel is the physical communication link between the terminal and the forwarding server, enabling bidirectional audio data transmission. It's a hardware and protocol combination built upon the audio transmission link information and the actual load status of the forwarding server, directly capable of carrying audio data transmission. Each transmission channel corresponds to a unique audio transmission link, and the channels are independent of each other, possessing data transmission isolation. Network status parameters are indicators characterizing the communication quality, transmission capacity, and stability of the audio transmission channel. These include end-to-end transmission capacity, data transmission stability, and basic connection quality, among other multi-dimensional characteristics, serving as the basis for determining whether a channel meets the requirements for stable audio data transmission.

[0046] Specifically, the process begins by quantitatively analyzing and determining the load levels of the collected primary load parameters of multiple forwarding servers. This clarifies the real-time business processing capacity and remaining load capacity of forwarding servers in each region, allowing for the selection of available forwarding server nodes capable of handling audio transmission tasks. Next, the scheduling server retrieves information from multiple generated audio transmission links. Based on the scheduling principle of "geographical proximity and load balancing," it matches the link information with available forwarding server nodes, prioritizing mapping the audio transmission links of terminals or device groups to the forwarding server corresponding to the terminal's access region. If the server's load in that region is close to a threshold, the link information is mapped to a backup forwarding server node in the same region. After node matching is complete, the scheduling server configures independent network transmission protocols and communication ports for each matching relationship based on the identifier association rules and data forwarding rules in the link information. It binds the transmission identifier and virtual device identifier to the physical communication port of the forwarding server, and simultaneously configures data reception and forwarding buffers for each channel in the forwarding server. For all matched link information, the scheduling server sequentially completes the channel setup process, generating multiple independent audio transmission channels. It then uniformly registers the node matching relationships, port configuration information, and identifier binding information for each channel, forming a complete transmission channel management ledger to ensure accurate transmission of audio data along the planned path. Following this, the scheduling server employs a collaborative data acquisition mode of "terminal-forwarding server bidirectional reporting - target server active probing" to obtain network status parameters. First, the terminal reports its own access-side network status data to the scheduling server at a preset 100ms interval, including available network bandwidth and Wi-Fi / 4G / 5G signal strength for the end-to-end transmission link between the terminal and the target forwarding server. Simultaneously, target forwarding servers distributed across various regions synchronously collect real-time data transmission status within the audio transmission channels at 100ms intervals, reporting stability indicators such as audio data packet loss rate and transmission latency. Second, based on the bidirectional reporting data, the scheduling server sends standardized probe data packets to each audio transmission channel to verify channel connectivity and data transmission round-trip latency, supplementing and correcting the accuracy of the data reported by the terminal and the forwarding server. The scheduling server performs structured processing on the collected multi-dimensional data, classifies and integrates it according to the audio transmission channel number, and forms a complete set of network status parameters including indicators such as network bandwidth, packet loss rate, signal strength, and transmission delay. All parameters are accompanied by a collection timestamp to ensure the timeliness and traceability of the data.

[0047] As can be seen, by combining the first load parameter of the forwarding server with the audio transmission link information to construct multiple independent audio transmission channels, the technical limitations of traditional audio transmission, such as lack of load consideration in channel construction, poor regional adaptability, and lack of isolation among multiple channels, are solved. This achieves intelligent, rational, and isolated construction of transmission channels. Its node matching based on the principle of "geographical proximity and load balancing" effectively shortens the cross-regional transmission path of audio data, reduces transmission latency, and simultaneously achieves load balancing distribution across the forwarding server cluster, avoiding the decrease in transmission efficiency caused by single-node overload and improving the overall resource utilization of the server cluster.

[0048] In one possible embodiment, constructing multiple audio transmission channels for transmitting audio data using the multiple first load parameters and the multiple audio transmission link information specifically includes the following steps: 341. Determine multiple load weight values ​​corresponding to the multiple forwarding servers based on the multiple first load parameters; 342. Determine the first regional attribute information corresponding to the terminal based on the network address information; 343. Filter forwarding server information corresponding to the first regional attribute information from the multiple audio transmission link information to obtain multiple forwarding server information; 344. Determine the priority parameters for regional audio scheduling in the multiple forwarding server information to obtain multiple audio scheduling parameters; 345. Based on the multiple load weight values, the multiple audio scheduling parameters are weighted and calculated to obtain multiple target audio scheduling parameters; 346. Determine the target forwarding servers corresponding to multiple audio transmission links based on the multiple target audio scheduling parameters, thereby obtaining multiple target forwarding servers; 347. Generate multiple transmission port information of the multiple target forwarding servers based on the multiple audio transmission link information and the transmission identifier; 348. Map the multiple audio transmission link information and the multiple transmission port information to the multiple target forwarding servers to obtain the multiple audio transmission channels, and allocate non-overlapping port segments to each of the multiple audio transmission channels to achieve port resource isolation between different audio transmission channels.

[0049] Among them, the load weight value is a quantitative representation of the real-time load status of the forwarding server, used to measure the server's ability to handle new transmission tasks. The first regional attribute information is the geographical area characteristic of the terminal accessing the wide area network, which is the basis for regional proximity scheduling. The audio scheduling parameters are the quantitative values ​​of scheduling priority in the regional dimension; the target audio scheduling parameters are the comprehensive scheduling basis after the integration of load and regional dimensions; the transmission port information are the dedicated communication port parameters allocated by the target forwarding server to each audio transmission link.

[0050] Specifically, the scheduling server first processes the primary load parameters (CPU utilization, bandwidth utilization, workload, etc.) of each forwarding server, converting different load indicators into values ​​between 0 and 1. Then, based on a preset load weight calculation model, it calculates the load weight value of each forwarding server; a higher value indicates stronger remaining server capacity. Next, based on data such as IP address location and access gateway region in the terminal's network address information, it determines the provincial / municipal geographical range of the terminal through geocoding parsing, i.e., the primary geographical attribute information. Using this primary geographical attribute information as a filtering condition, it selects forwarding servers from the list of forwarding servers associated with the audio transmission link information that are in the same or adjacent regions as the terminal, forming a candidate forwarding server information set. Then, based on indicators such as geographical distance and network latency, it configures regional audio scheduling priority parameters for the candidate forwarding servers; servers closer to the terminal and with lower latency have higher priority parameter values. Simultaneously, a weighted summation algorithm is employed to fuse the load weight value and the regional scheduling priority parameter according to a preset weight ratio (e.g., load accounts for 60%, region accounts for 40%), thereby obtaining the target audio scheduling parameters based on load and regional characteristics. The weighted calculation formula is as follows: ,in, For target scheduling parameters; This represents the load weight value. Priority parameters for regional scheduling; For hyperparameters; These are the preset weighting coefficients, and Based on the target audio scheduling parameters, the forwarding server with the highest parameter value is selected as the target forwarding server for each audio transmission link, achieving optimal node matching. The scheduling server, based on the uniqueness of the audio transmission link information and the exclusive characteristics of the transmission identifier, allocates an independent User Datagram Protocol (UDP) port segment to each target forwarding server, generating transmission port information containing port start values, end values, and port reuse rules. Finally, the audio transmission link information, transmission port information, and target forwarding server hardware resource information are mapped and bound. Independent data forwarding rules and buffers are configured for each mapping relationship in the forwarding server, ultimately generating multiple mutually isolated audio transmission channels.

[0051] As can be seen, the calculation of load weight values ​​enables accurate assessment of the forwarding server's operating status, ensuring that transmission tasks are preferentially allocated to nodes with strong remaining capacity, effectively reducing the risk of server overload and improving the overall resource utilization of the cluster. The principle of geographical proximity scheduling significantly shortens the cross-network transmission path of audio data, reduces the basic latency of WAN transmission, and improves the real-time performance of audio interaction. The integration of load and geographical dimensions achieves multi-objective optimization of scheduling decisions, ensuring transmission efficiency and server load balance. The allocation and mapping binding of UDP port segments physically isolates multiple transmission channels, completely resolving the port reuse conflict problem of audio data from different terminals or device groups in high-concurrency scenarios. In this way, intelligent, balanced, and isolated construction of WAN cross-regional audio transmission channels is realized, significantly improving the scheduling accuracy and transmission stability of the transmission control system, and providing reliable communication support for bidirectional audio interaction of large-scale IoT terminals.

[0052] In one possible embodiment, determining the multiple load weight values ​​corresponding to the multiple forwarding servers based on the multiple first load parameters specifically includes the following steps: 3411. Obtain multiple rated load parameters of the multiple forwarding servers, and record the multiple rated load parameters as multiple second load parameters; 3412. Determine the difference between each of the plurality of first load parameters and each of the plurality of second load parameters to obtain a plurality of load parameter differences; 3413. Calculate the load weight values ​​corresponding to the multiple forwarding servers based on the differences of the multiple load parameters, and obtain the multiple load weight values.

[0053] Among them, the rated load parameter (second load parameter) is the maximum load threshold that the forwarding server can handle under stable operating conditions, and is a benchmark indicator for measuring the server's theoretical carrying capacity; the load parameter difference is the quantitative difference between the real-time load and the rated load, which directly reflects the server's remaining load capacity.

[0054] Specifically, the scheduling server first retrieves the factory configuration parameters and performance test reports of each forwarding server, extracting its rated load parameters (second load parameters). These parameters include indicators such as the server's rated CPU utilization (set to 85%), rated bandwidth utilization, and the maximum number of audio transmission tasks it can handle. Different regions and different configurations of forwarding servers correspond to different second load parameters. The scheduling server establishes a second load parameter ledger by server number to ensure parameter traceability. The scheduling server calculates the difference between the real-time collected first load parameters and the corresponding second load parameters. For example, for a certain forwarding server, its CPU load parameter difference equals the rated CPU utilization minus the real-time CPU utilization, and its bandwidth load parameter difference equals the rated bandwidth utilization minus the real-time bandwidth utilization. Through calculation, a multi-dimensional set of load parameter differences is obtained. A positive difference indicates that the server has remaining load capacity, while a negative difference indicates that the server is overloaded. Then, the scheduling server normalizes the differences in multi-dimensional load parameters, converting the differences in different units into standardized values ​​in the range of 0-1. Based on the preset weight allocation rules (such as CPU load accounting for 40%, bandwidth load accounting for 30%, and task load accounting for 30%), the normalized differences are weighted and summed. Finally, the comprehensive load weight value of each forwarding server is obtained. The closer the weight value is to 1, the stronger the remaining carrying capacity of the server and the higher the feasibility of undertaking new transmission tasks.

[0055] As can be seen, the refined load weight calculation method overcomes the limitations of traditional load assessment methods, such as single indicators, lack of benchmarks, and incomparability across nodes. It achieves accurate, objective, and standardized assessment of the remaining capacity of forwarding servers. This effectively ensures the stable operation of the forwarding server cluster and improves the load balancing capability and resource utilization of the entire transmission control system.

[0056] In one possible embodiment, mapping the plurality of audio transmission link information and the plurality of transmission port information to the plurality of target forwarding servers to obtain the plurality of audio transmission channels specifically includes the following steps: 3481. Based on the multiple transmission port information, determine the association information between the transmission identifier and the transmission port to obtain transmission association information; 3482. Generate physical link mapping identifiers in the plurality of target forwarding servers based on the transmission association information; 3483. Based on the multiple audio transmission link information, determine the User Datagram Protocol (UDP) port segments used to isolate port resources between different audio transmission links in the multiple audio transmission links, thus obtaining multiple port segments; 3484. Based on the preset audio data forwarding rules, an audio transmission channel is constructed for the multiple port segments and the physical link mapping identifier to obtain the multiple audio transmission channels.

[0057] Among them, the transmission association information is a unique set of correspondences between transmission identifiers and transmission ports, which is the foundation for achieving precise binding between identifiers and physical ports. The physical link mapping identifier is a physical identifier within the target forwarding server to identify different transmission channels, used for channel differentiation and data forwarding control within the server; multiple port segments are UDP port segments, which are continuous and non-overlapping port intervals allocated for different audio transmission links. They are non-overlapping in numerical range, and each port segment corresponds to only one transmission identifier to form a unique physical link mapping relationship.

[0058] Specifically, the scheduling server first extracts parameters such as port number, port type, and port availability status from the transmission port information, as well as associated information such as the encoding rules and device group of the transmission identifier. It then binds each transmission identifier to a unique transmission port by establishing key-value pairs, forming a structured transmission association information of "transmission identifier-port number-home link," ensuring that each transmission identifier can be mapped to a physical transmission port. Next, based on the transmission association information, a corresponding physical link mapping identifier is generated in the local configuration of the target forwarding server. This identifier uses a locally recognizable encoding format and includes a transmission association information index, port configuration information, and link home information, stored in the forwarding server's local cache. This facilitates the server's rapid parsing of the transmission identifier in the audio data packet and matching it with the corresponding physical link mapping identifier. Then, based on the requirement of independence of audio transmission link information, the scheduling server allocates a dedicated UDP port segment to each link. The range of the port segment is selected as a dynamic / private port segment (49152-65535) to ensure that the port segments of different links do not overlap. For example, the 50000-50001 port segment is allocated to link 1, and the 50002-50003 port segment is allocated to link 2. Through the physical isolation of port segments, complete isolation of port resources between different audio transmission links is achieved, avoiding port reuse conflicts. Finally, the scheduling server retrieves the preset audio data forwarding rules (including packet header identifier parsing rules, port forwarding rules, data buffer allocation rules, etc.), associates each port segment with the corresponding physical link mapping identifier, and configures an independent data receiving buffer, forwarding priority, and error handling mechanism for each association in the target forwarding server. This completes the construction from logical link to physical transmission channel, generates multiple mutually isolated audio transmission channels, and synchronizes the channel configuration information to the management ledger of the scheduling server.

[0059] As can be seen, by uniquely binding the transmission identifier to the transmission port, audio data packets can be quickly matched with the corresponding transmission port through the packet header identifier, which greatly improves the accuracy and efficiency of data forwarding and avoids data forwarding errors caused by mismatch between identifier and port. The generation of physical link mapping identifiers enables the target forwarding server to quickly identify different channels locally, reduces the channel parsing latency inside the server, and improves data processing efficiency. The isolated allocation of UDP port segments physically eliminates port reuse conflicts of different audio transmission links, effectively solves the data confusion problem in high-concurrency scenarios, and ensures the independence of audio data transmission from multiple terminals or device groups.

[0060] Step S350: Based on the network status parameters, perform a stability judgment on the preset stable transmission threshold to obtain a stability status result.

[0061] Among them, the preset stable transmission threshold is a set of multi-dimensional quantitative judgment standards that the scheduling server pre-sets and stores for high-concurrency audio transmission scenarios in wide area networks. It is a quantitative judgment benchmark pre-set by the scheduling server based on the business requirements and measured data of wide area network audio transmission. It includes the critical values ​​of indicators such as network bandwidth, packet loss rate, signal strength, and transmission delay (e.g., bandwidth threshold of 100kbps and packet loss rate threshold of 1%), and is the basis for judging whether the channel meets the requirements for stable audio transmission. The stability status results include stable state and unstable state.

[0062] Specifically, the scheduling server first retrieves a pre-stored stable transmission threshold system, including indicators such as network bandwidth (end-to-end available bandwidth), packet loss rate, signal strength, and transmission delay. For the network status parameters of a single audio transmission channel, the scheduling server first extracts the indicator data: if the available network bandwidth of the channel is ≥100kbps and the packet loss rate is ≤1%, the core transmission status of the channel is initially determined to be up to standard. Based on this, further verification is performed: if the signal strength is ≥ a preset signal threshold (e.g., -85dBm) and the transmission delay is ≤ a preset delay threshold (e.g., 200ms), the final determination of the channel's stability status is "stable." If any of the core indicators does not meet the threshold requirements (e.g., bandwidth <100kbps or packet loss rate >1%), it is directly determined to be "unstable."

[0063] As can be seen, the stability judgment logic of this application focuses on the core quality requirements of audio transmission (bandwidth guarantee, low packet loss) while also taking into account the basic connection quality (signal, latency), thus avoiding misjudgment problems caused by single-dimensional judgment. The judgment process takes a single audio transmission channel as an independent unit, ensuring that the state judgment of different channels is independent of each other, adapting to the technical requirements of fine-grained scheduling in high-concurrency scenarios; at the same time, the threshold system is set based on measured data and business requirements, rather than abstract theoretical thresholds, to ensure the practicality and accuracy of the judgment results.

[0064] Step S360: If the stability state result is unstable, generate an audio encoding instruction based on the audio data, and / or determine the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network state parameters.

[0065] Among them, the audio encoding instruction is a bitrate control instruction issued by the scheduling server to change the encoding of audio data. The purpose is to reduce the data volume of a single frame by adjusting the bitrate, so as to adapt to the degraded network environment with low bandwidth and high packet loss. The audio adjustment sending parameters are data packet sending control parameters determined based on the real-time network status, including sending rate, sending interval, retransmission strategy, etc., which are used to optimize the data sending rhythm and reduce the risk of transmission congestion.

[0066] In one possible embodiment, the network status parameters include network bandwidth and packet loss rate; the step of generating audio encoding instructions based on the audio data specifically includes the following steps: A1. When the packet loss rate is greater than a preset packet loss rate threshold and the network bandwidth is less than a preset bandwidth threshold, a first audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to the preset first audio bitrate; the audio encoding instruction is determined according to the first audio bitrate adjustment instruction. A2. When the packet loss rate is less than or equal to the packet loss rate threshold and the network bandwidth is greater than or equal to the bandwidth threshold, a second audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to the preset second audio bitrate; the audio encoding instruction is determined according to the second audio bitrate adjustment instruction; the first audio bitrate is twice the second audio bitrate.

[0067] The preset packet loss rate threshold is the tolerable packet loss threshold for audio transmission (set to 1%), and the preset bandwidth threshold is the minimum bandwidth threshold to ensure basic audio transmission (set to 100kbps). The first audio bitrate adjustment instruction is a low bitrate adjustment instruction triggered when the network condition deteriorates, and the first preset bitrate is a low bitrate value adapted to the deteriorated network (specifically set to 8kbps). The second audio bitrate adjustment instruction is a high bitrate adjustment instruction triggered when the network condition is good, and the second preset bitrate is a high bitrate value adapted to the stable network (specifically set to 16kbps). The audio encoding instruction is a standardized control instruction that integrates the bitrate adjustment instruction, bitrate parameters, and execution rules, and can be directly parsed and executed by the forwarding server.

[0068] Specifically, when the scheduling server detects a real-time packet loss rate > 1% and a network bandwidth < 100kbps in the audio transmission channel, it determines that the current link cannot support high-bitrate audio transmission. It then issues a first audio bitrate adjustment command to the corresponding forwarding server, reducing the bitrate of the audio data sent to the terminal to a first preset bitrate of 8kbps. This bitrate value has been verified in practice to ensure audio quality under low bandwidth and high packet loss conditions. Simultaneously, the scheduling server integrates the bitrate parameters, execution target, and effective time of this command to generate a corresponding audio encoding command and stores it in the command ledger. When the real-time packet loss rate ≤ 1% and the network bandwidth ≥ 100kbps, it determines that the link meets the conditions for high-bitrate transmission. The scheduling server issues a second audio bitrate adjustment command to the forwarding server, controlling the audio data transmission bitrate to increase to a second preset bitrate of 16kbps (twice the first preset bitrate). This bitrate value can significantly improve audio quality while ensuring smooth transmission. Simultaneously, based on this command, an audio encoding command containing high bitrate parameters is generated, completing the reverse adaptation of the encoding strategy.

[0069] It is evident that the tiered audio bitrate adjustment strategy based on dual thresholds of bandwidth and packet loss rate effectively solves the problem of "imbalance between fluency and sound quality" in audio transmission caused by fluctuations in WAN link status, and achieves dynamic adaptation of audio encoding strategy and network status, significantly improving the system's adaptability to complex WAN link environments.

[0070] In one possible embodiment, the network status parameters include: network strength and network latency parameters; determining the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters specifically includes the following steps: B1. Obtain the first audio transmission rate and the first audio transmission interval parameters of the audio data; B2. Determine the network quality assessment value corresponding to each of the multiple audio transmission channels based on the network strength and the network latency parameters, thereby obtaining multiple network quality assessment values; B3. Determine the audio transmission level corresponding to the multiple network quality assessment values ​​to obtain multiple audio transmission levels; B4. Determine the target audio transmission rate and target audio transmission interval parameters corresponding to the target audio transmission level; the target audio transmission level is any one of the plurality of audio transmission levels; B5. Determine the adjusted transmission rate of the audio data based on the target audio transmission rate and the first audio transmission rate; and determine the adjusted transmission interval parameter of the audio data based on the first audio transmission interval parameter and the target audio transmission interval parameter. B6. Determine the audio adjustment transmission parameters based on the adjusted transmission rate, the adjusted transmission interval parameter, and the network latency parameter.

[0071] Among them, network strength is the wireless signal strength index of the terminal accessing the wide area network, network delay parameter is the end-to-end transmission delay index of audio data in the transmission channel, network quality assessment value is the quantitative score after normalized weighted calculation of network strength and network delay, used to characterize the quality of channel communication; audio transmission level is the gradient transmission adaptation level divided according to the network quality assessment value; first transmission rate and first transmission interval parameters are the initial default transmission configuration of audio data, target transmission rate and target transmission interval parameters are the standard transmission configuration matching the target transmission level, adjusted transmission rate and adjusted transmission interval parameters are the actual execution parameters after differential calibration, and audio adjusted transmission parameters are the final control parameter set integrating rate, interval and delay characteristics.

[0072] Specifically, the scheduling server obtains the original first transmission rate and first transmission interval parameters of the default configuration for audio data, normalizes the network strength and network latency parameters, and calculates an independent network quality assessment value for each audio transmission channel using a preset weighted model. Higher signal strength and lower network latency result in a higher network quality assessment value. Next, according to preset level classification rules, the network quality assessment value is mapped to the corresponding audio transmission level, forming a multi-gradient adaptation system. Then, a target audio transmission level matching the current channel's network quality is selected, and the preset standard target transmission rate and target transmission interval parameters for that level are retrieved. Then, using the original parameters as a baseline and the target parameters as the optimization direction, the adjusted transmission rate and adjusted transmission interval parameters adapted to the current network state are calculated, achieving a smooth parameter transition. Finally, the adjusted rate, interval parameters, and network latency parameters are integrated to comprehensively determine the final audio adjustment transmission parameters, ensuring that the transmission strategy perfectly matches the channel's transmission characteristics.

[0073] It is evident that the gradient-based and refined method for determining audio transmission parameters effectively solves the problems of audio stuttering, interruption, and excessive latency caused by weak signals and high latency in wide area networks (WANs). Network quality quantification assessment enables precise determination of the transmission channel status, and graded transmission levels allow for differentiated adjustment strategies for networks with varying degrees of degradation, ensuring transmission efficiency and audio smoothness. Parameter smoothing calibration avoids data packet loss or transmission blockage caused by sudden changes in transmission parameters, improving the stability of audio transmission. Thus, the system's adaptability to complex WAN environments is significantly enhanced, providing reliable transmission control guarantees for stable, low-latency transmission of audio data from cross-regional IoT terminals.

[0074] In one possible embodiment, determining the audio adjustment transmission parameters based on the adjusted transmission rate, the adjusted transmission interval parameter, and the network latency parameter specifically includes the following steps: B61. Obtain the audio transmission parameters of the plurality of audio transmission channels; and the second load parameters of the forwarding server corresponding to the plurality of audio transmission channels; B62. When the second load parameter is greater than or equal to the preset load parameter threshold, an audio task splitting instruction is generated according to the audio transmission parameters. B63. Based on the audio task splitting instruction, determine the audio transmission allocation task of the multiple audio transmission channels to obtain multiple first audio transmission tasks; B64. Determine the audio transmission rate of the plurality of first audio transmission tasks based on the network latency parameters to obtain a plurality of audio transmission rates; B65. Determine the audio adjustment transmission parameter based on the plurality of audio transmission rates, the adjusted transmission rate, and the adjusted transmission interval parameter.

[0075] Among them, the audio transmission parameters are the link status, transmission identifier, port segment configuration, and other operational parameters of the audio transmission channel; the second load parameters are the CPU utilization, bandwidth utilization, number of concurrent tasks, and other load indicators of the forwarding server in real time; the preset load parameter threshold is the server overload threshold preset by the scheduling server, used to determine whether the forwarding server is in an overloaded state; the audio task diversion instruction is the task scheduling instruction triggered when the server is overloaded, used to migrate some transmission tasks of the overloaded node to a lightly loaded node in the same area; the first audio transmission task is the audio transmission task reassigned after task diversion; the audio transmission rate is the sending rate adapted to network latency and the task allocation results after diversion; the audio adjustment sending parameters are the comprehensive control parameters that integrate task diversion, transmission rate, and sending interval.

[0076] Specifically, firstly, the scheduling server synchronously collects the audio transmission parameters of each audio transmission channel, as well as the second load parameters of the corresponding forwarding server, achieving synchronous perception of the network and server status. Simultaneously, the scheduling server compares the second load parameters with a preset load parameter threshold. When the load parameter is greater than or equal to the threshold, it determines that the forwarding server is overloaded and then generates an audio task distribution instruction based on the audio transmission parameters to prevent transmission blockage or interruption due to server overload. Next, based on the audio task distribution instruction, the original audio transmission tasks are redistributed according to the principle of "geographical proximity and load balancing," migrating some tasks from overloaded servers to idle servers in the same area, resulting in multiple evenly distributed first audio transmission tasks. Then, for each distributed first audio transmission task, an appropriate audio transmission rate is determined based on the network latency parameters of the corresponding transmission channel. The lower the network latency, the higher the transmission rate is configured; the higher the network latency, the lower the appropriate rate is used to ensure smooth transmission. Finally, the audio transmission rate after splitting and adaptation, the previously determined adjusted transmission rate, and the adjusted transmission interval parameters are fused and calculated to ultimately determine the audio adjustment transmission parameters that take into account network conditions, server load, and task allocation.

[0077] As can be seen, by integrating the load status of the forwarding server and the network status to determine the audio transmission parameters, the coordinated optimization of transmission parameter adjustment and server load balancing is achieved, effectively solving the audio stuttering and interruption problems caused by the superposition of network degradation and server overload. The load threshold judgment and task diversion mechanism can promptly alleviate the transmission pressure of overloaded servers, preventing the servers from reducing forwarding efficiency due to overload operation, and improving the resource utilization and operational stability of the server cluster. Combined with network latency reconfiguration of transmission rate, the diverted tasks can accurately adapt to the network characteristics of the corresponding channels, ensuring low latency and smooth audio transmission. The multi-dimensional parameter determination method makes the audio transmission parameter adjustment more in line with the actual transmission environment, greatly improving the system's adaptability to complex wide area network scenarios, and providing comprehensive protection for the stable and efficient transmission of cross-regional high-concurrency audio data.

[0078] Step S370: If the stability status result is stable, control the multiple forwarding servers to forward the audio data to the terminal through the multiple audio transmission channels.

[0079] The stability status result refers to a stable state, meaning that the network status parameters of the audio transmission channel, such as network bandwidth, packet loss rate, signal strength, and network latency, all meet the preset stable transmission threshold requirements, and the corresponding forwarding server load is within a reasonable range, indicating that the channel has the ability to continuously and stably carry audio data transmission. The multiple audio transmission channels are mutually isolated physical transmission channels built in the early stage based on virtual device identifiers, transmission identifiers, and UDP port segments. Controlling the forwarding server to execute forwarding actions is a standardized execution process in which the scheduling server issues instructions based on the channel status and preset forwarding rules, and the forwarding server completes the reception, parsing, and targeted delivery of audio data.

[0080] Specifically, when the scheduling server determines that the stability status of one or all audio transmission channels is stable, it immediately issues an audio data forwarding execution instruction to the target forwarding server in the corresponding area. Upon receiving the instruction, the forwarding server receives the audio data stream from the sending terminal in real time. By parsing the transmission identifier and virtual device identifier carried in the audio data packet header, it matches the pre-stored physical link mapping identifier and UDP port segment information to locate the audio transmission channel corresponding to the receiving terminal. In high-concurrency scenarios, audio data from different device groups is forwarded in parallel through mutually isolated audio transmission channels, eliminating port resource contention and data confusion issues between different channels. The forwarding server, according to preset audio data forwarding rules, sends the audio data to the target receiving terminal through the designated transmission port segment, while maintaining the temporal integrity of the data frames to ensure that the audio data received by the terminal is continuous and error-free. The entire forwarding process does not require dynamic optimization operations such as encoding adjustments, rewriting transmission parameters, or task splitting; efficient transmission can be achieved solely through stable transmission channels and standardized forwarding logic.

[0081] As can be seen, when the audio transmission channel is in a stable state, directly controlling the forwarding server to complete data forwarding eliminates the need for dynamic adjustment of encoding or sending parameters, effectively reducing system control overhead and data processing latency, and ensuring the real-time performance of audio interaction. Parallel forwarding based on a two-layer isolation mechanism fundamentally avoids the conflict and confusion of audio data from multiple terminals in high-concurrency scenarios, ensuring the orderliness of transmission. Furthermore, relying on geographically proximate and load-balanced forwarding channels, it further reduces cross-regional transmission latency and improves data delivery success rate.

[0082] With the above Figure 3 For embodiments that are consistent with those shown, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating an audio transmission control method for transmission failures provided in an embodiment of this application. As can be seen, the method specifically includes the following steps: S410. Obtain the response information and running status information of the response heartbeat frames of the multiple forwarding servers; S420. Based on the response information and the running status information, determine the transmission status of each of the multiple forwarding servers to obtain multiple transmission statuses; S430. When the target transmission state is in failure, the forwarding server corresponding to the target transmission state is recorded as the faulty forwarding server, and the faulty forwarding server is controlled to suspend the execution of the second audio transmission task. S440. Obtain the first forwarding server in the same area as the faulty forwarding server, and the network connectivity between the first forwarding server and the terminal; S450, Generate audio session transmission migration instructions based on network connectivity; S460. According to the audio session transmission migration instruction, cache the second audio transmission task corresponding to the faulty forwarding server to the first forwarding server; S470. Establish an audio transmission channel between the first forwarding server and the terminal to obtain an audio migration transmission channel; S480. The audio data corresponding to the second audio transmission task is forwarded to the terminal through the audio migration transmission channel.

[0083] The response information for the heartbeat frame consists of heartbeat response data fed back by the forwarding server at a preset period, including connectivity indicators such as heartbeat frame reception latency and frame loss rate; the operating status information includes hardware and software operating data such as the CPU operating status, network interface status, and process survival status of the forwarding server; the transmission status is divided into two categories: normal and fault, and the fault status is determined by the heartbeat frame timeout without response, the response frame loss rate exceeding the threshold, or abnormal operating status parameters; the first forwarding server is a backup forwarding node in the same region as the faulty forwarding server, which has the remaining load capacity to undertake audio transmission tasks; the audio session transmission migration instruction is a control instruction that triggers task migration, including information such as task identifier, caching rules, and channel reconstruction parameters; the second audio transmission task is the audio data forwarding task being executed by the faulty server; the audio migration transmission channel is a temporary audio transmission channel with isolation attributes reconstructed based on the fault switching requirements, and the audio migration transmission channel reuses the original virtual device identifier, transmission identifier, and UDP port segment mapping relationship to maintain the same identifier and port isolation attributes as the original audio transmission channel.

[0084] Specifically, the scheduling server sends heartbeat detection frames to all forwarding servers every 300ms, simultaneously collecting heartbeat response information and real-time operating status information from each server. At the same time, the scheduling server compares the response information with preset heartbeat thresholds (e.g., response latency > 500ms, frame drop rate > 5%), and, combined with the anomaly judgment logic of the operating status information, judges the transmission status of each forwarding server, clearly defining the boundary between faulty and normal nodes. If the transmission status of a forwarding server is determined to be faulty, it is marked as a faulty forwarding server, and a pause command is issued to stop it from executing the second audio transmission task, preventing the faulty node from continuously sending invalid data and causing abnormal terminal reception. Simultaneously, based on the topology information of the same area and the network connectivity between the first forwarding server and the terminal, a joint evaluation is performed to determine the target first forwarding server to undertake the migration task. Specifically, its network connectivity with the terminal can be verified through port probing, connectivity testing, etc., to ensure that it has the capability to undertake the task. Next, based on the connectivity test results, an audio session transmission migration instruction is generated. This instruction includes the identifier information of the second audio transmission task, data caching rules, channel reconstruction parameters, etc. According to the migration instruction, the incomplete second audio transmission task from the faulty server is fully cached to the first forwarding server. A sequential caching mechanism is used to ensure the temporal integrity of the audio data, preventing data loss or out-of-order delivery during task migration. Then, the channel construction logic of "identifier and port isolation" is reused to establish a dedicated audio migration transmission channel between the first forwarding server and the terminal. This channel has the same logical identifier binding and UDP port segment isolation characteristics as the original channel. Finally, the first forwarding server forwards the cached audio data corresponding to the second audio transmission task to the terminal through the audio migration transmission channel, completing the audio transmission recovery after the failover.

[0085] As can be seen, this embodiment effectively solves the technical bottleneck of transmission interruption caused by single point of failure of server in traditional audio transmission systems by means of forwarding server fault detection and audio task migration. Compared with the current technical method of simply switching between primary and backup nodes based on heartbeat detection, this embodiment determines the target node for migration by jointly evaluating the topology and network connectivity in the same area, avoiding the high latency problem caused by cross-region migration. At the same time, the original virtual device identifier, transmission identifier and UDP port segment mapping relationship are reused during the migration process, so that the terminal side does not need to rebuild the connection, realizing the seamless migration of audio sessions. And through sequential caching and queue reorganization mechanism, the timing consistency and data integrity of audio frames are guaranteed during the migration process. High-frequency heartbeat monitoring and multi-dimensional operational status awareness can quickly and accurately identify faulty servers, avoiding audio anomalies caused by the continuous participation of faulty nodes in transmission; the selection and connectivity verification of the first forwarding server in the same area ensure low latency and high reliability of task migration, meeting the real-time requirements of cross-regional transmission in wide area networks; the task caching mechanism and sequential forwarding strategy ensure the integrity and timing of audio data during fault switching, with no obvious audio dropouts or stuttering on the terminal side; the migration channel reuses a unified identifier and port isolation logic, which is highly compatible with the overall system architecture, balancing fault handling efficiency and transmission isolation.

[0086] For easier understanding, please refer to Figure 5 , Figure 5 This is a flowchart illustrating another control method for cross-regional audio data transmission based on a wide area network (WAN) provided in this application embodiment. This method achieves stable transmission and efficient management of audio data in a cross-regional environment through coordinated control of multiple steps, including initialization management of the audio transmission system, network status monitoring, transmission quality optimization, and global dynamic scheduling. Specifically, as... Figure 5As shown, in step S501, the "initialization, concurrent isolation and task allocation" process is executed first. First, when an audio transmission task starts, the system initializes and configures the relevant service modules and establishes a multi-task processing environment to support the parallel processing of multiple audio transmission tasks. Simultaneously, a concurrent isolation mechanism is used to isolate and manage different audio transmission tasks to avoid resource contention or mutual interference between them. Then, based on the audio transmission requests generated by each terminal device, the corresponding audio transmission tasks are allocated, allowing each audio task to be distributed to the corresponding processing node or forwarding server for processing, thus providing a basic operating environment for the stable transmission of subsequent audio data. After completing system initialization, the next process, step S502, is executed: "Dynamic bandwidth adaptation; terminals and forwarding server nodes report bandwidth usage, packet loss rate, and other data at 100ms intervals, and the scheduling server monitors in real time." Specifically, during cross-regional audio transmission, each terminal device and its corresponding forwarding server node periodically collects current network transmission status information, such as bandwidth usage and key indicators like packet loss rate. Network status data is reported periodically at 100ms intervals and transmitted to the scheduling server. By monitoring and analyzing this network status data in real time, the scheduling server can dynamically grasp the network quality of each audio transmission link, thus providing a basis for adjusting subsequent audio data transmission strategies. Next, step S503 is executed: "Transmission Quality Optimization: The forwarding server implements packet retransmission based on the RTP protocol, adds timestamps and checksums to audio data fragments, and the receiving end reassembles them in order." Specifically, during cross-regional audio data transmission, the forwarding server manages audio data transmission using the RTP (Real-time Transport Protocol). When audio data packets are detected to be lost or abnormal, the lost data can be retransmitted through the packet retransmission mechanism, thereby reducing the sound quality degradation caused by packet loss during audio data transmission. Simultaneously, before sending the audio data, the system fragments the audio data and adds a corresponding timestamp and checksum to each data fragment. The timestamp identifies the transmission order and time information of the data fragments, while the checksum verifies data integrity. After receiving multiple data fragments, the receiving end reassembles the audio data fragments in sequence according to the timestamp information and verifies the correctness of the data through a checksum, thereby recovering the complete audio data stream. During audio data transmission, to ensure the overall reliability and stability of the system, step S504, "Node Fault Tolerance and Failover: Load Balancer Real-Time Monitoring of Forwarding Node Status," is further executed. In this step, the load balancer continuously monitors the operating status of each forwarding server node, such as node processing load, network connection status, and operational stability.When an anomaly or performance degradation is detected in a forwarding node, the corresponding audio transmission task can be migrated to other available nodes through a load balancing mechanism, thereby achieving node-level fault tolerance and failover. Finally, step S505 is executed: "Global Dynamic Scheduling: The scheduling server constructs a global transmission quality view, tracking the latency and packet loss rate of each audio stream in real time." Specifically, the scheduling server summarizes and analyzes the status information reported by each terminal device, forwarding server node, and network link to construct a global transmission quality view for the entire system. This view reflects the key performance indicators of each audio data transmission link in real time, such as audio transmission latency and data packet loss rate. Based on this global view, the scheduling server can continuously track the running status of each audio transmission task and dynamically adjust the audio transmission path or task allocation strategy according to real-time network conditions to further improve the overall efficiency and stability of the audio transmission system.

[0087] Thus, this embodiment achieves stable cross-regional audio data transmission in a wide area network environment through the coordinated efforts of multiple steps, including initialization concurrency management, dynamic bandwidth monitoring, transmission quality optimization, node fault tolerance mechanism, and global dynamic scheduling.

[0088] For easier understanding, please refer to Figure 6 , Figure 6 This is a schematic diagram of a control method for cross-regional transmission of audio data based on a wide area network, as provided in an embodiment of this application. Specifically, as shown... Figure 6As shown, the structure mainly includes a scheduling server, multiple cross-regional deployed forwarding servers, and terminal devices distributed in different regions. It achieves stable transmission and unified scheduling control of cross-regional audio data through a wide area network, thus constructing an audio data transmission and control system for cross-regional scenarios. Specifically, a scheduling server is deployed at the center of the network system, primarily responsible for unified scheduling and control of nodes in various regions. In this embodiment, the scheduling server establishes network communication connections with the forwarding servers deployed in each region, uniformly managing the operating status, network load, and audio transmission status of each forwarding server. Based on the collected status information, it executes cross-regional audio transmission scheduling strategies, thereby achieving efficient forwarding and stable transmission of audio data between nodes in different regions. Further, the system is divided into multiple regional nodes, such as the Northern Region, the Eastern Region, and the Southern Region. Each region has a corresponding forwarding server deployed to receive, process, and forward audio data generated by terminal devices within its region. Specifically, a City X node is set up within the Northern Region. This City X node establishes an audio data transmission channel with terminal devices through forwarding servers deployed in that region. The terminal devices, such as cameras and walkie-talkies, connect to the forwarding server via the network and send the collected audio data to the forwarding server for processing. Similarly, multiple regional nodes are set up within the eastern region, such as the Z city node and K city node shown in the diagram. Each node accesses and forwards audio data from regional terminal devices through its corresponding forwarding server. For example, at the K city node, terminal devices and a forwarding server are deployed. The audio data generated by the terminal devices is first sent to the forwarding server in its region, and then the forwarding server transmits the audio data according to the control strategy of the scheduling server, thus realizing cross-regional transmission of audio data between different regional nodes. Furthermore, a Y city node is set up within the southern region. This Y city node also includes terminal devices and a forwarding server. The terminal devices can be, for example, mobile terminals or headsets. The audio data they generate is sent to the forwarding server at the Y city node via the network, and audio data transmission links are established with other regional nodes through the wide area network, thus realizing cross-regional audio interaction. Further, the scheduling server maintains communication connections with the forwarding servers in each region through the network and uses the control path shown by the dotted line to schedule and control each forwarding server. Specifically, the scheduling server can obtain the operating status information and network status information of each forwarding server, and schedule cross-regional audio data transmission based on the information obtained, such as selecting appropriate forwarding server nodes for data forwarding, dynamically adjusting audio transmission links, or reallocating audio transmission tasks.

[0089] In cross-regional audio data transmission, for example, when a node in City X in the north needs to communicate with a node in City Z in the east or a node in City Y in the south, the audio data generated by the terminal device is first sent to the forwarding server of the node in City X, and then transmitted through the wide area network to the forwarding server in the target area. Subsequently, the forwarding server in the target area forwards the audio data to the corresponding terminal device, thereby realizing audio data interaction between different regions. Furthermore, during audio data transmission, the scheduling server can dynamically optimize the audio transmission path based on the load status or network quality status of nodes in different regions. For example, when the forwarding server of a certain regional node is under high load, the scheduling server can select forwarding servers of other regional nodes as backup forwarding nodes to distribute the audio transmission task, thereby ensuring the stability and real-time performance of the overall audio transmission.

[0090] In summary, by setting up a scheduling server, cross-regional forwarding servers, and terminal devices in a wide area network environment, and by implementing a unified scheduling and control mechanism to achieve audio data transmission between nodes in different regions, a control method supporting cross-regional audio communication has been constructed. This method ensures the stability of audio data transmission while achieving efficient collaboration among multiple regional nodes, thereby improving the transmission efficiency and system stability of cross-regional audio communication.

[0091] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the server includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0092] This application embodiment can divide the server into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0093] When dividing each function into modules according to its corresponding function. Figure 7This is a functional block diagram of a control device for cross-regional audio data transmission based on a wide area network (WAN) according to an embodiment of this application. The WAN-based cross-regional audio data transmission control device 700 is applied to a scheduling server of an audio transmission control system. The audio transmission control system also includes multiple forwarding servers located in multiple regions, and terminals communicatively connected to the multiple forwarding servers. The multiple forwarding servers are communicatively connected to the scheduling server. The WAN-based cross-regional audio data transmission control device 700 includes: The acquisition unit 710 is used to acquire the device information and network address information of the terminal; and the multiple first load parameters of the multiple forwarding servers; The determining unit 720 is configured to generate a virtual device identifier and a transmission identifier for transmitting audio data based on the device information and the network address information; and determine the audio transmission link information for transmitting the audio data based on the transmission identifier and the device identifier, thereby obtaining multiple audio transmission link information. Processing unit 730 is configured to construct multiple audio transmission channels for transmitting audio data by the multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information; and to obtain network status parameters of the multiple audio transmission channels; The control unit 740 is configured to determine the stability of a preset stable transmission threshold based on the network status parameters, and obtain a stability status result; if the stability status result is unstable, generate an audio encoding instruction based on the audio data, and / or determine the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters; if the stability status result is stable, control the plurality of forwarding servers to forward the audio data to the terminal through the plurality of audio transmission channels.

[0094] In one possible embodiment, the processing unit 730, in constructing multiple audio transmission channels for transmitting audio data via multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information, is specifically used for: Based on the plurality of first load parameters, determine the plurality of load weight values ​​corresponding to the plurality of forwarding servers; The first geographic attribute information corresponding to the terminal is determined based on the network address information; From the multiple audio transmission link information, forwarding server information corresponding to the first regional attribute information is filtered to obtain multiple forwarding server information; The priority parameters for regional audio scheduling in the multiple forwarding server information are determined to obtain multiple audio scheduling parameters; Based on the multiple load weight values, the multiple audio scheduling parameters are weighted and calculated to obtain multiple target audio scheduling parameters; Based on the multiple target audio scheduling parameters, the target forwarding servers corresponding to multiple audio transmission links are determined, thus obtaining multiple target forwarding servers; Based on the multiple audio transmission link information and the transmission identifier, multiple transmission port information of the multiple target forwarding servers is generated; The multiple audio transmission link information and the multiple transmission port information are mapped to the multiple target forwarding servers to obtain the multiple audio transmission channels, and non-overlapping port segments are allocated to each of the multiple audio transmission channels to achieve port resource isolation between different audio transmission channels.

[0095] In one possible embodiment, the processing unit 730, in determining the plurality of load weight values ​​corresponding to the plurality of forwarding servers based on the plurality of first load parameters, is specifically configured to: Obtain multiple rated load parameters of the multiple forwarding servers, and record the multiple rated load parameters as multiple second load parameters; Determine the difference between each of the plurality of first load parameters and each of the plurality of second load parameters to obtain a plurality of load parameter differences; The load weight values ​​corresponding to the multiple forwarding servers are calculated based on the differences in the multiple load parameters, and the multiple load weight values ​​are obtained.

[0096] In one possible embodiment, the processing unit 730, in mapping the plurality of audio transmission link information and the plurality of transmission port information to the plurality of target forwarding servers to obtain the plurality of audio transmission channels, is specifically used for: Based on the multiple transmission port information, the association information between the transmission identifier and the transmission port is determined to obtain the transmission association information; A physical link mapping identifier is generated in the plurality of target forwarding servers based on the transmission association information; Based on the information of the multiple audio transmission links, a User Datagram Protocol (UDP) port segment is determined to isolate port resources between different audio transmission links in the multiple audio transmission links, resulting in multiple port segments; Based on preset audio data forwarding rules, audio transmission channels are constructed using the multiple port segments and the physical link mapping identifiers to obtain the multiple audio transmission channels.

[0097] In one possible embodiment, the network status parameters include network bandwidth and packet loss rate, and the control unit 740, in generating audio encoding instructions based on the audio data, is specifically configured to: When the packet loss rate is greater than a preset packet loss rate threshold and the network bandwidth is less than a preset bandwidth threshold, a first audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to the preset first audio bitrate; the audio encoding instruction is determined according to the first audio bitrate adjustment instruction; When the packet loss rate is less than or equal to the packet loss rate threshold and the network bandwidth is greater than or equal to the bandwidth threshold, a second audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to a preset second audio bitrate; the audio encoding instruction is determined according to the second audio bitrate adjustment instruction; the first audio bitrate is twice the second audio bitrate.

[0098] In one possible embodiment, the network status parameters include: network strength and network latency parameters; the control unit 740, in determining the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters, is specifically used for: Obtain the first audio transmission rate and the first audio transmission interval parameters of the audio data; Based on the network strength and the network latency parameters, a network quality assessment value is determined for each of the multiple audio transmission channels, resulting in multiple network quality assessment values. Determine the audio transmission level corresponding to the multiple network quality assessment values ​​to obtain multiple audio transmission levels; Determine the target audio transmission rate and target audio transmission interval parameters corresponding to the target audio transmission level; the target audio transmission level is any one of the plurality of audio transmission levels. The adjusted transmission rate of the audio data is determined based on the target audio transmission rate and the first audio transmission rate; and the adjusted transmission interval parameter of the audio data is determined based on the first audio transmission interval parameter and the target audio transmission interval parameter. The audio adjustment transmission parameters are determined based on the adjusted transmission rate, the adjusted transmission interval parameter, and the network latency parameter.

[0099] In one possible embodiment, the control unit 740, in determining the audio adjustment transmission parameters based on the adjustment transmission rate, the adjustment transmission interval parameter, and the network latency parameter, is specifically configured to: Obtain the audio transmission parameters of the plurality of audio transmission channels; and the second load parameters of the forwarding server corresponding to the plurality of audio transmission channels; When the second load parameter is greater than or equal to the preset load parameter threshold, an audio task splitting instruction is generated according to the audio transmission parameters; Based on the audio task splitting instruction, the audio transmission allocation task of the multiple audio transmission channels is determined, resulting in multiple first audio transmission tasks; The audio transmission rate of the plurality of first audio transmission tasks is determined based on the network latency parameter, thereby obtaining a plurality of audio transmission rates; The audio adjustment transmission parameters are determined based on the plurality of audio transmission rates, the adjusted transmission rate, and the adjusted transmission interval parameters.

[0100] In one possible embodiment, the control unit is further configured to: Obtain the response information and running status information of the heartbeat frames from the multiple forwarding servers; Based on the response information and the running status information, the transmission status of each of the multiple forwarding servers is determined, and multiple transmission statuses are obtained; When the target transmission state is in failure, the forwarding server corresponding to the target transmission state is recorded as the faulty forwarding server, and the faulty forwarding server is controlled to suspend the execution of the second audio transmission task. Obtain the first forwarding server in the same area as the faulty forwarding server, and the network connectivity between the first forwarding server and the terminal; Generate audio session transport migration instructions based on network connectivity; According to the audio session transmission migration instruction, the second audio transmission task corresponding to the faulty forwarding server is cached in the first forwarding server; An audio transmission channel is established between the first forwarding server and the terminal to obtain an audio migration transmission channel; The audio data corresponding to the second audio transmission task is forwarded to the terminal through the audio migration transmission channel.

[0101] As can be seen, this application provides a control device for cross-regional audio data transmission based on a wide area network, applied to a scheduling server of an audio transmission control system. The audio transmission control system further includes multiple forwarding servers located in multiple regions, a terminal communicating with the multiple forwarding servers, and the multiple forwarding servers communicating with the scheduling server. The method includes: acquiring device information and network address information of the terminal, and multiple first load parameters of the multiple forwarding servers; generating a virtual device identifier and a transmission identifier for transmitting audio data based on the device information and network address information; determining audio transmission link information for transmitting audio data based on the transmission identifier and the virtual device identifier, obtaining multiple audio transmission link information; constructing audio transmission channels for the multiple forwarding servers to transmit audio data based on the multiple first load parameters and the multiple audio transmission link information, obtaining multiple audio transmission channels; acquiring network status parameters of the multiple audio transmission channels; performing a stability judgment on a preset stable transmission threshold based on the network status parameters, obtaining a stability status result; if the stability status result is unstable, generating an audio encoding instruction based on the audio data, and / or determining audio adjustment transmission parameters for the audio data in the multiple audio transmission channels based on the network status parameters; if the stability status result is stable, controlling the multiple forwarding servers to forward the audio data to the terminal through the multiple audio transmission channels. Thus, on the one hand, by scheduling the server to isolate concurrency and allocate tasks, and by building an audio transmission channel to transmit audio data, data conflicts can be avoided when multiple devices transmit audio. On the other hand, the forwarding server dynamically converts the bitrate of the audio signal and transmits it based on network status parameters (transmission bandwidth, packet loss rate, etc.), and controls the sending tasks of the forwarding server, thereby reducing audio stuttering issues during interaction and improving transmission stability.

[0102] This application also provides a control system for cross-regional audio data transmission based on a wide area network. This control system can execute some or all of the steps of any of the methods described in the above method embodiments.

[0103] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include a server.

[0104] It should be noted that, for the sake of simplicity, the above embodiments are all described as a series of actions. Those skilled in the art should understand that this application is not limited to the described order of actions, as some steps in the embodiments of this application can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of this application.

[0105] In the above embodiments, the descriptions of each embodiment in this application have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0106] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0107] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0108] The various devices and products described in the above embodiments include modules / units that can be software modules / units, hardware modules / units, or a combination of both. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of the embodiments of this application. It should be understood that the above descriptions are merely specific implementations of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of the embodiments of this application should be included within the scope of protection of the embodiments of this application.

Claims

1. A control method for cross-regional transmission of audio data based on a wide area network, characterized in that, A scheduling server for an audio transmission control system, the audio transmission control system further comprising multiple forwarding servers located in multiple areas, and a terminal communicatively connected to the multiple forwarding servers, the multiple forwarding servers being communicatively connected to the scheduling server, the method comprising: Obtain the device information and network address information of the terminal; and the multiple first load parameters of the multiple forwarding servers; A virtual device identifier and a transmission identifier for transmitting audio data are generated based on the device information and the network address information; Based on the transmission identifier and the virtual device identifier, the audio transmission link information for transmitting the audio data is determined, resulting in multiple audio transmission link information. Based on the multiple first load parameters and the multiple audio transmission link information, multiple audio transmission channels for transmitting audio data are constructed by the multiple forwarding servers; and network status parameters of the multiple audio channels are obtained; The stability of the preset stable transmission threshold is determined based on the network state parameters to obtain the stability state result. If the stability state result is unstable, an audio encoding instruction is generated based on the audio data, and / or, the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels are determined based on the network state parameters; If the stability status result is a stable state, control the multiple forwarding servers to forward the audio data to the terminal through the multiple audio transmission channels; The step of constructing multiple audio transmission channels for transmitting audio data by multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information includes: Based on the plurality of first load parameters, determine the plurality of load weight values ​​corresponding to the plurality of forwarding servers; The first geographic attribute information corresponding to the terminal is determined based on the network address information; From the plurality of audio transmission link information, forwarding server information corresponding to the first regional attribute information is filtered to obtain a plurality of forwarding server information; The priority parameters for regional audio scheduling in the multiple forwarding server information are determined to obtain multiple audio scheduling parameters; Based on the multiple load weight values, the multiple audio scheduling parameters are weighted and calculated to obtain multiple target audio scheduling parameters; Based on the multiple target audio scheduling parameters, the target forwarding servers corresponding to multiple audio transmission links are determined, thus obtaining multiple target forwarding servers; Based on the multiple audio transmission link information and the transmission identifier, multiple transmission port information of the multiple target forwarding servers is generated; The multiple audio transmission link information and the multiple transmission port information are mapped to the multiple target forwarding servers to obtain the multiple audio transmission channels, and non-overlapping port segments are allocated to each of the multiple audio transmission channels to achieve port resource isolation between different audio transmission channels.

2. The method as described in claim 1, characterized in that, The step of determining the multiple load weight values ​​corresponding to the multiple forwarding servers based on the multiple first load parameters includes: Obtain multiple rated load parameters of the multiple forwarding servers, and record the multiple rated load parameters as multiple second load parameters; Determine the difference between each of the plurality of first load parameters and each of the plurality of second load parameters to obtain a plurality of load parameter differences; The load weight values ​​corresponding to the multiple forwarding servers are calculated based on the differences in the multiple load parameters, and the multiple load weight values ​​are obtained.

3. The method as described in claim 1, characterized in that, The step of mapping the multiple audio transmission link information and the multiple transmission port information to the multiple target forwarding servers to obtain the multiple audio transmission channels includes: Based on the multiple transmission port information, the association information between the transmission identifier and the transmission port is determined to obtain the transmission association information; A physical link mapping identifier is generated in the plurality of target forwarding servers based on the transmission association information; Based on the information of the multiple audio transmission links, a User Datagram Protocol (UDP) port segment is determined to isolate port resources between different audio transmission links in the multiple audio transmission links, resulting in multiple port segments; Based on preset audio data forwarding rules, audio transmission channels are constructed using the multiple port segments and the physical link mapping identifiers to obtain the multiple audio transmission channels.

4. The method according to any one of claims 1-3, characterized in that, The network status parameters include network bandwidth and packet loss rate; the step of generating audio encoding instructions based on the audio data includes: When the packet loss rate is greater than a preset packet loss rate threshold and the network bandwidth is less than a preset bandwidth threshold, a first audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to the preset first audio bitrate; the audio encoding instruction is determined according to the first audio bitrate adjustment instruction; When the packet loss rate is less than or equal to the packet loss rate threshold and the network bandwidth is greater than or equal to the bandwidth threshold, a second audio bitrate adjustment instruction is sent to the plurality of forwarding servers to control the plurality of forwarding servers to send the audio data to the terminal according to a preset second audio bitrate; the audio encoding instruction is determined according to the second audio bitrate adjustment instruction; the first audio bitrate is twice the second audio bitrate.

5. The method according to any one of claims 1-3, characterized in that, The network status parameters include: network strength and network latency parameters; determining the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters includes: Obtain the first audio transmission rate and the first audio transmission interval parameters of the audio data; Based on the network strength and the network latency parameters, a network quality assessment value is determined for each of the multiple audio transmission channels, resulting in multiple network quality assessment values. Determine the audio transmission level corresponding to the multiple network quality assessment values ​​to obtain multiple audio transmission levels; Determine the target audio transmission rate and target audio transmission interval parameters corresponding to the target audio transmission level; the target audio transmission level is any one of the plurality of audio transmission levels. The adjusted transmission rate of the audio data is determined based on the target audio transmission rate and the first audio transmission rate; and the adjusted transmission interval parameter of the audio data is determined based on the first audio transmission interval parameter and the target audio transmission interval parameter. The audio adjustment transmission parameters are determined based on the adjusted transmission rate, the adjusted transmission interval parameter, and the network latency parameter.

6. The method as described in claim 5, characterized in that, Determining the audio adjustment transmission parameters based on the adjusted transmission rate, the adjusted transmission interval parameter, and the network latency parameter includes: Obtain the audio transmission parameters of the plurality of audio transmission channels; and the load parameters of the forwarding servers corresponding to the plurality of audio transmission channels; When the load parameter of the forwarding server is greater than or equal to a preset load parameter threshold, an audio task splitting instruction is generated according to the audio transmission parameters. Based on the audio task splitting instruction, the audio transmission allocation task of the multiple audio transmission channels is determined, resulting in multiple first audio transmission tasks; The audio transmission rate of the plurality of first audio transmission tasks is determined based on the network latency parameter, thereby obtaining a plurality of audio transmission rates; The audio adjustment transmission parameters are determined based on the plurality of audio transmission rates, the adjusted transmission rate, and the adjusted transmission interval parameters.

7. The method as described in claim 1, characterized in that, The method further includes: Obtain the response information and running status information of the heartbeat frames from the multiple forwarding servers; Based on the response information and the running status information, the transmission status of each of the multiple forwarding servers is determined, and multiple transmission statuses are obtained; When the target transmission state is in failure, the forwarding server corresponding to the target transmission state is recorded as the faulty forwarding server, and the faulty forwarding server is controlled to suspend the execution of the second audio transmission task. Obtain the first forwarding server in the same area as the faulty forwarding server, and the network connectivity between the first forwarding server and the terminal; Generate audio session transport migration instructions based on network connectivity; According to the audio session transmission migration instruction, the second audio transmission task corresponding to the faulty forwarding server is cached in the first forwarding server; An audio transmission channel is established between the first forwarding server and the terminal to obtain an audio migration transmission channel; The audio data corresponding to the second audio transmission task is forwarded to the terminal through the audio migration transmission channel.

8. A control device for cross-regional transmission of audio data based on a wide area network, characterized in that, A scheduling server for an audio transmission control system, the audio transmission control system further comprising multiple forwarding servers located in multiple areas, and a terminal communicatively connected to the multiple forwarding servers, the multiple forwarding servers being communicatively connected to the scheduling server, the device comprising: The acquisition unit is used to acquire the device information and network address information of the terminal; and multiple first load parameters of the multiple forwarding servers; The determining unit is configured to generate a virtual device identifier and a transmission identifier for transmitting audio data based on the device information and the network address information; and determine the audio transmission link information for transmitting the audio data based on the transmission identifier and the virtual device identifier, thereby obtaining multiple audio transmission link information. The processing unit is configured to construct multiple audio transmission channels for transmitting audio data from the multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information; and to obtain network status parameters of the multiple audio transmission channels; The control unit is configured to determine the stability of a preset stable transmission threshold based on the network status parameters, and obtain a stability status result; if the stability status result is unstable, generate an audio encoding instruction based on the audio data, and / or determine the audio adjustment transmission parameters of the audio data in the plurality of audio transmission channels based on the network status parameters; if the stability status result is stable, control the plurality of forwarding servers to forward the audio data to the terminal through the plurality of audio transmission channels; The step of constructing multiple audio transmission channels for transmitting audio data by multiple forwarding servers based on the multiple first load parameters and the multiple audio transmission link information includes: Based on the plurality of first load parameters, determine the plurality of load weight values ​​corresponding to the plurality of forwarding servers; The first geographic attribute information corresponding to the terminal is determined based on the network address information; From the multiple audio transmission link information, forwarding server information corresponding to the first regional attribute information is filtered to obtain multiple forwarding server information; The priority parameters for regional audio scheduling in the multiple forwarding server information are determined to obtain multiple audio scheduling parameters; Based on the multiple load weight values, the multiple audio scheduling parameters are weighted and calculated to obtain multiple target audio scheduling parameters; Based on the multiple target audio scheduling parameters, the target forwarding servers corresponding to multiple audio transmission links are determined, thus obtaining multiple target forwarding servers; Based on the multiple audio transmission link information and the transmission identifier, multiple transmission port information of the multiple target forwarding servers is generated; The multiple audio transmission link information and the multiple transmission port information are mapped to the multiple target forwarding servers to obtain the multiple audio transmission channels. Each audio transmission channel is assigned a non-overlapping port segment to achieve port resource isolation between different audio transmission channels.

9. A control system for cross-regional transmission of audio data based on a wide area network, characterized in that, The control system for cross-regional transmission of audio data based on a wide area network is used to execute instructions for the steps in the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video service response method and electronic equipment

    CN113259375A

  • Data transmission method, and communication apparatus and system

    WO2019196801A1