Audio transmission control system and method and audio equipment

Through the closed-loop feedback mechanism of generating TSI allocation plans through the switching scheduling layer, embedding and parsing TSIs at the connection layer, and synchronizing the terminal processing layer, the problems of insufficient scheduling capabilities on the terminal side and the lack of AVTP protocol and time slot coordination mechanism are solved, and end-to-end alignment and adaptive scheduling of audio transmission are achieved, thereby improving network adaptability and synchronization.

CN120639718AActive Publication Date: 2025-09-12SHANGHAI DAYIN INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511114245.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-12
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

The existing technology lacks scheduling capabilities on the terminal side, lacks the AVTP protocol and time slot coordination mechanism, and the terminal's dynamic response capability is insufficient, resulting in the problems of disconnection between scheduling information and terminals in audio transmission, lack of binding between AVTP encapsulation and time slot scheduling, and insufficient dynamic scheduling capabilities.

Method used

The switching scheduling layer generates a TSI allocation scheme containing the mapping relationship between GCL and TSI. The connection layer embeds and parses TSI. The terminal processing layer synchronizes based on TSI to form a closed-loop feedback mechanism, achieve cross-layer collaboration, and ensure that the audio stream strictly follows the TSN scheduling cycle transmission.

Benefits of technology

It achieves end-to-end alignment and adaptive scheduling of AVTP frames and TSN time slots, improves the system's adaptability to dynamic network environments, ensures the synchronization of audio processing and network transmission, and avoids freezes and distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639718A_ABST
    Figure CN120639718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio control, and discloses an audio transmission control system and method and audio equipment. The system comprises an exchange scheduling layer which dynamically generates a TSI allocation scheme based on a network performance index and issues the TSI allocation scheme to a connection layer; the connection layer comprises a transmitting end module and a receiving end module; the sending end module is deployed on audio sending equipment, and the TSI is embedded into an AVTP frame extension field according to a TSI allocation scheme and is sent to a network; the receiving end module is deployed in audio receiving equipment, analyzes TSI in an AVTP frame from a network, generates a network performance index, and feeds back the network performance index to the exchange scheduling layer; the terminal processing layer is deployed on audio receiving equipment, and dynamically adjusts the starting time of an audio processing algorithm according to the analyzed TSI, so that the starting time is aligned with a scheduling time window mapped by the TSI; and the exchange scheduling layer dynamically updates the TSI allocation scheme of the next period by using the fed-back network performance index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio control technology, and in particular to an audio transmission control system, method and audio equipment. Background Art

[0002] Time-Sensitive Networking (TSN), with its deterministic transmission capabilities for Ethernet, can meet the demands of professional audio and real-time control applications for data transmission under microsecond or even nanosecond time constraints. Currently, TSN's foundational standards (such as IEEE 802.1AS time synchronization and 802.1Qbv traffic scheduling) have been implemented in this field, specifically in core aspects such as switch time-aware shaping, AVTP protocol encapsulation, and terminal clock synchronization and flow control.

[0003] However, the inventors have at least discovered that the existing solutions have technical problems such as the lack of scheduling capabilities on the terminal side, the lack of AVTP protocol and time slot coordination mechanism, and the insufficient dynamic response capabilities of the terminal. Summary of the Invention

[0004] One purpose of the present application is to provide an audio transmission control system, method and audio equipment, at least to solve the technical problems of the related technology in the lack of scheduling capability on the terminal side, the lack of AVTP protocol and time slot coordination mechanism, and the insufficient dynamic response capability of the terminal.

[0005] To achieve the above objectives, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application provide an audio transmission control system, the system comprising: a switching scheduling layer, for dynamically generating a TSI allocation scheme based on network performance indicators, and issuing the TSI allocation scheme to a connection layer; the TSI allocation scheme comprises a mapping relationship between GCL and TSI, for binding the audio stream to the scheduling time window of the TSN scheduling period; a connection layer, comprising a sending end module and a receiving end module; the sending end module is deployed on an audio sending device, for embedding the TSI into an AVTP frame extension field according to the TSI allocation scheme and sending it to the network; the receiving end module is deployed on an audio receiving device, for parsing the TSI in the AVTP frame from the network, generating a network performance indicator, and feeding back the network performance indicator to the switching scheduling layer; a terminal processing layer, deployed on the audio receiving device, for dynamically adjusting the start time of the audio processing algorithm according to the parsed TSI, so that the start time is aligned with the scheduling time window mapped by the TSI; wherein the switching scheduling layer dynamically updates the TSI allocation scheme of the next cycle using the fed-back network performance indicator to form a closed-loop control.

[0006] In the second aspect, some embodiments of the present application also provide an audio transmission control method, which is applied to the system as described above, and the method includes: generating a basic TSI allocation scheme containing a mapping relationship between GCL and TSI and issuing it, so that the audio stream is bound to the TSN scheduling time window; embedding the TSI into the AVTP frame extension field according to the TSI allocation scheme, and sending it to the TSN network in the mapped scheduling time window; parsing the TSI and global timestamp in the received AVTP frame, generating network performance indicators and feedback; dynamically aligning the start time of the audio processing algorithm and the scheduling time window based on the parsed TSI; and dynamically optimizing the TSI allocation scheme for the next cycle using the fed-back network performance indicators.

[0007] In a third aspect, some embodiments of the present application further provide an audio device, comprising: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, cause the processor to perform the steps of the method described above.

[0008] Compared with the related art, in the solution provided by the embodiment of the present application, since the switching scheduling layer can generate a TSI allocation scheme containing the mapping relationship between GCL and TSI and send it down to the connection layer, the receiving end module parses the TSI in the AVTP frame and passes it to the terminal processing layer. Therefore, the terminal processing layer can directly associate the GCL time slot through TSI, so that the terminal changes from passive reception to active perception of the scheduling time window, completely breaking the isolation between the scheduling information and the terminal, and realizing the penetrating transmission of network scheduling to the terminal, thereby solving the technical problem of the disconnection between network scheduling and terminal perception in the related art; since the sending end module of the connection layer embeds the TSI into the AVTP frame extension field according to the TSI allocation scheme. Therefore, TSI, as the unique identifier of the GCL time slot, can bind the AVTP frame to a specific scheduling time window during the encapsulation stage, realize end-to-end mapping of "frame-time slot", ensure that the audio stream strictly follows the TSN scheduling cycle transmission, thereby solving the technical problem of the lack of binding between AVTP encapsulation and time slot scheduling in the related art. Since the switching scheduling layer dynamically updates the TSI allocation scheme for the next cycle based on the network performance indicators fed back by the receiving end. Therefore, the closed-loop feedback mechanism can enable the scheduling scheme to respond to changes in network status in real time, replacing manual static configuration, and realizing an adaptive cycle of network status → performance feedback → scheduling optimization, which greatly improves the system's adaptability to dynamic network environments, thereby solving the technical problem of the lack of dynamic scheduling capabilities in related technologies. Since the terminal processing layer can dynamically adjust the start time of the audio processing algorithm based on the parsed TSI so that it is strictly aligned with the scheduling time window mapped by the TSI, the network scheduling time window information can be directly passed to the terminal algorithm layer through the TSI, so that the terminal processing is transformed from local independent operation to dynamic adaptation with network scheduling, ensuring that the audio processing rhythm is synchronized with the network transmission time slot, avoiding problems such as audio freeze and distortion caused by asynchrony, thereby solving the technical problem of decoupling terminal processing and network scheduling in related technologies.

[0009] In summary, this embodiment uses a closed-loop layered architecture in which the switching scheduling layer generates a TSI allocation plan, the connection layer embeds and parses TSI, the terminal processing layer synchronizes based on TSI, and performance feedback drives dynamic optimization. This architecture uses TSI as the core carrier of cross-layer collaboration, systematically solving problems in related technologies such as network and terminal disconnection, static scheduling, and unmeasurable performance. Ultimately, it achieves end-to-end alignment of AVTP frames and TSN time slots, adaptive scheduling, deep collaboration between terminals and networks, and continuous optimization of full-link performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0011] Figure 1 This is an exemplary schematic diagram of an audio transmission control system provided in Example 1 of the present application; Figure 2 This is an exemplary schematic diagram of an audio transmission control system provided in Example 2 of the present application; Figure 3 This is an exemplary schematic diagram of an audio transmission control system provided in Example 8 of the present application; Figure 4 This is an exemplary flow chart of an audio transmission control method provided in Example 9 of the present application; Figure 5 This is an exemplary structural diagram of an audio device provided in Example 11 of the present application. DETAILED DESCRIPTION

[0012] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0013] The following terms are used in this document.

[0014] Time-Sensitive Networking, or TSN for short, is a set of protocol standards designed to provide deterministic transmission capabilities for traditional Ethernet networks, ensuring that critical data is reliably transmitted under precise time constraints at the microsecond or even nanosecond level.

[0015] Gate Control List (GCL) is a network control mechanism that manages the packet forwarding permissions of ports on network devices (such as switches and routers). It filters, allows, or denies traffic in and out of ports based on preset rules, thereby achieving network access control, traffic regulation, and security protection.

[0016] Precision Time Protocol, also known as Precision Time Protocol (PTP), is a time synchronization protocol used for high-precision frequency and phase synchronization between network nodes.

[0017] Audio Video Transport Protocol, also known as Audio Video Transport Protocol (AVTP), is a link layer transmission protocol that is mainly used to transmit audio and video data in the network. It can ensure high-quality, low-latency and synchronous transmission of audio data.

[0018] Time Slot Index (TSI) is a number or identification information used to identify a specific time slot in a time division multiplexing system or related communication architecture.

[0019] Distributed Audio Station, also known as Distributed Audio Station (DAS), is a related device or site in a distributed audio system.

[0020] Example 1 Example 1 of the present application relates to an audio transmission control system. Figure 1 As shown, the system may include: The switching scheduling layer is used to dynamically generate a TSI allocation plan based on network performance indicators and send the TSI allocation plan to the connection layer; the TSI allocation plan includes a mapping relationship between the gate control list GCL and the time slot index TSI, which is used to bind the audio stream to the scheduling time window of the TSN scheduling period; The connection layer includes a transmitter module (Talker) and a receiver module (Listener); the transmitter module is deployed on the audio transmitting device and is used to embed the TSI into the extended field of the audio transmission protocol AVTP frame according to the TSI allocation scheme and send the TSI to the network; the receiver module is deployed on the audio receiving device and is used to parse the TSI in the AVTP frame from the network, generate network performance indicators, and feed back the network performance indicators to the switching and scheduling layer; The terminal processing layer is deployed on the audio receiving device and is used to dynamically adjust the start time of the audio processing algorithm according to the parsed TSI so that the start time is aligned with the scheduling time window mapped by the TSI; The switching scheduling layer dynamically updates the TSI allocation scheme for the next cycle using the fed-back network performance indicators, thus forming a closed-loop control.

[0021] Specifically, the network performance indicators may include but are not limited to: packet loss rate, delay distribution, etc.; in this embodiment, accurate allocation of network bandwidth can be achieved by binding the audio stream to a specific time window of the TSN scheduling cycle; the switching scheduling layer can send the TSN allocation plan to the connection layer through the second interface (interface B), and continuously optimize the scheduling strategy based on the real-time network status feedback from the receiving module to form a closed-loop control.

[0022] Among them, GCL defines the opening and closing time of the transmission gate corresponding to each queue. Specifically, according to the IEEE802.1Qbv standard, the time-aware shaping mechanism TAS can achieve precise time-based control of queue transmission switches through GCL. For example, if GCL defines that during odd time intervals, the gating mechanism corresponding to queue 7 is in the open state and receives data frames; during even time intervals, the gating mechanism corresponding to queue 6 is in the open state and receives data frames, and so on. Each queue performs an opening or closing operation at a specified time based on the pre-configuration of GCL to determine whether the data packets in the queue can be transmitted, realizing the isolation of transmission of queues with different priorities and ensuring the forwarding of high-priority services based on precise time.

[0023] The TSI allocation scheme is used to match specific time slots with different data streams by comprehensively considering various factors in the network, such as traffic characteristics (constant bit rate requirements of audio streams, burst traffic characteristics of control streams and event streams, etc.), service quality requirements, etc.

[0024] Specifically, during the TSN network communication process, the TSI allocation plan issued by the switching scheduling layer can be received by the sending end module, and the TSI can be embedded in the custom extension field of the frame payload header when encapsulating the AVTP frame, so as to achieve accurate binding of audio data and network scheduling time slots, and send the AVTP frame with TSI identification to the network; the receiving end module can parse the TSI and timestamp information in the AVTP frame, calculate the end-to-end transmission delay, jitter and other network performance indicators based on the TSI and timestamp information, and feed back the indicators to the switching scheduling layer through the third interface (interface C).

[0025] Exemplarily, the connection layer may further include an AVTP protocol engine, which may perform AVTP encapsulation of audio data at the transmitting end module, perform AVTP decapsulation at the receiving end module, and output the decapsulated data to the terminal processing layer.

[0026] Specifically, the terminal processing layer is directly responsible for processing and playing audio signals. Specifically, by parsing the TSI information in the AVTP frame, the start time of the audio processing algorithm can be dynamically adjusted to ensure that the execution of the audio processing algorithm is strictly aligned with the network scheduling time window. Exemplarily, the audio processing algorithm may include but is not limited to: echo cancellation algorithm, noise reduction algorithm, etc. Through this time slot-aware processing mechanism, the terminal processing layer can achieve coordination between the audio processing algorithm and network transmission, which is conducive to improving the audio playback quality in complex network environments.

[0027] Exemplarily, the switching scheduling layer may dynamically update the TSI allocation scheme of the next period according to the timing model.

[0028] The three-layer architecture described above enables data exchange through standardized interfaces (Interface B and Interface C), forming a closed-loop control system of scheduling, execution, feedback, and optimization. This ensures that the system continuously provides high-quality audio transmission services in dynamic network environments. For example, in a smart conference system scenario, this system achieves end-to-end deterministic transmission through three layers of collaborative closed-loop control: the switching scheduling layer dynamically generates a TSI allocation plan based on real-time network performance indicators, binding the main speaker's audio stream to the highest priority scheduling time slot through the GCL and TSI mapping relationship; the transmission module of the connection layer embeds the TSI into the AVTP frame extension field and sends it to the network. The receiving module parses the TSI, calculates the network jitter value, and feeds it back to the scheduling layer; the terminal processing layer dynamically adjusts the start time of the echo cancellation algorithm of each participating terminal based on the parsed TSI, strictly aligning it with the scheduling time window mapped by the TSI to ensure lip synchronization; when the network is congested, the switching scheduling layer can immediately update the TSI allocation plan based on feedback indicators (such as migrating non-critical streams and increasing the bandwidth of the main audio stream), forming a closed-loop control.

[0029] It's important to emphasize that this embodiment pioneers the use of TSI as a cross-layer synchronization mechanism. By dynamically generating a TSI allocation scheme (including the mapping between GCL and TSI) based on network performance indicators at the switching scheduling layer, each AVTP audio frame is identified as a unique identifier within the TSN scheduling cycle's scheduling window, thus overcoming the existing barrier of terminals being unaware of network scheduling status.

[0030] That is to say, in this embodiment, the life cycle of TSI runs through all layers of the system: after the switching scheduling layer generates the TSI allocation plan, it is sent to the sending end module of the connection layer through the control channel; the sending end module embeds the TSI into the AVTP frame extension field according to the TSI allocation plan and sends it to the network; after the receiving end module parses the TSI and related time information in the AVTP frame, it feeds back the network performance indicators to the switching scheduling layer on the one hand, and passes the TSI to the terminal processing layer on the other hand; the terminal processing layer accurately synchronizes the execution of the audio processing algorithm based on TSI, and the switching scheduling layer dynamically updates the TSI allocation plan for the next cycle based on the feedback, forming a closed-loop control to achieve efficient collaboration among all layers.

[0031] Those skilled in the art will appreciate that TSN applications in the professional audio field face several limitations: First, there's a disconnect between network scheduling and terminal awareness. While TSN switches support 802.1Qbv gated scheduling, only core devices possess time awareness, leaving terminal devices unable to access scheduling information. This results in the terminal being completely unaware of network time slot status. Second, AVTP encapsulation is not tied to time slot scheduling. In this technology, AVTP frame encapsulation (compliant with the IEEE 1722 standard) only packages audio and video data and doesn't associate it with GCL time slots, making it impossible to accurately map audio streams to scheduled time windows. Third, there's a lack of dynamic scheduling capabilities. In this technology, GCL configurations are often static, requiring manual modification and terminal synchronization, making them incapable of responding to network status changes. Fourth, terminal processing is decoupled from network scheduling. Existing terminal audio processing algorithms rely solely on local configurations and lack awareness of network time slots or scheduling cadence, easily leading to asynchrony between terminal processing and network transmission.

[0032] It is not difficult to find that compared with the related art, the embodiment of the present application provides an audio transmission control system. Since the switching scheduling layer can generate a TSI allocation scheme containing the mapping relationship between GCL and TSI and send it down to the connection layer, the receiving end module parses the TSI in the AVTP frame and passes it to the terminal processing layer. Therefore, the terminal processing layer can directly associate the GCL time slot through the TSI, so that the terminal changes from passive reception to active perception of the scheduling time window, completely breaking the isolation between scheduling information and the terminal, and realizing the penetrating transmission of network scheduling to the terminal, thereby solving the technical problem of the disconnection between network scheduling and terminal perception in the related art. Because the transmitter module of the connection layer embeds the TSI into the AVTP frame extension field according to the TSI allocation scheme, the TSI, as the unique identifier of the GCL timeslot, can bind the AVTP frame to a specific scheduling time window during the encapsulation phase, achieving end-to-end "frame-timeslot" mapping and ensuring that the audio stream strictly follows the TSN scheduling cycle. This solves the technical problem of the lack of binding between AVTP encapsulation and timeslot scheduling in related technologies.

[0033] Because the switching scheduling layer dynamically updates the TSI allocation plan for the next cycle based on network performance indicators (latency, jitter, etc.) fed back by the receiving end, the closed-loop feedback mechanism enables the scheduling plan to respond to network status changes in real time, replacing manual static configuration and implementing an adaptive cycle of network status → performance feedback → scheduling optimization. This significantly improves the system's adaptability to dynamic network environments, thereby resolving the technical problem of the lack of dynamic scheduling capabilities in related technologies.

[0034] Since the terminal processing layer can dynamically adjust the start time of the audio processing algorithm based on the parsed TSI to strictly align it with the scheduling time window mapped by the TSI, the network scheduling time window information can be directly passed to the terminal algorithm layer through TSI, so that the terminal processing can be switched from local independent operation to dynamic adaptation with network scheduling, ensuring that the audio processing rhythm is synchronized with the network transmission time slot, avoiding audio freezes, distortion and other problems caused by asynchrony, thereby solving the technical problem of decoupling terminal processing and network scheduling in related technologies.

[0035] In summary, this embodiment uses a closed-loop layered architecture in which the switching scheduling layer generates a TSI allocation plan, the connection layer embeds and parses TSI, the terminal processing layer synchronizes based on TSI, and performance feedback drives dynamic optimization. This architecture uses TSI as the core carrier of cross-layer collaboration, systematically solving problems in related technologies such as network and terminal disconnection, static scheduling, and unmeasurable performance. Ultimately, it achieves end-to-end alignment of AVTP frames and TSN time slots, adaptive scheduling, deep collaboration between terminals and networks, and continuous optimization of full-link performance.

[0036] Example 2 Example 2 of the present application relates to an audio transmission control system. Example 2 is an improvement based on Example 1. The specific improvement is: in this embodiment, if Figure 2 As shown, the system may further include a time synchronization layer.

[0037] Optionally, the time synchronization layer may include: A clock source module, configured to generate a clock synchronization signal and perform redundant switching of a master clock source, thereby providing a unified global time reference for the system; A protocol execution module, configured to transmit the clock synchronization signal to a boundary clock, and synchronize the TSN switching device, audio sending device, and audio receiving device in the boundary clock network to eliminate clock deviations between devices; The output module is used to output clock synchronization signals to the switching scheduling layer and the terminal processing layer to ensure the time alignment of the TSI allocation scheme and the reference synchronization of the start time of the audio processing algorithm.

[0038] Specifically, the clock source module can be built based on the Precision Time Protocol (PTP) by deploying a grandmaster clock as the system time reference source. Optionally, the clock source module can utilize the Global Navigation Satellite System (GNSS) or a local high-stability crystal oscillator to generate a stable and reliable clock synchronization signal. In some cases, these two methods can also be used together as the reference source for the master clock. Furthermore, the clock source module supports master / backup clock redundancy switching. If the master clock source fails, it can automatically switch to the backup clock source based on actual needs, thereby providing a unified and continuous global time reference for the entire system and ensuring the continuity and stability of time synchronization.

[0039] Specifically, the protocol execution module can also be based on the IEEE 802.1AS protocol specification. For example, the protocol execution module can transmit the clock synchronization signal generated by the clock source module to a boundary clock (BoundaryClock). The boundary clock then forwards the synchronization message and compensates for link delay, passing the clock synchronization signal step by step to ordinary clocks such as TSN switching devices, audio transmitters, and audio receivers. During this process, a delay measurement mechanism can be used to accurately calculate and compensate for message transmission delay, thereby eliminating clock deviations between different devices and ensuring that the clocks of all devices remain consistent.

[0040] Specifically, the output module can output a clock synchronization signal to each system layer via the first interface (Interface A). For the switching scheduling layer, this clock synchronization signal can drive the time alignment of GCL table generation and TSI allocation schemes, ensuring that the scheduling strategy is based on a unified time standard. For the terminal processing layer, this clock synchronization signal can provide a unified benchmark for the startup time of the audio processing algorithm, keeping it synchronized with the global network time.

[0041] It is understandable that in the relevant technology, there is a lack of a dedicated time synchronization layer mechanism, and its time synchronization mostly relies on the single scheduling of the main clock source or the basic clock transmission method. There is no independent clock source module to perform redundant switching of the main clock source to ensure the stability of the global time base, nor is there a dedicated protocol execution module to accurately transmit the clock synchronization signal to the TSN switching equipment, audio sending equipment and audio receiving equipment in the network through the boundary clock. As a result, clock deviations are easily generated between devices, making it difficult to achieve high-precision time unification at the system level, and unable to meet the strict requirements of audio transmission for time synchronization.

[0042] It is not difficult to find that compared with the related art, in the embodiment of the present application, the time synchronization layer can provide a nanosecond-level precision time base for the dynamic scheduling of the switching scheduling layer and the window alignment of the terminal processing layer through the collaboration of the clock source module, the protocol execution module and the output module. Specifically, the clock source module generates a clock synchronization signal and performs redundant switching of the master clock source, which can ensure the stability and continuity of the global time base and lay the foundation for the time unification of the entire system; the protocol execution module transmits the clock synchronization signal to the boundary clock, thereby synchronizing the TSN switching equipment, audio sending equipment and audio receiving equipment in the network, which can effectively eliminate the clock deviation between devices and ensure the consistency of each device on the time scale; the output module outputs the clock synchronization signal to the switching scheduling layer and the terminal processing layer, which can enable the switching scheduling layer to achieve accurate time alignment based on a unified time base when generating and updating the TSI allocation plan, and at the same time ensure that the terminal processing layer has a reliable time base as a reference when adjusting the start time of the audio processing algorithm, thereby achieving the dynamic scheduling of the switching scheduling layer and the window alignment of the terminal processing layer in nanosecond precision, which is conducive to achieving high timeliness and synchronization of audio transmission.

[0043] Example 3 Embodiment 3 of the present application relates to an audio transmission control system. Embodiment 3 is an improvement based on Embodiment 1, and the specific improvement is that: in this embodiment, a specific implementation method of a switching scheduling layer is provided.

[0044] Optionally, the switching scheduling layer may include: Configuration management unit, used to manage network topology and configuration parameters of TSN switches; The scheduling management unit is used to generate a GCL according to the service quality requirements of the audio stream and establish a mapping relationship between the GCL and the TSI to generate a basic TSI allocation plan; The plan sending unit is used to send the basic TSI allocation plan to the connection layer as the execution basis of the first round of scheduling cycle.

[0045] Specifically, the configuration management unit can automatically identify all TSN switches in the network and their connections through the link layer discovery protocol, building a real-time topology view. It also supports centralized configuration and management of switch port parameters (such as bandwidth, priority, and forwarding rules). For example, port status and statistical information (such as traffic rate and error counts) can be queried through SNMP or NETCONF protocols. During system initialization, the configuration management unit can generate an initial network configuration template and support dynamic runtime adjustments to accommodate changes in network topology or service requirements. These configuration parameters may include, but are not limited to, gate list periods, time window lengths, and queue mapping rules.

[0046] Specifically, the scheduling management unit can generate a GCL based on the IEEE 802.1Qbv standard. Exemplarily, the scheduling management unit can employ a hybrid scheduling algorithm to generate a basic TSI allocation scheme. For example, it can allocate fixed time slots for audio streams to ensure a constant bit rate, implementing static scheduling; reserve flexible time slots for control and event streams to support bursty traffic, implementing dynamic scheduling; and implement eight levels of traffic priority based on the IEEE 802.1p priority mapping table.

[0047] Furthermore, the scheduling management unit can use network performance metrics fed back by the receiving module, combined with the precise time reference PTP, to establish a mapping between GCL and TSI, generating an initial TSI allocation plan. This mapping supports differentiated scheduling strategies based on service type (audio / control / event), dynamically adjusting time slot allocation based on traffic characteristics.

[0048] Specifically, the solution delivery unit can deliver the basic TSI allocation solution to the connection layer and protocol bridging layer via the second interface (Interface B). This delivery process can utilize a reliable transmission mechanism to ensure consistent configurations received by all terminal nodes. The TSI allocation solution includes complete GCL and TSI mapping rules. Before the first scheduling cycle, the solution delivery unit can perform version verification and consistency checks to ensure that all terminals are synchronized and enabled with the new configuration. Furthermore, the solution delivery unit can support an incremental update mechanism, transmitting only the modified components when network status changes, reducing bandwidth consumption and configuration delays.

[0049] Optionally, in some embodiments, the switching scheduling layer may further include an intelligent prediction and optimization module for dynamically optimizing the TSI allocation scheme based on network performance indicators and predicted network load. The intelligent prediction and optimization module may further include: An indicator receiving unit, configured to receive network performance indicators from the receiving end module; Load prediction unit, used to predict the network load of the next cycle; A scheme optimization unit, configured to dynamically optimize the mapping relationship between GCL and TSI in the basic TSI allocation scheme according to the network performance indicator and the predicted network load; The plan output unit is used to output the updated TSI allocation plan to the scheduling management unit.

[0050] Specifically, the indicator receiving unit can continuously receive network performance indicators from the receiving module via a third interface (Interface C). Exemplarily, these network performance indicators can include at least packet loss rate, delay distribution, and bit error rate, and may also include multi-dimensional indicators such as topology change and control message rate. The indicator receiving unit can parse data frames encapsulated using the IEEE 802.1Qch protocol and encoded in TLV (Type-Length-Value) format in real time. Furthermore, the indicator receiving unit can exemplarily include a built-in data verification mechanism to automatically identify outliers and perform smoothing to ensure the reliability of the input data.

[0051] Specifically, the load prediction unit can embed a time series prediction model such as a long short-term memory network (LSTM) or an autoregressive integrated moving average (ARIMA) model. Based on the PTP clock provided by the first interface (interface A), it performs time series analysis on historical indicator data to predict network load trends and congestion risks within a future timeframe (e.g., 20-50ms). For audio traffic, LSTM can be used to capture its periodic fluctuation characteristics, while for control messages, ARIMA can be used to model its burst characteristics. The time series prediction model can be trained using an online learning mechanism, with incremental training triggered every 1,000 new data points, supporting adaptive adjustment of model parameters. In addition, the prediction results can be output as a multidimensional vector, for example, containing bandwidth demand forecasts for each priority level of traffic, congestion probability distribution, etc.

[0052] Specifically, the solution optimization unit can adopt a multi-objective optimization framework, with minimizing end-to-end delay, maximizing bandwidth utilization, and meeting service quality requirements as optimization goals, and call the particle swarm optimization algorithm PSO to solve the optimal scheduling solution.

[0053] For example, the time slot ratios for audio, control, and event streams can be dynamically adjusted based on the predicted load. The IEEE 802.1p priority mapping table can then be adjusted based on the real-time bit error rate, while reserving a minimum bandwidth of 10% to 20% for high-priority traffic. Constraints can be introduced during the optimization process, such as a maximum number of frames transmitted per time slot (≤8 frames), a minimum time slot switching interval (≥1μs), and an upper limit on end-to-end latency (≤10ms for audio streams).

[0054] Specifically, the solution output unit can convert the optimized TSI allocation solution into a standard format and send it to the scheduling management unit via the second interface (Interface B). The output content may include: the updated GCL table (time slot opening and closing times accurate to the microsecond level), the mapping relationship between TSI and service type (including priority adjustment rules), and the solution effective timestamp synchronized with the PTP clock.

[0055] It should be noted that this embodiment may also be an improvement based on Embodiment 2.

[0056] It is not difficult to find that in the embodiment of the present application, the switching scheduling layer can realize dynamic perception of network topology, precise scheduling of differentiated traffic and reliable execution of scheduling schemes through the collaborative work of the configuration management unit, the scheduling management unit and the scheme issuing unit, thus forming a complete end-to-end scheduling control link. Specifically, the configuration management unit can solve the problem of lack of terminal node scheduling capability in traditional solutions by automatically identifying network topology and centrally managing switch parameters, so that the system can quickly adapt to network changes; the scheduling management unit generates GCL based on service quality requirements and establishes a mapping relationship between GCL and TSI, which can realize classified scheduling of different types of traffic such as audio streams, ensure that the differentiated requirements of constant bit rate and burst traffic are met, and effectively solve the problem of lack of AVTP protocol and time slot coordination mechanism; the scheme issuing unit can ensure the consistency of configuration received by all terminal nodes through reliable transmission mechanism and version verification, and provide accurate execution basis for the first round of scheduling cycle, thereby solving the problem that scheduling instructions in traditional solutions are difficult to synchronize to terminals.

[0057] Example 4 Embodiment 4 of the present application relates to an audio transmission control system. Embodiment 4 is an improvement based on Embodiment 3, and the specific improvement is that: in this embodiment, a specific implementation of the solution optimization unit is provided.

[0058] Optionally, the solution optimization unit may specifically include: a state snapshot generation subunit and a feature conversion subunit; The state snapshot generation subunit is used to periodically collect four types of data flows to generate a full network state snapshot; the four types of data flows include end-to-end performance flow, network device status flow, traffic statistics flow and application layer status flow; The feature conversion subunit is used to determine a graph structure feature vector based on the snapshot of the entire network state; wherein the graph structure feature vector is used to characterize the topology state of the entire network and serves as an input feature for dynamically optimizing the mapping relationship between GCL and TSI in the basic TSI allocation scheme.

[0059] Exemplarily, the specific composition and sources of the four types of data streams are as follows: The end-to-end performance flow can be measured by the connection layer at each receiving-end module, including the end-to-end latency, network jitter, and packet loss rate of each key audio stream. The network device status flow can be reported by the network interface cards of TSN switches and end nodes, including physical layer and link layer information such as the buffer occupancy rate, link bit error rate, and link up / down status of each port. The traffic statistics flow can be collected by TSN switches and end nodes, including statistics such as the real-time bandwidth, packet size distribution, and inter-frame arrival interval of each data stream. The application layer status flow can be fed back by the processing layer, including information such as the internal buffer status of the audio algorithm (such as whether overflow / underflow occurs) and changes in the signal-to-noise ratio (SNR). As can be seen, the coordinated collection and analysis of these four types of data flows has for the first time achieved a direct correlation between network performance and the quality of end-user experience (QoE).

[0060] Exemplarily, the state snapshot generation subunit can set a fixed collection period (e.g., 100ms / time) based on the unified PTP clock reference of the system to ensure time synchronization of each data source. Specifically, for end-to-end performance flow, the end-to-end delay, network jitter value, and packet loss rate of each key audio stream synchronously calculated and cached by the receiving module when parsing the AVTP frame can be obtained through the built-in performance measurement interface of the receiving module of the connection layer; for network device status flow, the physical layer and link layer information such as the buffer occupancy rate of each port and the link bit error rate can be collected through the management interface of the TSN switch (e.g., NETCONF protocol) and the status reporting mechanism of the terminal node network card; for traffic statistics flow, real-time bandwidth, packet size distribution and other statistical data can be collected from the traffic statistics module of the TSN switch and the packet capture interface of the terminal node; for application layer status flow, the audio algorithm buffer status, signal-to-noise ratio and other information can be obtained through the status feedback channel of the terminal processing layer.

[0061] Furthermore, the four types of data streams can be timestamp aligned (accurate to microseconds) and standardized according to a preset format to form a structured full-network status snapshot that includes the full-network device status, link performance, traffic characteristics, and application layer status. For example, at the PTP clock synchronization time of 10:00:00.000000 (microsecond level), the state snapshot generation subunit can align the end-to-end delay (500μs) and jitter (20μs) of the audio stream measured by the connection layer receiving module at 10:00:00.000002, the port buffer occupancy (30%) and link bit error rate (1e-6) reported by the TSN switch at 10:00:00.000001, the real-time bandwidth (10Mbps) and interframe interval (1ms) collected by the terminal node at 10:00:00.000003, and the buffer normal status (no overflow) and signal-to-noise ratio (45dB) reported by the terminal processing layer at 10:00:00.000000 to the timestamp of 10:00:00.000000. The data is then standardized according to a preset format: latency and jitter are expressed as integer μs, buffer occupancy is expressed as a percentage, bit error rate is expressed in scientific notation, bandwidth is uniformly expressed in Mbps, and status is marked with a Boolean value (normally true) or an enumeration value. Finally, a structured snapshot is formed, which includes: "timestamp + device identifier + link performance indicators + traffic statistics + application layer status", such as: {"timestamp":"10:00:00.000000","device_1":{"port_cache":"30%","link_ber":1e-6},"audio_stream_1":{"delay":500,"jitter":20},"traffic_stats":{"bandwidth":10,"frame_interval":1},"app_state":{"buffer_ok":true,"snr":45}}, which fully presents the status of the entire network at that moment.

[0062] Exemplarily, the feature conversion subunit can use a timestamp-aligned snapshot of the entire network state as a basis and, through feature engineering, process the raw data into feature vectors usable by the time series prediction model. This allows for time series data such as bandwidth to be time-series processed, preserving its dynamic characteristics. Furthermore, the network topology can be abstracted into a graph structure. Specifically, TSN switches and end nodes (audio transmitters and receivers) can be used as nodes, and physical links between devices as edges. Based on this, a feature vector can be extracted for each node. For example, for TSN switch nodes, this includes device state features such as port buffer occupancy, while for end nodes, it includes application-layer features such as the audio algorithm buffer status. Link performance features such as link latency, jitter, bandwidth, and bit error rate are extracted for each edge. These features are then quantized into numerical vectors through normalization (e.g., mapping buffer occupancy to a normalized value between 0 and 1). Ultimately, a graph-structured feature vector is constructed, centered around a node feature matrix, an edge feature matrix, and an adjacency matrix. This represents the entire network topology and serves as input features for dynamically optimizing the mapping between GCL and TSI in the basic TSI allocation scheme.

[0063] It should be noted that this embodiment may also be an improvement based on embodiment 1 and / or embodiment 2.

[0064] It is not difficult to find that in the embodiment of the present application, the state snapshot generation subunit periodically collects four types of data streams such as end-to-end performance flow and network device status flow and generates a full network status snapshot, which can comprehensively capture the real-time dynamics of device status, link performance, traffic characteristics and application layer status in the network, thereby providing complete and synchronized original data support for subsequent optimization, avoiding scheduling decision deviations caused by information fragmentation; and the feature conversion subunit converts the full network status snapshot into a graph structure feature vector representing the topological state of the entire network. Through structured processing, complex network status data can be efficiently parsed by the time series prediction model, which not only retains the topological association between devices and links, but also realizes data standardization and feature extraction. This enables dynamic optimization of the GCL and TSI mapping relationship. It can make decisions based on global and accurate network status characteristics, thereby achieving the purpose of improving the adaptability of the TSI allocation scheme to network changes, while also ensuring the certainty and service quality of audio stream transmission in the TSN network.

[0065] Example 5 Embodiment 5 of the present application relates to an audio transmission control system. Embodiment 5 is an improvement based on Embodiment 4, and the specific improvement is that: in this embodiment, another specific implementation of the solution optimization unit is provided.

[0066] Optionally, the solution optimization unit may further include: a state prediction subunit, a schedule generation subunit and a strategy execution subunit; The state prediction subunit is used to use the graph structure feature vector as input to predict the bandwidth occupancy, delay, network jitter value and packet loss rate of each link in the future scheduling period; the bandwidth occupancy, delay, network jitter value and packet loss rate form the predicted link performance index; The scheduling generation subunit is implemented based on the reinforcement learning framework, including: receiving the predicted link performance index output by the state prediction subunit; generating an updated GCL and TSI mapping table by the reinforcement learning agent; and calculating the reward value according to the multi-objective optimization formula; The strategy execution subunit is used to strengthen the learning agent through the reward value training, and send the updated TSI allocation plan output by the learning agent to the plan output unit.

[0067] Specifically, in this embodiment, the solution optimization unit constructs a two-stage hybrid AI model architecture through the collaboration of the state prediction subunit, the scheduling generation subunit and the strategy execution subunit to achieve accurate prediction of the network state and dynamic optimization of the scheduling strategy.

[0068] The state prediction subunit, which operates in the first stage, can be implemented using a graph neural network (GNN) model, or alternatively, a GNN variant such as a spatiotemporal graph convolutional network (ST-GCN). This allows for deep mining of the relationships between nodes (switches, terminals) and edges (links) in the network topology using graph convolution operations, taking a graph feature vector (i.e., a graph-structured snapshot of the network state) as input. For example, it can learn how congestion propagates from upstream switches to downstream nodes, or how performance fluctuations at a single node affect neighboring devices. Compared to the limitations of traditional LSTM or ARIMA models, which can only process time series data, the state prediction subunit leverages its topology-awareness to accurately capture the spatial correlations and temporal dynamics of network state. Ultimately, it outputs predicted link performance metrics for one or more future scheduling periods, including key parameters such as link bandwidth utilization, end-to-end latency for each path, network jitter, and packet loss rate.

[0069] The scheduling generation subunit acts in the second stage and can use advanced reinforcement learning algorithms such as proximal policy optimization (PPO) to build a framework. Specifically, the predicted link performance indicators output by the state prediction subunit in the first stage can be used as state input, and the process of the reinforcement learning agent generating the GCL gate list and TSI mapping table can be defined as a high-dimensional action space. The action effects are quantified through a multi-objective reward function to form a closed-loop learning mechanism of state-action-reward. For example, the multi-objective reward function can be specifically: , where R is the reward value, R 时延 The quantized value of the corresponding delay, R 抖动 The quantized value of the corresponding jitter, P 丢包The penalty value corresponding to packet loss, The weights are adjustable. Compared with traditional schemes, traditional schedulers rely on fixed optimization algorithms such as PSO. Their decision-making logic is limited by preset mathematical models and easily falls into local optimality in complex dynamic network environments, which only enables passive computational scheduling. However, the scheduling generation subunit provided in this embodiment, through the trial-and-error reward mechanism of reinforcement learning, enables the intelligent agent to discover non-intuitive optimal strategies through continuous exploration, achieving a transition from calculating scheduling solutions to generating dynamic strategies, thereby more flexibly adapting to network changes and improving the global optimality of the scheduling solution.

[0070] Specifically, the strategy execution subunit can reversely optimize the decision logic of the RL agent through the reward value, and at the same time pass the generated updated TSI allocation plan to the plan output unit, forming a closed loop of prediction-decision-feedback-iteration. It can be seen that the hybrid architecture provided by this embodiment can not only improve the prediction accuracy by leveraging the topology perception capability of GNN, but also optimize the scheduling strategy through the dynamic learning characteristics of RL, ultimately achieving adaptive adjustment of the mapping relationship between GCL and TSI in complex network environments, ensuring low latency, low jitter and high reliability of audio stream transmission.

[0071] It should be noted that this embodiment may also be an improvement based on any one or more of Embodiments 1 to 3.

[0072] It is not difficult to find that in the embodiment of the present application, the scheme optimization unit forms a closed-loop mechanism of prediction-decision-optimization through the coordinated operation of the state prediction subunit, the scheduling generation subunit and the strategy execution subunit. Among them, the state prediction subunit takes the graph structure feature vector as input, can accurately predict the future link performance indicators, and provide a basis for scheduling decisions, thereby avoiding the lag that may be caused by formulating strategies based on historical data; the scheduling generation subunit is based on the reinforcement learning framework, combines the prediction indicators to generate GCL and TSI mapping tables and calculates the reward value through a multi-objective formula. Compared with the traditional fixed algorithm, the "trial and error-reward" mechanism of reinforcement learning can explore better non-intuitive strategies and break through the local optimal limitations; the strategy execution subunit can use the reward value to continuously train the intelligent agent and output an updated plan, so that the scheduling strategy can dynamically adapt to network changes. The three work together to achieve adaptive optimization of the TSI allocation scheme, improve the certainty and service quality of audio stream transmission in the TSN network, and effectively guarantee the transmission requirements of low latency and low jitter.

[0073] Example 6 Embodiment 6 of the present application relates to an audio transmission control system. Embodiment 6 is an improvement based on Embodiment 5, and the specific improvement is that: in this embodiment, a specific implementation of the state prediction subunit is provided.

[0074] Optionally, the state prediction subunit is implemented using a graph neural network model; the graph neural network model is optimized through a closed-loop self-supervision mechanism: The predicted link performance indicator output by the graph neural network model is compared with the actual collected link performance true value after the end of the next scheduling cycle to generate an error signal; wherein, the error signal is back-propagated to the graph neural network model to drive the online adjustment of the parameters of the graph neural network model to achieve continuous iteration of the prediction capability.

[0075] Specifically, the state prediction subunit can adopt a GNN model to achieve continuous optimization through a closed-loop self-supervision mechanism: the GNN model first outputs a predicted link performance indicator, and after the next scheduling cycle ends, the system compares the predicted value with the actual collected link performance true value to generate an error signal; then, the error signal is fed back to the GNN model through a backpropagation mechanism, automatically driving the model parameters to be adjusted online, forming a complete closed loop of prediction-comparison-correction. It can be seen that the technical solution provided by this embodiment does not require manual labeling of true value labels, which can reduce operation and maintenance costs; the continuous self-calibration of the GNN model is achieved through real-time error feedback, which can enable the GNN model to dynamically adapt to complex scenarios such as network topology changes and traffic fluctuations; after long-term iteration, the prediction accuracy can be significantly improved to ensure that the output link performance indicators (bandwidth occupancy, latency, etc.) are closer to the actual network status.

[0076] For example, during one scheduling cycle, the GNN model, based on the current graph structure feature vector, predicts that link L1's bandwidth usage is 50 Mbps, latency is 200 μs, and jitter is 10 μs. At the end of the next scheduling cycle, the system actually collects bandwidth usage of L1 as 55 Mbps, latency is 210 μs, and jitter is 12 μs. This comparison generates an error signal (bandwidth error 5 Mbps, latency error 10 μs, jitter error 2 μs). This error signal is fed back to the GNN model via a backpropagation mechanism, automatically adjusting the graph convolutional layer weight parameters related to link L1 (for example, increasing sensitivity to the cache occupancy characteristics of adjacent switches). In subsequent cycles, the optimized GNN model corrects its predictions for L1's bandwidth usage to 54 Mbps and latency to 208 μs, significantly reducing the deviation from the actual values. This closed-loop prediction-comparison-correction process continuously improves prediction accuracy.

[0077] Optionally, in some embodiments, in the reinforcement learning framework of the scheduling generation subunit, the applicant designs the reward value calculation formula as a hybrid mechanism that takes into account both performance improvement and optimal maintenance. Specifically, the reward value calculation formula is: ; Where ΔD represents the end-to-end delay optimization rate, Indicates the current jitter value, Indicates the current packet loss rate. represents the minimum value, Indicates the current bandwidth utilization of key links; is a configurable weight coefficient, and R represents the reward value.

[0078] Specifically, in this embodiment, the reward calculation formula is used to evaluate the scheduling strategy under the reinforcement learning framework, quantifying network performance through four dimensions and weighted integration: Latency Optimization (α·ΔD): ΔD represents the end-to-end latency optimization rate, i.e., the decrease in current latency compared to the baseline state. This is weighted by α to encourage performance improvement. As long as latency can be reduced (the larger the ΔD), a high reward is given, driving the model to continuously optimize latency. Jitter Control ( ): is the current jitter value. The smaller the jitter, the The larger the value, the more it is weighted by β, encouraging optimality. That is, the smaller the jitter, the larger the reciprocal. Even if the network is already in a low-jitter state (such as in live broadcast scenarios), maintaining low jitter can still earn high rewards, avoiding the exploration of worse solutions due to insufficient rewards. Packet loss suppression ( ): is the current packet loss rate, is a minimum value (such as , which can avoid the denominator being invalid when the packet loss rate is 0), take the reciprocal γ weighting is also used to encourage optimality. That is, the lower the packet loss rate, the larger the reciprocal. This makes it suitable for scenarios sensitive to packet loss, such as file transfers. Bandwidth Utilization ( ): Represents the bandwidth utilization of key links, directly through Weighting encourages bandwidth efficiency and can balance the contradiction between quality and resource utilization (for example, avoiding bandwidth waste under low load).

[0079] The above formula is configurable by weight It can flexibly adapt to the priority requirements of different services for latency improvement, jitter maintenance, packet loss maintenance, and bandwidth utilization (such as increasing β for live broadcast and γ for file transfer). It can not only solve the problem of insufficient optimal state rewards caused by only using relative optimization rate, but also guide reinforcement learning strategies through multi-dimensional weighted guidance to continuously optimize network performance and stably maintain a high-quality state, thereby maximizing the reward-driven scheduling effect.

[0080] For example, in the application scenario of an audio transmission control system, take an actual scheduling optimization process as an example: the switching scheduling layer needs to dynamically generate a TSI allocation plan to ensure audio stream transmission. The sending module of the connection layer embeds the TSI into the AVTP frame according to the plan and sends it. The receiving module parses the TSI and feeds back network performance indicators (initial delay D0=90ms, jitter J0=12ms, packet loss rate L0=1.2%, critical link bandwidth utilization U0=55%). After reinforcement learning optimization, the delay optimization rate ΔD=(90-72) / 90=20% (delay reduced to 72ms) and the current jitter value are measured in the next cycle. =8ms, current packet loss rate =0.6%, current key link bandwidth utilization =68%, configure the configurable weight coefficient α=0.3 (emphasis on latency optimization), β=0.25 (focus on jitter control), γ=0.25 (emphasis on packet loss suppression), =0.2 (considering bandwidth utilization), minimum value =0.0001, and substituting them into the above reward calculation formula, we can get the contribution of delay optimization item 0.3×20%=0.06, the contribution of jitter control item 0.25×1 / 8=0.03125, the contribution of packet loss suppression item 0.25×0.6%+0.00011≈0.0416, and the contribution of bandwidth utilization item 0.2×68%=0.136. The total reward value R≈0.06+0.03125+0.0416+0.136=0.26885 The switching scheduling layer can use this reward value in combination with the reinforcement learning framework. If the reward is high, the decision logic for generating the current TSI allocation plan will be strengthened. If the network performance deteriorates (such as increased latency and increased packet loss rate resulting in low rewards), the corresponding strategy will be suppressed. This can drive continuous exploration of better TSI allocation plans. Through the closed loop of indicator feedback-reward calculation-plan optimization, the stable and efficient transmission of audio streams in the TSN network is guaranteed. The terminal processing layer can also accurately align the scheduling time window based on the optimized TSI to adjust the start time of the audio processing algorithm, thereby improving the overall quality of audio transmission.

[0081] Optionally, in some embodiments, the learning agent may be obtained by adopting an offline-online hybrid training mechanism: In the offline phase, in the digital twin simulation environment, different types of traffic loads are simulated to generate training data, and the initial weight of the learning agent is determined based on the training data; In the online stage, online learning is performed based on the initial weights.

[0082] Specifically, during the offline phase, a digital twin simulation environment can be used to simulate massive and diverse traffic load scenarios (such as concurrent audio streams and sudden traffic surges), generating training data covering various network states. This data is then used to train and determine the initial weights of the learning agent. During the online phase, the learning agent can conduct continuous online learning based on the initial weights obtained from offline training, combined with real-time interaction data generated during actual network operation (such as link performance changes and reward feedback from scheduling strategies), and continuously optimize model parameters through dynamic iteration. This approach not only leverages the simulation advantages of the digital twin environment, enabling the learning agent to master basic strategies for dealing with complex scenarios early in deployment, but also adapts to the dynamic characteristics of the real network through online learning, facilitating the continuous evolution of decision-making capabilities.

[0083] Optionally, in some embodiments, the updated TSI allocation scheme output by the scheme output unit includes two structured configuration files: GCL file, used to define the opening and closing times of the queue gates on each port of each TSN switch; TSI mapping table, used to define the association between each audio stream and the corresponding TSI identifier; The configuration file is sent to all TSN switches and terminal nodes to ensure that the new configuration is enabled synchronously across the entire network to achieve closed-loop control.

[0084] Specifically, the GCL file can clearly define the opening and closing times of the queue gates of each port of each TSN switch, ensuring transmission certainty by precisely controlling the forwarding timing of data frames; the TSI mapping table establishes an association between each audio stream and the corresponding TSI identifier, ensuring the precise binding of streams to time slice resources. After these two configuration files are obtained by the switching scheduling layer through the control channel, they are synchronously sent to all TSN switches and terminal nodes to achieve atomic configuration updates and network-wide activation. In this way, the generated optimization strategy can be quickly converted into a specific configuration that can be executed by the device, forming a complete closed loop from strategy output to network-wide execution, which is conducive to ensuring the dynamic adaptation of the TSN network to the needs of audio stream transmission.

[0085] It should be noted that this embodiment may also be an improvement based on any one or more of Embodiments 1 to 4.

[0086] It is not difficult to find that in the embodiment of the present application, the state prediction subunit can adopt a graph neural network model and be optimized through a closed-loop self-supervision mechanism. Since the predicted link performance indicators output by the model can be compared with the actual true value in the next scheduling cycle to generate an error signal, and the error automatically adjusts the model parameters through back propagation, this can enable the model to achieve dynamic iteration of parameters without manual labeling, that is, dynamic scenarios such as network topology changes and traffic fluctuations will be converted into correction signals in real time, promoting the prediction logic to continuously adapt to the actual environment; the continuous improvement of prediction accuracy can provide the scheduling generation subunit with a link performance input that is closer to reality, avoiding the failure of the scheduling strategy due to prediction deviation; furthermore, the obtained optimization scheme can more effectively match the audio stream transmission requirements, ensure the low-latency and low-jitter transmission of audio data in the TSN network, and form a virtuous cycle of accurate prediction → scheduling optimization → reliable transmission.

[0087] Example 7 Embodiment 7 of the present application relates to an audio transmission control system. Embodiment 7 is an improvement based on Embodiment 1, and the specific improvement is that: in this embodiment, a specific implementation of a receiving end module is provided.

[0088] Specifically, the receiving end module may include: Frame parsing module, used to parse TSI and global timestamp in AVTP frames; A performance calculation module, configured to calculate end-to-end transmission delay, network jitter value, and packet loss rate based on frame arrival time and the global timestamp; The indicator feedback module is used to feed back the network performance indicator to the switching scheduling layer.

[0089] Specifically, the frame parsing module can locate the custom extension field in the AVTP frame payload header, extract the TSI field therefrom, and parse the PTP timestamp, which records the global time of the sending module.

[0090] Specifically, the performance calculation module calculates the transmission delay of a single frame by comparing the frame arrival time (obtained from the local PTP clock) with the global timestamp carried within the frame. Using a sliding window algorithm, it continuously tracks delay variations, generating a delay distribution histogram, statistical extreme values, and network jitter in real time. For packet loss rate calculation, a sequence number comparison mechanism supports continuous sequence number recovery across network shards. For example, a statistical report containing 1000 frames of performance data can be generated every 20ms.

[0091] The extension and read-write mechanism of the TSI field can be implemented in the following way: 16 bits of space are reserved for storing TSI in the custom extension field of the payload header of the AVTP frame (following the IEEE1722 standard). After the sending module receives the mapping relationship between GCL and TSI issued by the switching scheduling layer, it writes the corresponding TSI into the extension field when encapsulating the AVTP frame; the receiving module locates and reads this field during the decapsulation process, and associates it with the PTP timestamp in the frame to provide a time slot alignment basis for the terminal processing layer. The entire process does not change the main structure of the AVTP protocol, and only realizes the binding of TSI and frame data and end-to-end transmission through custom extension.

[0092] Specifically, the indicator feedback module can feed back the performance calculation results to the switching scheduling layer through the third interface (interface C). The feedback data is encapsulated in TLV format and includes: end-to-end delay distribution (frame proportion in 10 intervals), jitter statistical extreme values ​​(maximum, minimum, average), packet loss rate (subdivided by service type), and bit error rate (based on AVTP frame CRC check results). A double buffer mechanism can be used to decouple data collection and transmission, supporting lossless collection of indicators under burst traffic. The feedback cycle is configurable (default 100ms) and can be adaptively adjusted based on network load. When the network is congested, the feedback priority of key indicators (such as audio stream delay) is automatically increased to ensure the timeliness of scheduling optimization.

[0093] For example, when an AVTP frame containing audio data arrives at the connection layer of the receiving end, the frame parsing module can first parse it and extract the TSI identifier (such as "TSI_001") and global timestamp (such as 10:00:00.123456) carried in the frame; then, the performance calculation module can compare the global timestamp with the local time when the frame actually arrives at the receiving end (such as 10:00:00.123656), calculate the end-to-end transmission delay to be 200μs, and calculate the jitter to be 15μs based on the arrival time deviation of 10 consecutive frames. At the same time, by detecting missing sequence number frames (such as the failure of the fifth frame to arrive), it is determined that the packet loss rate is 10%; the indicator feedback module can package network performance indicators including "packet loss rate 10%, delay 200μs, jitter 15μs" and feedback them to the switching scheduling layer through the internal interface.

[0094] Optionally, in some embodiments, the terminal processing layer may be specifically used to: Based on the parsed TSI and the corresponding global timestamp, the start time of the audio processing algorithm is dynamically adjusted to align the start time with the scheduling time window mapped by the TSI; Among them, the dynamic adjustment includes at least one of the following operations: triggering the audio processing algorithm according to the scheduling time window boundary mapped by TSI; adjusting the audio buffer depth according to the end-to-end transmission delay; dynamically switching the audio redundancy error correction strategy according to the packet loss rate; and dynamically modifying the anti-jitter parameters of the audio processing algorithm according to the network jitter value.

[0095] Specifically, the terminal processing layer receives AVTP frames with TSI from the connection layer through the fourth interface (interface D), parses the TSI and PTP timestamp, and compares them with the local PTP clock to calculate the frame arrival time deviation and time slot position. Based on this, the parameters of the audio processing algorithm (such as echo cancellation, dynamic EQ, noise reduction, etc.) are dynamically adjusted to achieve deep coordination with network scheduling. Specifically, the algorithm can be triggered according to the scheduling time window boundary mapped by TSI to ensure that echo cancellation, noise reduction and other processing are started during the period when data transmission is most stable, thereby improving processing efficiency; based on the end-to-end transmission delay distribution, the audio buffer depth can be optimized through the delay prediction model to avoid jamming or data overflow caused by delay fluctuations; the audio redundancy error correction strategy can be dynamically switched according to changes in packet loss rate to enhance data fault tolerance; the anti-jitter parameters such as filter coefficients can be adaptively adjusted according to the network jitter value to improve the algorithm's adaptability to network instability. Ultimately, by periodically receiving TSI and timestamp updates, a closed-loop dynamic adjustment mechanism for algorithm parameters is formed, ensuring that the terminal processing process is strictly aligned with the network scheduling time slot, which is conducive to comprehensively ensuring the coordination and stability of audio processing and network transmission.

[0096] It should be noted that this embodiment may also be an improvement based on any one or more embodiments from Embodiment 2 to Embodiment 6.

[0097] Understandably, in related technologies, audio transmission systems lack a systematic performance feedback loop. They lack a real-time collection mechanism for key metrics like latency and jitter during end-to-end transmission, and lack an effective path for systematically feeding these metrics back to the scheduling core. This prevents network administrators from obtaining real-time quality data on audio streams during transmission, making it difficult to accurately assess transmission stability and dynamically adjust scheduling strategies based on actual performance. This makes it difficult for the system to quickly identify transmission quality issues when faced with changes in network load or emergencies, and consequently, fails to guarantee the certainty and reliability of audio transmission.

[0098] It is not difficult to find that in the embodiment of the present application, while the receiving end module parses TSI, it combines the frame arrival time and the global timestamp to generate network performance indicators and feeds them back to the switching scheduling layer. Therefore, TSI can become the anchor point for performance testing, so that indicators such as delay and jitter can be accurately associated with specific scheduling time slots, and a complete link of transmission → analysis → timing → feedback is constructed, providing data support for scheduling optimization, realizing the quantification and traceability of transmission quality, and solving the technical problem of the blank end-to-end performance testing mechanism in related technologies.

[0099] Example 8 Example 8 of the present application relates to an audio transmission control system. Example 8 is an improvement based on Example 1. The specific improvement is that: in this embodiment, the system also includes a protocol bridge layer, which is deployed on the audio sending device and the audio receiving device. In this way, the audio transmission control system is specifically as follows: Figure 3 The five-tier distributed architecture shown.

[0100] In this embodiment, the audio transmission control system is logically divided into five functional layers, with clear interfaces and data flows to achieve efficient collaboration between each layer. This separation of the data plane and the control plane ensures high cohesion and low coupling between modules, facilitating expansion and maintenance. The control plane, primarily composed of the switching and scheduling layer, is responsible for global network awareness, intelligent scheduling, and centralized management. The data plane encompasses the time synchronization layer, the connection layer, the terminal processing layer, and the protocol bridging layer, each responsible for specific tasks such as establishing time synchronization, data encapsulation and decapsulation, transmission, and audio processing. This layered collaboration ensures stable system operation.

[0101] Optionally, the protocol bridging layer may include: Protocol receiving module, used to receive data frames of non-TSN bus protocols; A format conversion module, configured to parse and re-encapsulate the data frame into a TSN-compatible AVTP frame; A time window binding module is used to bind the re-encapsulated audio stream to the TSN scheduling time window according to the TSI allocation plan issued by the switching scheduling layer; The protocol restoration module is used to restore the encapsulated data from the TSN backbone network into the original non-TSN protocol format data frame.

[0102] Specifically, the non-TSN bus protocols may include but are not limited to: CAN, A2B, FlexRay, etc. The protocol receiving module can ensure that all types of non-TSN data can smoothly enter the protocol bridge layer for subsequent processing by adapting to the physical interfaces and communication protocols of different heterogeneous buses. Specifically, the format conversion module can parse the non-TSN data frames transmitted by the protocol receiving module, extract the valid data and control information therein, and then re-encapsulate this information into TSN-compatible AVTP frames according to the specifications of the TSN network. In this way, the format conversion of non-TSN protocol data to TSN protocol data is realized, so that data that could originally only be transmitted on heterogeneous buses can be transparently transmitted on the TSN backbone network.

[0103] Specifically, to ensure that the converted audio stream complies with the TSN network's scheduling requirements, the time window binding module binds the re-encapsulated AVTP frames to the corresponding TSN scheduling time windows based on the TSI allocation plan issued by the switching scheduling layer. This binding ensures that these data frames can be transmitted within the TSN network's predetermined time slots, aligning with the network's overall scheduling mechanism and ensuring real-time and deterministic data transmission.

[0104] Specifically, the protocol restoration module can parse the AVTP frames from the TSN backbone network, extract the original non-TSN protocol data information, and then restore it to the corresponding non-TSN protocol format data frame, and then send it through the corresponding heterogeneous bus interface. In this way, the reverse conversion of TSN protocol data to non-TSN protocol data can be achieved, ensuring that data can be transmitted normally in the heterogeneous bus network, further improving the interconnection and interoperability between the TSN network and heterogeneous bus networks.

[0105] It should be noted that this embodiment may also be an improvement based on any one or more embodiments from Embodiment 2 to Embodiment 7.

[0106] It is not difficult to find that in the embodiment of the present application, the protocol bridging layer, through the collaborative work of the above-mentioned modules, with the help of the TSN backbone network interface and multiple heterogeneous bus interfaces, can realize the two-way transmission and protocol conversion of data between the TSN network and the heterogeneous bus network, thereby ensuring the system's compatibility with multiple vehicle-mounted or industrial buses.

[0107] It is worth mentioning that all modules involved in this embodiment are logical modules. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovation of this application, this embodiment does not include units that are not closely related to solving the technical problem proposed by this application. However, this does not mean that other units do not exist in this embodiment.

[0108] Example 9 Embodiment 9 of the present application relates to a method for controlling audio transmission, which can be applied to the system described in any one or more of Embodiments 1 to 8.

[0109] like Figure 4 As shown, the method may include at least the following steps: Step S101: Generate and distribute a basic TSI allocation scheme containing a mapping relationship between GCL and TSI, so that the audio stream is bound to the TSN scheduling time window; Step S102, embedding the TSI into the AVTP frame extension field according to the TSI allocation scheme, and sending it to the TSN network in the mapped scheduling time window; Step S103: parse the TSI and global timestamp in the received AVTP frame, generate network performance indicators and provide feedback; Step S104, dynamically aligning the start time of the audio processing algorithm with the scheduling time window based on the parsed TSI; Step S105, dynamically optimizing the TSI allocation scheme for the next cycle using the fed-back network performance indicators; The above steps are continuously executed in a loop to achieve adaptive scheduling.

[0110] The following is a detailed description of each of the above steps.

[0111] For step S101, illustratively, after the system is started, the time synchronization layer can establish a unified time reference through the PTP protocol to complete time synchronization; then, the switching scheduling layer can generate a basic TSI allocation scheme containing a mapping relationship between GCL and TSI based on preset configurations such as the quality of service requirements of the audio stream, and send it to all sending modules and receiving modules in the network through the second interface (interface B), so that the audio stream is bound to the corresponding TSN scheduling time window.

[0112] Regarding step S102, the transmitting module can, for example, receive and encapsulate the audio data according to the issued basic TSI allocation scheme, embedding the corresponding TSI into a custom extension field of the AVTP frame. After encapsulation, the connection layer can send the AVTP frame over the TSN network according to the precise time slot specified by the GCL, ensuring that the audio stream is transmitted within the predetermined time window.

[0113] Regarding step S103, illustratively, after receiving AVTP frames from the TSN network, the receiving module parses them, extracts the TSI and global timestamp, and transmits them to the terminal processing layer. Simultaneously, the connection layer calculates network performance indicators such as end-to-end transmission delay and network jitter based on the frame arrival time and global timestamp, calculates the audio packet loss rate, and feeds these network performance indicators back to the switching and scheduling layer via a third interface (Interface C).

[0114] For step S104, exemplarily, after the terminal processing layer receives the TSI and the corresponding global timestamp from the connection layer, it can dynamically adjust the start time of the audio processing algorithm (such as echo cancellation, noise reduction, etc.) based on this information, so that the start time is accurately aligned with the scheduling time window mapped by the TSI, thereby achieving deep synchronization between the terminal audio processing algorithm and the network status.

[0115] For step S105, illustratively, the intelligent prediction and optimization module of the switching scheduling layer can collect and analyze the network performance indicators fed back from the receiving end module, learn and predict in combination with historical data, and determine whether the current scheduling scheme can meet the quality requirements of audio transmission; based on the analysis and prediction results, dynamically recalculate and generate the optimized TSI allocation scheme for the next cycle (including the updated GCL and TSI mapping relationship), and then send it to each node through the second interface (interface B).

[0116] The above steps are executed in a continuous loop, forming an adaptive, intelligent closed-loop control system. By continuously generating solutions, transmitting data, analyzing feedback, and optimizing solutions, the system can respond to changes in network status in real time, continuously ensuring high-quality audio transmission and achieving adaptive scheduling.

[0117] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0118] It is not difficult to find that this embodiment is a method embodiment corresponding to Example 1, and this embodiment can be implemented in conjunction with Example 1. The relevant technical details mentioned in Example 1 are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in Example 1.

[0119] Example 10 Embodiment 10 of the present application relates to a specific application example of an audio transmission control system and method.

[0120] Exemplarily, the method and system can be applied to an in-vehicle surround audio system. The deployment and operation process of the in-vehicle surround audio system is as follows: the system deploys PTP master / backup clocks and local crystal oscillators in the main switch and auxiliary switch on the vehicle to form a time synchronization layer, and the switching scheduling layer sends the initial GCL / TSI to the in-vehicle AVTP transmitter module through the management port; in the connection layer, the transmitter module embeds TSI in the PCM audio stream, the receiver module is responsible for aligning the playback, and the distributed audio station DAS nodes of each channel in the cockpit act as the terminal processing layer to dynamically run noise reduction and three-dimensional surround algorithms without the need for a protocol bridge layer. After the system is started, the PTP synchronization is initialized to establish a unified time base, and then the GCL and TSI mapping is sent through the management port. Subsequently, the AVTP transmitter module encapsulates and sends the audio frame, the receiver module unpacks and calibrates the playback and feeds back the timing data, and the switching scheduling layer updates the GCL / TSI based on the feedback, forming an iterative and cyclical adaptive scheduling process.

[0121] Exemplarily, the method and system can be applied to industrial plant broadcasting systems. The deployment architecture of the industrial plant broadcasting system is as follows: the time synchronization layer deploys multi-source PTP in the plant core switch and the ring network switch to ensure the uniformity of the time base of the entire network; the switching scheduling layer triggers the dynamic update of GCL based on real-time traffic monitoring indicators to achieve flexible scheduling; the transmitting end module of the connection layer completes the AVTP broadcast frame encapsulation at the ring network node; each speaker node of the terminal processing layer performs audio processing such as echo cancellation and reverberation compensation; the protocol bridging layer can integrate the industrial monitoring system according to demand to achieve docking with the industrial bus. Its operating process is basically the same as that of the in-vehicle surround audio system. It only adapts to the docking requirements of the industrial bus in the protocol interaction link, and ensures the real-time and stability of the broadcast audio through the collaboration of each layer.

[0122] Exemplarily, the method and system can be applied to multi-room multimedia in smart buildings. The deployment architecture of multi-room multimedia in smart buildings is as follows: the time synchronization layer realizes the clock synchronization between the building center room and the switches on each floor, providing a unified time base for the system; the switching scheduling layer allocates GCL and TSI by floor to achieve hierarchical and precise scheduling; the connection layer connects with the smart speaker through the floor AVTP transmitter module / receiver module, responsible for the encapsulation and parsing of audio and video streams; the smart speaker node of the terminal processing layer performs dynamic EQ and scene recognition to optimize the audio output effect; the protocol bridge layer integrates the building management system to achieve linkage with other systems in the building. Its operation process focuses on multi-room synchronous playback and interactive control, and ensures the time consistency of audio playback in different rooms through the collaboration of various layers, and supports users' interactive operations on multimedia content to enhance the multimedia experience of smart buildings.

[0123] This application has at least the following beneficial effects: First, the end-to-end synchronization accuracy is significantly improved. The system controls the synchronization error between AVTP frames and GCL time slots to sub-microsecond level through deep binding of TSI and PTP timestamps; Second, the modular expansion capability is enhanced. Each layer adopts an independent design and interacts through standardized interfaces, supporting single-layer technology upgrades (such as upgrading the time synchronization layer from PTPv2 to IEEE802.1AS-Rev) without affecting other layers; Third, the complexity of system maintenance is reduced. Standardized interface definitions enable remote configuration management and fault diagnosis, and combined with the performance optimization suggestions of the intelligent prediction and optimization modules, the operation and maintenance efficiency and functional scalability are greatly improved.

[0124] Example 11 Embodiment 11 of the present application relates to an audio device. The audio device may be a digital computer in various forms, such as a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, etc. The audio device may also be a mobile device in various forms, such as a personal digital assistant, a cellular phone, a smartphone, a wearable device, and other similar computing devices.

[0125] The audio device includes: one or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor performs the steps of the method provided in any one or more of the above embodiments. Figure 5 An exemplary structural diagram of the audio device is disclosed. The audio device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are interconnected using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the audio device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple audio devices can be connected, with each device providing some of the necessary operations. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0126] The audio device may further include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means, with the bus connection being used as an example in the figure.

[0127] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the audio device. Input devices such as a touch screen, keypad, mouse, trackpad, touchpad, pointer, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0128] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.

[0129] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.

[0131] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.

Claims

1. An audio transmission control system, characterized in that: The system comprises: The switching scheduling layer is used to dynamically generate a TSI allocation plan based on network performance indicators and send the TSI allocation plan to the connection layer; the TSI allocation plan includes a mapping relationship between GCL and TSI, which is used to bind the audio stream to the scheduling time window of the TSN scheduling cycle; The connection layer includes a transmitting module and a receiving module; the transmitting module is deployed on the audio transmitting device and is used to embed the TSI into the AVTP frame extension field according to the TSI allocation scheme and send the TSI to the network; the receiving module is deployed on the audio receiving device and is used to parse the TSI in the AVTP frame from the network, generate a network performance indicator, and feed back the network performance indicator to the switching scheduling layer; The terminal processing layer is deployed on the audio receiving device and is used to dynamically adjust the start time of the audio processing algorithm according to the parsed TSI so that the start time is aligned with the scheduling time window mapped by the TSI; The switching scheduling layer dynamically updates the TSI allocation scheme for the next cycle using the fed-back network performance indicators, thus forming a closed-loop control.

2. The system according to claim 1, wherein: The system further includes: a time synchronization layer; the time synchronization layer includes: A clock source module, configured to generate a clock synchronization signal and perform redundant switching of a master clock source, thereby providing a unified global time reference for the system; A protocol execution module, configured to transmit the clock synchronization signal to a boundary clock, and synchronize the TSN switching device, audio sending device, and audio receiving device in the boundary clock network to eliminate clock deviations between devices; The output module is used to output clock synchronization signals to the switching scheduling layer and the terminal processing layer to ensure the time alignment of the TSI allocation scheme and the reference synchronization of the start time of the audio processing algorithm.

3. The system according to claim 1, wherein: The switching scheduling layer includes: Configuration management unit, used to manage network topology and configuration parameters of TSN switches; The scheduling management unit is used to generate a GCL according to the service quality requirements of the audio stream and establish a mapping relationship between the GCL and the TSI to generate a basic TSI allocation plan; The plan sending unit is used to send the basic TSI allocation plan to the connection layer as the execution basis of the first round of scheduling cycle.

4. The system according to claim 3, characterized in that The switching scheduling layer also includes an intelligent prediction and optimization module, which includes: An indicator receiving unit, configured to receive network performance indicators from the receiving end module; Load prediction unit, used to predict the network load of the next cycle; A scheme optimization unit, configured to dynamically optimize the mapping relationship between GCL and TSI in the basic TSI allocation scheme according to the network performance indicator and the predicted network load; The plan output unit is used to output the updated TSI allocation plan to the scheduling management unit.

5. The system according to claim 4, characterized in that The solution optimization unit specifically includes: a state snapshot generation subunit and a feature conversion subunit; The state snapshot generation subunit is used to periodically collect four types of data flows to generate a full network state snapshot; the four types of data flows include end-to-end performance flow, network device status flow, traffic statistics flow and application layer status flow; The feature conversion subunit is used to determine a graph structure feature vector based on the snapshot of the entire network state; wherein the graph structure feature vector is used to characterize the topology state of the entire network and serves as an input feature for dynamically optimizing the mapping relationship between GCL and TSI in the basic TSI allocation scheme.

6. The system according to claim 5, characterized in that The solution optimization unit also includes: a state prediction subunit, a scheduling generation subunit and a strategy execution subunit; The state prediction subunit is used to use the graph structure feature vector as input to predict the bandwidth occupancy, delay, network jitter value and packet loss rate of each link in the future scheduling period; the bandwidth occupancy, delay, network jitter value and packet loss rate form the predicted link performance index; The scheduling generation subunit is implemented based on the reinforcement learning framework, including: receiving the predicted link performance index output by the state prediction subunit; generating an updated GCL and TSI mapping table by the reinforcement learning agent; and calculating the reward value according to the multi-objective optimization formula; The strategy execution subunit is used to strengthen the learning agent through the reward value training, and send the updated TSI allocation plan output by the learning agent to the plan output unit.

7. The system according to claim 6, characterized in that The state prediction subunit is specifically implemented using a graph neural network model; the graph neural network model is optimized through a closed-loop self-supervision mechanism: The predicted link performance indicator output by the graph neural network model is compared with the actual collected link performance true value after the end of the next scheduling cycle to generate an error signal; wherein, the error signal is back-propagated to the graph neural network model to drive the online adjustment of the parameters of the graph neural network model to achieve continuous iteration of the prediction capability.

8. The system according to claim 6, wherein: In the reinforcement learning framework of the scheduling generation subunit, the calculation formula of the reward value is: ; Where ΔD represents the end-to-end delay optimization rate, Indicates the current jitter value, Indicates the current packet loss rate. represents the minimum value, Indicates the current bandwidth utilization of key links; is a configurable weight coefficient, and R represents the reward value.

9. The system according to claim 6, wherein: The learning agent is specifically obtained by adopting an offline-online hybrid training mechanism: In the offline phase, in the digital twin simulation environment, different types of traffic loads are simulated to generate training data, and the initial weight of the learning agent is determined based on the training data; In the online stage, online learning is performed based on the initial weights.

10. The system according to any one of claims 4 to 9, characterized in that The updated TSI allocation plan output by the plan output unit includes two structured configuration files: GCL file, used to define the opening and closing times of the queue gates on each port of each TSN switch; TSI mapping table, used to define the association between each audio stream and the corresponding TSI identifier; The configuration file is sent to all TSN switches and terminal nodes to ensure that the new configuration is enabled synchronously across the entire network to achieve closed-loop control.

11. The system according to claim 1, wherein: The receiving end module includes: Frame parsing module, used to parse TSI and global timestamp in AVTP frames; A performance calculation module, configured to calculate end-to-end transmission delay, network jitter value, and packet loss rate based on frame arrival time and the global timestamp; The indicator feedback module is used to feed back the network performance indicator to the switching scheduling layer.

12. The system according to claim 11, wherein: The terminal processing layer is specifically used for: Based on the parsed TSI and the corresponding global timestamp, the start time of the audio processing algorithm is dynamically adjusted to align the start time with the scheduling time window mapped by the TSI; Among them, the dynamic adjustment includes at least one of the following operations: triggering the audio processing algorithm according to the scheduling time window boundary mapped by TSI; adjusting the audio buffer depth according to the end-to-end transmission delay; dynamically switching the audio redundancy error correction strategy according to the packet loss rate; and dynamically modifying the anti-jitter parameters of the audio processing algorithm according to the network jitter value.

13. The system according to claim 1, wherein: The system also includes a protocol bridging layer, which is deployed on the audio sending device and the audio receiving device, including: Protocol receiving module, used to receive data frames of non-TSN bus protocols; A format conversion module, configured to parse and re-encapsulate the data frame into a TSN-compatible AVTP frame; The time window binding module is used to bind the re-encapsulated audio stream to the TSN scheduling time window according to the TSI allocation plan issued by the switching scheduling layer; The protocol restoration module is used to restore the encapsulated data from the TSN backbone network into the original non-TSN protocol format data frame.

14. An audio transmission control method, characterized in that: The method is applied to the system according to any one of claims 1 to 13, and the method at least comprises: Generate and distribute a basic TSI allocation plan containing the mapping between GCL and TSI, so that the audio stream is bound to the TSN scheduling time window; Embed the TSI into the AVTP frame extension field according to the TSI allocation scheme and send it to the TSN network in the mapped scheduling time window; Parse the TSI and global timestamp in the received AVTP frame, generate network performance indicators and provide feedback; Dynamically align the start time and scheduling time window of the audio processing algorithm based on the parsed TSI; The TSI allocation scheme for the next cycle is dynamically optimized using the feedback network performance indicators.

15. An audio device, characterized in that The audio device includes: one or more processors; and A memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method of claim 14.

Citation Information

Patent Citations

  • Centralized network configuration entity and time sensitive network control system comprising same

    CN114830611A

  • Process layer network communication method, device and system based on TSN, and TSN controller

    CN115484140A

  • Real-time interactive digital human system supporting high concurrency and implementation method thereof

    CN120179081A

  • Industrial real-time data transmission guarantee method and system based on TSN

    CN120343043A

  • Joint traffic routing and scheduling method for removing non-deterministic interrupt for TSN network used in industrial IoT

    US20250112873A1