Audio transmission control system, method and audio device
By generating a TSI allocation scheme through the exchange scheduling layer, embedding and parsing TSIs in the connection layer, and dynamically adjusting the audio processing algorithm in the terminal processing layer, the problems of insufficient scheduling capability on the terminal side and lack of AVTP protocol and time slot coordination mechanism are solved. End-to-end alignment and adaptive scheduling of audio transmission are achieved, improving the dynamic response capability and synchronization of the system.
Patent Information
- Application Number
- CN202511114245.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing technologies lack terminal-side scheduling capabilities, AVTP protocol and time slot coordination mechanisms, and terminal dynamic response capabilities, resulting in problems such as disconnect between network scheduling and terminal perception, lack of binding between AVTP encapsulation and time slot scheduling, and lack of dynamic scheduling capabilities in audio transmission.
The TSI allocation scheme, which includes the mapping relationship between GCL and TSI, is generated by the exchange scheduling layer. The AVTP frame extension field is embedded through the connection layer. After the receiving module parses the TSI, it is passed to the terminal processing layer to form a closed-loop control, realizing the precise binding and dynamic adjustment of audio stream and network scheduling time window.
It achieves end-to-end alignment of audio streams with TSN scheduling cycles, improves the system's adaptability to dynamic network environments, ensures synchronization between audio processing and network transmission, avoids stuttering and distortion, and realizes deep collaboration between terminals and networks and end-to-end performance optimization.
Smart Images

Figure CN120639718B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio control, and in particular to an audio transmission control system and method and an audio device. BACKGROUND
[0002] Time Sensitive Network (TSN) can meet the demand of professional audio and real-time control field for data transmission under microsecond or even nanosecond time constraint due to its deterministic transmission capability for Ethernet. The basic standard set of TSN (such as IEEE 802.1AS time synchronization, 802.1Qbv traffic scheduling, etc.) has been applied in this field, which is specifically reflected in the following aspects: switch time-aware shaping, AVTP protocol encapsulation, and terminal clock synchronization and traffic control.
[0003] However, the inventors have at least found that the existing scheme has the technical problems of lack of terminal scheduling capability, lack of AVTP protocol and time slot coordination mechanism, and insufficient terminal dynamic response capability. SUMMARY
[0004] An object of the present application is to provide an audio transmission control system and method and an audio device, at least to solve the technical problems of lack of terminal scheduling capability, lack of AVTP protocol and time slot coordination mechanism, and insufficient terminal dynamic response capability in the related art.
[0005] To achieve the above object, some embodiments of the present application provide the following aspects:
[0006] In a first aspect, some embodiments of the present application provide an audio transmission control system, which comprises: a switch scheduling layer, configured to dynamically generate a TSI allocation scheme based on a network performance index and to issue the TSI allocation scheme to a link layer; the TSI allocation scheme comprises a mapping relationship between a GCL and a TSI, and is used to bind an audio stream to a scheduling time window of a TSN scheduling cycle; the link layer comprises a sending end module and a receiving end module; the sending end module is deployed on an audio sending device and is configured to embed a TSI into an AVTP frame extension field according to the TSI allocation scheme and to send the TSI to a network; the receiving end module is deployed on an audio receiving device and is configured to parse the TSI in an AVTP frame from the network, to generate a network performance index, and to feed back the network performance index to the switch scheduling layer; and a terminal processing layer is deployed on the audio receiving device and is configured to dynamically adjust a start time of an audio processing algorithm according to the parsed TSI, so that the start time is aligned with a scheduling time window mapped by the TSI; wherein the switch scheduling layer dynamically updates a TSI allocation scheme of a next cycle by using the fed back network performance index, to form a closed loop control.
[0007] In a second aspect, some embodiments of the present application further provide an audio transmission control method, which is applied to the system as described above, and the method comprises: generating a basic TSI allocation scheme containing a GCL and TSI mapping relationship and issuing, so that an audio stream is bound to a TSN scheduling time window; embedding a TSI in an AVTP frame extension field according to the TSI allocation scheme, and sending to a TSN network in a mapped scheduling time window; parsing the TSI and global timestamp in the received AVTP frame, generating a network performance index and feeding back; dynamically aligning the starting time of an audio processing algorithm with the scheduling time window based on the parsed TSI; and dynamically optimizing the TSI allocation scheme of the next period by using the feedback network performance index.
[0008] In a third aspect, some embodiments of the present application further provide an audio device, which comprises: one or more processors; and a memory storing computer program instructions which, when executed, cause the processor to perform the steps of the method as described above.
[0009] Compared with the related art, in the scheme provided by the embodiments of the present application, the exchange scheduling layer can generate a TSI allocation scheme containing the mapping relationship between the GCL and the TSI and issue it to the adaptation layer, and the terminal processing layer can directly associate the GCL time slot through the TSI after the receiving end module parses the TSI in the AVTP frame. Therefore, the terminal can be switched from passive reception to active sensing of the scheduling time window, completely breaking the isolation between the scheduling information and the terminal, realizing the penetration of network scheduling to the terminal, and thus the technical problem of disconnection between network scheduling and terminal sensing in the related art can be solved. Since the sending end module of the adaptation layer embeds the TSI in the AVTP frame extension field according to the TSI allocation scheme, the TSI, as the unique identifier of the GCL time slot, can bind the AVTP frame to a specific scheduling time window at the encapsulation stage, realize end-to-end mapping of "frame-time slot", and ensure that the audio stream strictly follows the TSN scheduling period for transmission, so that the technical problem of lack of binding between AVTP encapsulation and time slot scheduling in the related art can be solved. Since the exchange scheduling layer dynamically updates the TSI allocation scheme for the next period based on the network performance indicators fed back by the receiving end, the closed-loop feedback mechanism can enable the scheduling scheme to respond to network state changes in real time, replace manual static configuration, realize the adaptive cycle of network state->performance feedback->scheduling optimization, and greatly improve the adaptability of the system to dynamic network environments, so that the technical problem of lack of dynamic scheduling capability in the related art can be solved. Since the terminal processing layer can dynamically adjust the start time of the audio processing algorithm according to the parsed TSI, so that it is strictly aligned with the scheduling time window mapped by the TSI, the network scheduling time window information can be directly transmitted to the terminal algorithm layer through the TSI, the terminal processing is switched from local independent running to dynamic adaptation with network scheduling, the audio processing rhythm is synchronized with the network transmission time slot, and problems such as audio lag and distortion caused by asynchronization are avoided, so that the technical problem of decoupling between terminal processing and network scheduling in the related art can be solved.
[0010] In summary, the embodiments of the present application generate a TSI allocation scheme through an exchange scheduling layer, embed and parse the TSI through an adaptation layer, synchronize based on the TSI through a terminal processing layer, and drive dynamic optimization through a closed-loop hierarchical architecture of performance feedback. The TSI is used as the core carrier for cross-layer collaboration, and the problems of disconnection between the network and the terminal, static scheduling, and unmeasurable performance in the related art are systematically solved, and finally the end-to-end alignment of the AVTP frame and the TSN time slot, adaptive scheduling, deep collaboration between the terminal and the network, and continuous optimization of the performance of the entire link are realized. BRIEF DESCRIPTION OF DRAWINGS
[0011] One or more embodiments are illustrated by way of example with reference to the drawings, which are not limiting of the embodiments and are merely meant to illustrate the general structure of the embodiments. Like reference numerals refer to like elements throughout the drawings. The drawings are not to scale.
[0012] Figure 1 An exemplary schematic diagram of an audio transmission control system provided for Embodiment 1 of the present application;
[0013] Figure 2 An exemplary schematic diagram of an audio transmission control system provided for Embodiment 2 of the present application;
[0014] Figure 3 An exemplary schematic diagram of an audio transmission control system provided for Embodiment 8 of the present application;
[0015] Figure 4 An exemplary flow chart of an audio transmission control method provided for Embodiment 9 of the present application;
[0016] Figure 5 An exemplary structural diagram of an audio device provided for Embodiment 11 of the present application. DETAILED DESCRIPTION
[0017] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0018] The following terms are used herein.
[0019] Time-Sensitive Networking (TSN) is a set of protocol standards designed to give traditional Ethernet deterministic transmission capabilities, ensuring reliable transmission of critical data within microseconds or even nanoseconds of precise time constraints.
[0020] Gate Control List (GCL) is a network control mechanism mainly used to manage the data packet forwarding rights of network device (such as switch, router) ports. Through pre-set rules, it filters, allows or denies the traffic in and out of the port, thereby achieving network access control, traffic monitoring and security protection.
[0021] Precision Time Protocol (PTP) is a time synchronization protocol for high-precision frequency synchronization and phase synchronization between network nodes.
[0022] Audio Video Transport Protocol, abbreviated as AVTP, is a link layer transmission protocol, which is mainly used for transmitting audio and video data in the network, and can ensure high-quality, low-delay and synchronous transmission of audio data.
[0023] Time Slot Index, abbreviated as TSI, refers to the number or identification information used to identify a specific time slot in a time division multiplexing system or related communication architecture.
[0024] Distributed Audio Station, abbreviated as DAS, is a related device or station in a distributed audio system.
[0025] Embodiment 1
[0026] Embodiment 1 of the present application relates to an audio transmission control system. As shown in Figure 1 , the system can include:
[0027] The exchange scheduling layer is used to dynamically generate a TSI allocation scheme based on network performance indicators and issue the TSI allocation scheme to the adaptation layer; the TSI allocation scheme includes a mapping relationship between a GCL and a TSI, and is used to bind an audio stream to a scheduling time window of a TSN scheduling cycle;
[0028] The adaptation layer includes a talker and a listener; the talker is deployed in an audio sending device and is used to embed a TSI in an AVTP frame extension field according to the TSI allocation scheme and send it to the network; the listener is deployed in an audio receiving device and is used to parse the TSI in the AVTP frame from the network, generate network performance indicators, and feed back the network performance indicators to the exchange scheduling layer;
[0029] The terminal processing layer is deployed in the audio receiving device and is used to dynamically adjust the start time of an audio processing algorithm according to the parsed TSI, so that the start time is aligned with the scheduling time window mapped by the TSI.
[0030] The exchange scheduling layer dynamically updates the TSI allocation scheme of the next cycle using the feedback network performance indicators to form a closed-loop control.
[0031] Specifically, the network performance indicators can include, but are not limited to, packet loss rate, delay distribution, etc.; in this embodiment, the network bandwidth can be accurately allocated by binding the audio stream to a specific time window of the TSN scheduling cycle; the exchange scheduling layer can distribute the TSN allocation scheme to the adaptation layer through the second interface (interface B), and continuously optimize the scheduling strategy according to the real-time network state feedback of the receiving end module, forming a closed-loop control.
[0032] Wherein, GCL defines the opening and closing time of the transmission gate corresponding to each queue. Specifically, according to the IEEE802.1Qbv standard, the time-aware shaping mechanism TAS can realize accurate time-based control of the queue transmission switch through GCL. For example, if GCL defines that during the odd time interval, the queue 7 corresponding gate mechanism is in the open state to receive data frames; during the even interval, the queue 6 corresponding gate mechanism is in the open state to receive data frames, and so on. Each queue executes the open or close operation at the specified time according to the pre-configuration of GCL to determine whether the data packets in the queue can be transmitted, realizing the isolation of transmission of different priority queues and guaranteeing the accurate time-based forwarding of high priority services.
[0033] Wherein, the TSI allocation scheme is used to correspond specific time slots to different data streams by comprehensively considering various factors in the network, such as traffic characteristics (constant bit rate requirements of audio stream, burst traffic characteristics of control stream and event stream, etc.), quality of service requirements, etc.
[0034] Specifically, in the TSN network communication process, the TSI allocation scheme distributed by the exchange scheduling layer can be received by the sending end module, and the TSI can be embedded in the custom extension field of the frame payload header when encapsulating the AVTP frame, so as to realize the accurate binding of audio data and network scheduling time slot, and send the AVTP frame with TSI identifier to the network; the receiving end module can analyze the TSI and timestamp information in the AVTP frame, calculate the end-to-end transmission delay, jitter and other network performance indicators according to the TSI and timestamp information, and feed back the indicators to the exchange scheduling layer through the third interface (interface C).
[0035] Exemplarily, the adaptation layer can also include an AVTP protocol engine, which can perform AVTP encapsulation of audio data at the sending end module; perform AVTP decapsulation at the receiving end module, and output the decapsulated data to the terminal processing layer.
[0036] Specifically, the terminal processing layer is configured to be directly responsible for processing and playing of the audio signal. Specifically, the TSI information in the AVTP frame can be parsed to dynamically adjust the starting time of the audio processing algorithm, so as to ensure that the execution of the audio processing algorithm is strictly aligned with the network scheduling time window. Exemplarily, the audio processing algorithm can include, but is not limited to, an echo cancellation algorithm, a noise reduction algorithm, etc. Through the time slot sensing processing mechanism, the terminal processing layer can realize the cooperation of the audio processing algorithm and the network transmission, and is beneficial to improving the audio playing quality in a complex network environment.
[0037] Exemplarily, the exchange scheduling layer can dynamically update the TSI allocation scheme of the next period according to the timing model.
[0038] The three-layer architecture realizes data interaction through standardized interfaces (interfaces B and C), forms a closed-loop control system of scheduling-execution-feedback-optimization, and can ensure that the system continuously provides high-quality audio transmission services in a dynamic network environment. For example, in the intelligent conference system scenario, the system realizes end-to-end deterministic transmission through the three-layer cooperative closed-loop control: the exchange scheduling layer dynamically generates a TSI allocation scheme based on real-time network performance indicators, binds the audio stream of the main speaker to the highest priority scheduling time slot through the GCL and TSI mapping relationship; the sending end module of the interface layer embeds the TSI in the AVTP frame extension field and sends it to the network, and the receiving end module calculates the network jitter value after parsing the TSI and feeds back to the scheduling layer; the terminal processing layer dynamically adjusts the starting time of the echo cancellation algorithm of each participant terminal according to the parsed TSI, so that it is strictly aligned with the scheduling time window of the TSI mapping, and ensures the lip synchronization; when the network is congested, the exchange scheduling layer can immediately update the TSI allocation scheme based on the feedback indicators (such as migrating non-critical streams and increasing the bandwidth of the main audio stream), forming a closed-loop control.
[0039] It should be emphasized that the embodiment innovatively constructs the TSI as a cross-layer synchronization mechanism. Through the exchange scheduling layer, the TSI allocation scheme (including the mapping relationship between GCL and TSI) is dynamically generated based on network performance indicators, and the scheduling time window of the TSN scheduling period to which each AVTP audio frame belongs is uniquely identified in the form of a unique identifier, thereby breaking the barrier that the terminal is not aware of the network scheduling state in the prior art.
[0040] That is, in this embodiment, the life cycle of TSI runs through all layers of the system: after the exchange scheduling layer generates the TSI allocation scheme, it is issued to the sending module of the adaptation layer through the control channel; the sending module embeds the TSI into the AVTP frame extension field according to the TSI allocation scheme and sends it to the network; after the receiving module parses the TSI and related time information in the AVTP frame, on the one hand, it feeds back the network performance indicators to the exchange scheduling layer, and on the other hand, it transfers the TSI to the terminal processing layer; the terminal processing layer is based on the execution of the TSI precise synchronization audio processing algorithm, and the exchange scheduling layer dynamically updates the TSI allocation scheme for the next period according to the feedback, forming a closed loop control, and realizing efficient cooperation of all layers.
[0041] Those skilled in the art can understand that the related art has many limitations in the application of TSN in the professional audio field: first, the network scheduling is disconnected from the terminal perception. In the related art, although the TSN switch supports 802.1Qbv gating scheduling, only the core device has time perception capability, and the terminal device cannot obtain scheduling information, resulting in that the terminal is completely unaware of the network time slot state; second, AVTP encapsulation and time slot scheduling are not bound. In the related art, the encapsulation of the AVTP frame (complying with the IEEE1722 standard) only completes the packaging of audio and video data, and is not associated with the GCL time slot, so the audio stream cannot be accurately mapped to the scheduling time window; third, dynamic scheduling capability is missing. In the related art, the GCL is mostly statically configured, and modification needs to be manually operated and needs to be synchronized with the terminal, and cannot respond to network state changes; fourth, terminal processing and network scheduling are decoupled. The audio processing algorithm of the existing terminal only relies on local configuration for running, and cannot perceive the network time slot or the scheduling rhythm, which easily leads to the desynchronization of terminal processing and network transmission.
[0042] It is not difficult to find that, compared with the related art, the audio transmission control system provided by the embodiments of the present application can generate a TSI allocation scheme containing the mapping relationship between GCL and TSI and issue it to the adaptation layer, and the receiving module transfers the TSI to the terminal processing layer after parsing the TSI in the AVTP frame. Therefore, the terminal processing layer can directly associate the GCL time slot through the TSI, so that the terminal changes from passive reception to active perception of the scheduling time window, completely breaking the isolation between scheduling information and the terminal, realizing the penetrating transmission of network scheduling to the terminal, and thus the technical problem of disconnection between network scheduling and terminal perception in the related art can be solved.
[0043] Since the sending module of the adaptation layer embeds the TSI into the AVTP frame extension field according to the TSI allocation scheme, the TSI, as the unique identifier of the GCL time slot, can bind the AVTP frame to a specific scheduling time window at the encapsulation stage, realize end-to-end mapping of "frame-time slot", and ensure that the audio stream strictly follows the TSN scheduling period for transmission, so that the technical problem of lack of binding between AVTP encapsulation and time slot scheduling in the related art can be solved.
[0044] Since the exchange scheduling layer dynamically updates the TSI allocation scheme of the next period based on the network performance indicators (delay, jitter, etc.) fed back by the receiving end, the closed-loop feedback mechanism can enable the scheduling scheme to respond to network state changes in real time, replace manual static configuration, realize an adaptive cycle of network state→performance feedback→scheduling optimization, and greatly improve the adaptability of the system to dynamic network environments, thereby solving the technical problem of lack of dynamic scheduling capability in the related art.
[0045] Since the terminal processing layer can dynamically adjust the start time of the audio processing algorithm according to the parsed TSI, so that it is strictly aligned with the scheduling time window mapped by the TSI, the network scheduling time window information can be directly transmitted to the terminal algorithm layer through the TSI, the terminal processing is changed from local independent running to dynamic adaptation with network scheduling, the audio processing rhythm is ensured to be synchronized with the network transmission time slot, and problems such as audio stall and distortion caused by asynchronization are avoided, thereby solving the technical problem of decoupling of terminal processing and network scheduling in the related art.
[0046] In summary, the embodiment generates a TSI allocation scheme through an exchange scheduling layer, embeds and parses the TSI through a link layer, synchronizes the terminal processing layer based on the TSI, and drives dynamic optimization through performance feedback, thereby realizing an adaptive scheduling, deep collaboration between the terminal and the network, and continuous optimization of the performance of the whole link.
[0047] Embodiment 2
[0048] Embodiment 2 of the present application relates to an audio transmission control system. Embodiment 2 is an improvement based on embodiment 1, and the specific improvement is that in the present embodiment, as shown in Figure 2 The system can further include a time synchronization layer.
[0049] Optionally, the time synchronization layer can include:
[0050] A clock source module for generating a clock synchronization signal and performing redundant switching of a master clock source to provide a unified global time reference for the system;
[0051] A protocol execution module for transmitting the clock synchronization signal to a boundary clock, synchronizing TSN exchange devices, audio sending devices and audio receiving devices in the network through the boundary clock, so as to eliminate the clock deviation between devices;
[0052] The output module is configured to output a clock synchronization signal to the exchange scheduling layer and the terminal processing layer, so as to ensure time alignment of the TSI allocation scheme and reference synchronization of the starting time of the audio processing algorithm.
[0053] Specifically, the clock source module can be constructed based on a Precision Time Protocol (PTP), and a grand master clock can be deployed as a system time reference source. Alternatively, the clock source module can generate a stable and reliable clock synchronization signal by using a global navigation satellite system (GNSS) or a local high-stability crystal oscillator. In some examples, the two methods can be used in combination as reference sources of the grand master clock. Meanwhile, the clock source module supports primary / backup clock redundancy switching, and can automatically switch to a backup clock source when the primary clock source fails, so as to provide a unified and continuous global time reference for the entire system and ensure the continuity and stability of time synchronization.
[0054] Specifically, the protocol execution module can also be based on the IEEE 802.1AS protocol specification. The protocol execution module can transmit the clock synchronization signal generated by the clock source module to a boundary clock (BoundaryClock), which then forwards the synchronization message and compensates for the link delay, and transmits the clock synchronization signal to the TSN exchange device, the audio sending device, and the audio receiving device, and other ordinary clocks. In the above process, a delay measurement mechanism can be used to accurately calculate the message transmission delay and compensate for it, so as to eliminate the clock deviation between different devices and keep the clocks of all devices consistent.
[0055] Specifically, the output module can output the clock synchronization signal to each level of the system through a first interface (interface A). For the exchange scheduling layer, the clock synchronization signal can drive the generation of the GCL table and the time alignment of the TSI allocation scheme, so as to ensure that the scheduling strategy is based on a unified time standard. For the terminal processing layer, the clock synchronization signal can provide a unified reference for the starting time of the audio processing algorithm, so that the audio processing algorithm is synchronized with the global network time.
[0056] It can be understood that in the related art, there is a lack of a dedicated time synchronization layer mechanism, and time synchronization is mostly dependent on single scheduling of a grand master clock source or a basic clock transmission mode. There is no independent clock source module to perform redundancy switching of the grand master clock source to ensure the stability of the global time reference, and there is no dedicated protocol execution module to accurately transmit the clock synchronization signal to the TSN exchange device, the audio sending device, and the audio receiving device in the network through the boundary clock, resulting in clock deviation between devices, difficulty in achieving high-precision time unification at the system level, and inability to meet the stringent requirements of audio transmission on time synchronization.
[0057] It can be found that, compared with the related art, in the embodiment of the application, the time synchronization layer can provide a time reference with nanosecond-level precision for the dynamic scheduling of the exchange scheduling layer and the window alignment of the terminal processing layer through the cooperation of the clock source module, the protocol execution module and the output module. Specifically, the clock source module generates a clock synchronization signal and performs redundant switching of the main clock source, which can ensure the stability and continuity of the global time reference and lay the foundation for the time unification of the entire system; the protocol execution module transmits the clock synchronization signal to the boundary clock, and then synchronizes the TSN exchange device, the audio sending device and the audio receiving device in the network, which can effectively eliminate the clock deviation between devices and ensure the consistency of each device in the time scale; the output module outputs the clock synchronization signal to the exchange scheduling layer and the terminal processing layer, which can enable the exchange scheduling layer to achieve accurate time alignment based on the unified time reference when generating and updating the TSI allocation scheme, and ensure that the terminal processing layer has a reliable time reference as a reference when adjusting the starting time of the audio processing algorithm, so as to realize the cooperation of the dynamic scheduling of the exchange scheduling layer and the window alignment of the terminal processing layer in nanosecond-level precision, which is conducive to realizing the high timeliness and synchronization of audio transmission.
[0058] Embodiment 3
[0059] Embodiment 3 of the application relates to an audio transmission control system. Embodiment 3 is an improvement based on embodiment 1, and the specific improvement is that in the embodiment, a specific implementation mode of the exchange scheduling layer is provided.
[0060] Optionally, the exchange scheduling layer can include:
[0061] A configuration management unit is configured to manage the network topology and the configuration parameters of the TSN switch.
[0062] A scheduling management unit is configured to generate a GCL according to the quality of service requirement of the audio stream, and establish a mapping relationship between the GCL and the TSI to generate a basic TSI allocation scheme.
[0063] A scheme issuing unit is configured to issue the basic TSI allocation scheme to the adaptation layer as the basis for execution of the first scheduling period.
[0064] In particular, the configuration management unit can automatically identify all TSN switches in the network and their connection relationships through a link layer discovery protocol, and construct a real-time topology view. At the same time, it supports centralized configuration and management of switch port parameters (such as bandwidth, priority, forwarding rules). For example, the port status and statistical information (such as traffic rate, error count) can be queried through the SNMP or NETCONF protocol. During the system initialization phase, the configuration management unit can generate an initial network configuration template and support dynamic adjustment at runtime to adapt to changes in network topology or changes in business requirements. The configuration parameters can include but are not limited to: gate list period, time window length, and queue mapping rules.
[0065] In particular, the scheduling management unit can generate a GCL according to the IEEE 802.1Qbv standard. For example, the scheduling management unit can use a hybrid scheduling algorithm to generate a basic TSI allocation scheme. For example: allocate fixed time slots for audio streams to guarantee constant bit rate and achieve static scheduling; reserve flexible time slots for control streams and event streams to support burst traffic and achieve dynamic scheduling; and implement 8-level traffic priority based on the IEEE 802.1p priority mapping table.
[0066] Further, the scheduling management unit can establish a mapping relationship between the GCL and the TSI by receiving network performance indicators fed back by the receiving end module, combined with the precise time reference PTP, to generate an initial TSI allocation scheme. This mapping relationship can support differentiated scheduling policies based on service type (audio / control / event), and can dynamically adjust time slot allocation according to traffic characteristics.
[0067] In particular, the scheme issuing unit can issue the basic TSI allocation scheme to the adaptation layer and the protocol bridge layer through the second interface (interface B). The issuing process can use a reliable transmission mechanism to ensure the consistency of the configuration received by all terminal nodes. The TSI allocation scheme includes complete GCL and TSI mapping rules. Before the first scheduling period is executed, the scheme issuing unit can perform version checking and consistency checking to ensure that all terminals synchronize to use the new configuration. At the same time, the scheme issuing unit can support an incremental update mechanism to only transmit the changed part when the network state changes, reducing bandwidth consumption and configuration delay.
[0068] Optionally, in some embodiments, the switching scheduling layer can further include an intelligent prediction and optimization module for dynamically optimizing the TSI allocation scheme based on network performance indicators and predicted network load. The intelligent prediction and optimization module can further include:
[0069] An indicator receiving unit for receiving network performance indicators from the receiving end module;
[0070] a load prediction unit configured to predict a network load of a next period;
[0071] a scheme optimization unit configured to dynamically optimize a GCL and TSI mapping relationship in the basic TSI allocation scheme according to the network performance indicators and the predicted network load;
[0072] a scheme output unit configured to output the updated TSI allocation scheme to the scheduling management unit.
[0073] Specifically, the indicator receiving unit can continuously receive network performance indicators from the receiving end module through a third interface (interface C). Exemplarily, the network performance indicators can at least include a packet loss rate, a delay distribution, and a bit error rate, and can also include multi-dimensional indicators such as a topology change and a control message rate. The indicator receiving unit can analyze data frames in real time, which are encapsulated using an IEEE 802.1Qch protocol and encoded in a TLV (Type-Length-Value) format. In addition, the indicator receiving unit can internally embed a data verification mechanism, so as to automatically identify abnormal values and perform smoothing processing, thereby ensuring the reliability of the input data.
[0074] Specifically, the load prediction unit can internally embed a time series prediction model such as a long short-term memory network (LSTM) or an autoregressive integrated moving average model (ARIMA), perform time series analysis on historical indicator data based on a PTP clock provided by a first interface (interface A), and predict network load trends and congestion risks in a future time (for example, 20-50 ms). For audio traffic, an LSTM can be used to capture its periodic fluctuation characteristics, and for control messages, an ARIMA model can be used to capture its burst characteristics. The training of the time series prediction model can use an online learning mechanism, and an incremental training is triggered once every 1000 new data points, supporting adaptive adjustment of model parameters. In addition, the prediction results can be output in the form of a multi-dimensional vector, for example, can include bandwidth demand prediction values of each priority traffic, congestion probability distribution, and the like.
[0075] Specifically, the scheme optimization unit can use a multi-objective optimization framework, with the optimization objectives being to minimize end-to-end delay, maximize bandwidth utilization, and meet quality of service requirement constraints, and can call a particle swarm optimization algorithm (PSO) to solve the optimal scheduling scheme.
[0076] For example, the time slot proportion of audio / control / event streams can be first adjusted dynamically according to the predicted load, and then an IEEE 802.1p priority mapping table can be adjusted based on real-time bit error rate, while reserving 10%-20% of the guaranteed bandwidth for high-priority traffic. In the optimization process, constraint conditions can be introduced, such as a single time slot maximum transmission frame number constraint (≤8 frames), a time slot switching minimum interval constraint (≥1 μs), an end-to-end delay upper limit constraint (audio stream ≤10 ms), and the like.
[0077] Specifically, the scheme output unit can specifically convert the optimized TSI allocation scheme into a standard format and issue it to the scheduling management unit through a second interface (interface B). The output content can include: an updated GCL table (accurate to the time slot opening and closing time of the microsecond level), a mapping relationship between TSI and service type (including priority adjustment rules), and a scheme effective timestamp for PTP clock synchronization.
[0078] It should be noted that the present embodiment can also be an improvement based on embodiment 2.
[0079] As can be seen, in the present embodiment, the exchange scheduling layer can realize dynamic perception of network topology, accurate scheduling of differentiated traffic, and reliable execution of scheduling schemes through the collaborative work of the configuration management unit, the scheduling management unit, and the scheme issuing unit, forming a complete end-to-end scheduling control link. Specifically, the configuration management unit can solve the problem of lack of terminal node scheduling capability in traditional schemes by automatically identifying network topology and centrally managing switch parameters, enabling the system to quickly adapt to network changes; the scheduling management unit can realize the classification and scheduling of different types of traffic such as audio streams by generating GCL based on quality of service requirements and establishing a mapping relationship between GCL and TSI, ensuring that the differentiated needs of constant bit rate and burst traffic are met, effectively solving the problem of lack of AVTP protocol and time slot coordination mechanism; the scheme issuing unit can ensure the consistency of configurations received by all terminal nodes through reliable transmission mechanism and version checking, providing accurate execution basis for the first round of scheduling period, thereby solving the problem of difficulty in synchronizing scheduling instructions to terminals in traditional schemes.
[0080] Embodiment 4
[0081] Embodiment 4 of the present application relates to an audio transmission control system. Embodiment 4 is an improvement based on embodiment 3, and the specific improvement is that in the present embodiment, a specific implementation of the scheme optimization unit is provided.
[0082] Optionally, the scheme optimization unit can specifically include: a state snapshot generation subunit and a feature conversion subunit;
[0083] The state snapshot generation subunit is configured to periodically collect four types of data streams to generate a full-network state snapshot; the four types of data streams include end-to-end performance streams, network device state streams, traffic statistics streams, and application layer state streams;
[0084] The feature conversion subunit is configured to determine a graph structure feature vector according to the full-network state snapshot; wherein the graph structure feature vector is used to represent the full-network topology state as an input feature for dynamically optimizing the GCL and TSI mapping relationship in the basic TSI allocation scheme.
[0085] For example, the specific composition and source of the four types of data flow are as follows:
[0086] The end-to-end performance flow can be measured by the adaptation layer at each receiving end module, including the end-to-end delay, network jitter value and packet loss rate of each key audio stream; the network device state flow can be reported by the TSN switch and the network card of the terminal node, including the buffer occupancy rate, link error rate, link on-off state and other physical layer and link layer information of each port; the traffic statistics flow can be collected by the TSN switch and the terminal node, including real-time bandwidth, packet size distribution, inter-frame arrival interval and other statistical data of each data flow; the application layer state flow can be fed back by the processing layer, including the internal buffer state of the audio algorithm (such as whether overflow / underflow occurs), signal-to-noise ratio (SNR) change and other information. It can be seen that the cooperative collection and analysis of the four types of data flow directly relates the network performance to the quality of experience (QoE) of the end user for the first time.
[0087] For example, the state snapshot generation sub-unit can set a fixed collection period (such as 100 ms / time) based on the unified PTP clock reference of the system to ensure the time synchronization of each data source. Specifically, for the end-to-end performance flow, the end-to-end delay, network jitter value and packet loss rate of each key audio stream calculated and cached by the receiving end module when analyzing the AVTP frame can be obtained through the performance measurement interface built-in the receiving end module of the adaptation layer; for the network device state flow, the buffer occupancy rate, link error rate and other physical layer and link layer information of each port can be collected through the management interface (such as the NETCONF protocol) of the TSN switch and the state reporting mechanism of the terminal node network card; for the traffic statistics flow, real-time bandwidth, packet size distribution and other statistical data can be collected from the traffic statistics module of the TSN switch and the packet capture interface of the terminal node; for the application layer state flow, the buffer state of the audio algorithm and the signal-to-noise ratio can be obtained through the state feedback channel of the terminal processing layer.
[0088] Further, the four types of data streams can be timestamped (accurate to the microsecond level) and standardized according to a preset format, forming a structured network-wide state snapshot containing device status, link performance, traffic characteristics, and application layer status. For example, at the PTP clock synchronization time of 10:00:00.000000 (microsecond level), the state snapshot generation subunit can align the audio stream end-to-end delay (500 μs) and jitter (20 μs) measured by the adapter layer receiving end module at 10:00:00.000002, the port cache occupancy rate (30%) and link error rate (1e-6) reported by the TSN switch at 10:00:00.000001, the real-time bandwidth (10 Mbps) and interframe interval (1 ms) collected by the terminal node at 10:00:00.000003, and the buffer normal state (no overflow) and signal-to-noise ratio (45 dB) feedback by the terminal processing layer at 10:00:00.000000 to the timestamp of 10:00:00.000000. Then, standardized according to the preset format: delay and jitter retain integer μs values, cache occupancy rate is expressed in percentage, error rate uses scientific notation, bandwidth unit is unified to Mbps, state is marked with Boolean value (normal is true) or enumeration value, and finally integrated to form a structured snapshot containing "timestamp + device identification + link performance indicators + traffic statistics + application layer state", such as: { "timestamp": "10:00:00.000000", "device_1": { "port_cache": "30%", "link_ber": 1e-6}, "audio_stream_1": { "delay": 500, "jitter": 20}, "traffic_stats": { "bandwidth": 10, "frame_interval": 1}, "app_state": { "buffer_ok": true, "snr": 45}}. The complete network state at that moment is presented.
[0089] Exemplarily, the feature conversion subunit can process the original data into a feature vector available for the time series prediction model through feature engineering based on the timestamp-aligned full-network state snapshot: on the one hand, the time series data such as bandwidth can be time-series processed to retain its dynamic change characteristics; on the other hand, the network topology can be abstracted into a graph structure. Specifically, the TSN switch and terminal node (audio sending device and receiving device) can be taken as nodes, and the physical link between devices can be taken as edges. On this basis, a feature vector can be extracted for each node, such as the device state features including port buffer occupancy rate of the TSN switch node, and the application layer features including the audio algorithm buffer state of the terminal node; the link performance features including link delay, jitter, bandwidth, and error rate can be extracted for each edge, and various features can be quantified into numerical vectors (such as the buffer occupancy rate is mapped to a 0-1 interval normalized value) through standardization processing, and finally a graph structure feature vector with the node feature matrix, edge feature matrix, and adjacency matrix as the core is constructed to represent the full-network topology state as the input feature of the GCL and TSI mapping relationship in the dynamic optimization basis TSI allocation scheme.
[0090] It should be noted that the embodiment can also be improved on the basis of embodiment 1 and / or embodiment 2.
[0091] It can be found that, in the embodiment, the state snapshot generation subunit can comprehensively capture the real-time dynamics of device state, link performance, traffic characteristics, and application layer state by periodically collecting four types of data streams including end-to-end performance flow and network device state flow and generating a full-network state snapshot, thereby providing complete and synchronous original data support for subsequent optimization and avoiding scheduling decision deviation caused by information fragmentation; and the feature conversion subunit converts the full-network state snapshot into a graph structure feature vector representing the full-network topology state, and the complex network state data can be efficiently analyzed by the time series prediction model through structured processing, which not only retains the topology association of devices and links, but also realizes data standardization and feature extraction, so that the dynamic optimization of the GCL and TSI mapping relationship can be based on global and accurate network state features to make decisions, thereby improving the adaptability of the TSI allocation scheme to network changes, and also ensuring the determinism and quality of service of audio stream transmission in the TSN network.
[0092] Embodiment 5
[0093] Embodiment 5 of the present application relates to an audio transmission control system. Embodiment 5 is an improvement on the basis of embodiment 4, and the specific improvement is that in the embodiment, another specific implementation manner of the scheme optimization unit is provided.
[0094] Optionally, the scheme optimization unit can further include a state prediction subunit, a scheduling generation subunit, and a strategy execution subunit.
[0095] The state prediction subunit is configured to take the graph structure feature vector as input to predict bandwidth occupation, time delay, network jitter value and packet loss rate of each link in a future scheduling period; and the bandwidth occupation, time delay, network jitter value and packet loss rate form the predicted link performance index.
[0096] The scheduling generation subunit is implemented based on a reinforcement learning framework, and includes: receiving the predicted link performance index output by the state prediction subunit; generating an updated GCL and TSI mapping table by a reinforcement learning agent; and calculating a reward value according to a multi-objective optimization formula.
[0097] The policy execution subunit is configured to train the reinforcement learning agent by the reward value, and send the updated TSI allocation scheme output by the learning agent to the scheme output unit.
[0098] Specifically, in the embodiment, the scheme optimization unit cooperates with the state prediction subunit, the scheduling generation subunit and the policy execution subunit to construct a two-stage hybrid AI model architecture, so as to realize accurate prediction of network state and dynamic optimization of scheduling strategy.
[0099] The state prediction subunit acts on the first stage and can be implemented by using a graph neural network (GNN) model, and can also be implemented by using a GNN variant such as a space-time graph convolutional network (ST-GCN). In this way, the graph structure feature vector (i.e., a graph-structured network state snapshot) can be taken as input, and the correlation between nodes (switches and terminals) and edges (links) in the network topology can be deeply mined by graph convolution operation. For example, the propagation path of congestion from an upstream switch to a downstream node, the influence law of performance fluctuation of a single node on adjacent devices, and the like can be learned. Compared with the limitation of traditional LSTM or ARIMA model that can only process time series data, the state prediction subunit can accurately capture the spatial correlation and time sequence dynamics of the network state by virtue of the perception ability of the topology structure, and finally output the predicted link performance index in one or more scheduling periods in the future, including key parameters such as bandwidth occupation of each link, end-to-end time delay of each path, network jitter value and packet loss rate.
[0100] The scheduling generation subunit acts on the second stage and can be implemented by using an advanced reinforcement learning algorithm such as proximal policy optimization (PPO). Specifically, the predicted link performance index output by the state prediction subunit in the first stage can be taken as state input, the process of generating a GCL gating list and a TSI mapping table by a reinforcement learning agent can be defined as a high-dimensional action space, and a multi-objective reward function can be used to quantify the action effect to form a closed-loop learning mechanism of state-action-reward. For example, the multi-objective reward function can be: wherein R is a reward value, R 时延quantized value of the corresponding delay, R 抖动 quantized value of the corresponding jitter, P 丢包 penalty value corresponding to packet loss, is an adjustable weight. Compared with the traditional scheme, the traditional scheduler relies on fixed optimization algorithms such as PSO, and its decision logic is limited to the preset mathematical model, which is easy to fall into local optimum in a complex dynamic network environment, and can only achieve passive computational scheduling; while the scheduling generation subunit provided in the embodiment can enable the agent to discover non-intuitive optimal strategies in continuous exploration through the trial-and-error-reward mechanism of reinforcement learning, realizing the transition from computational scheduling scheme to dynamic strategy, thereby more flexibly adapting to network changes and improving the global optimality of the scheduling scheme.
[0101] Specifically, the strategy execution subunit can optimize the decision logic of the RL agent through the reward value, and at the same time pass the generated updated TSI allocation scheme to the scheme output unit to form a closed loop of prediction-decision-feedback-iteration. It can be seen that the hybrid architecture provided in the embodiment can not only improve the prediction accuracy with the help of the topology perception ability of GNN, but also optimize the scheduling strategy through the dynamic learning characteristics of RL, and finally realize the adaptive adjustment of the GCL and TSI mapping relationship in a complex network environment, ensuring low latency, low jitter and high reliability of audio stream transmission.
[0102] It should be noted that the embodiment can also be improved on the basis of any one or more of embodiments 1 to 3.
[0103] It is not difficult to find that in the embodiment of the application, the scheme optimization unit forms a closed loop mechanism of prediction-decision-optimization through the cooperative operation of the state prediction subunit, the scheduling generation subunit and the strategy execution subunit. Among them, the state prediction subunit takes the graph structure feature vector as input and can accurately predict future link performance indicators to provide basis for scheduling decision, thereby avoiding the hysteresis that may be caused by formulating strategies based on historical data; the scheduling generation subunit generates a GCL and TSI mapping table based on a reinforcement learning framework and calculates a reward value through a multi-objective formula, compared with traditional fixed algorithms, the "trial-and-error-reward" mechanism of reinforcement learning can explore a better non-intuitive strategy, breaking through the limitation of local optimum; the strategy execution subunit can use the reward value to continuously train the agent and output an updated scheme, so that the scheduling strategy can dynamically adapt to network changes. The three work together to realize adaptive optimization of the TSI allocation scheme, improve the determinism and quality of service of audio stream transmission in the TSN network, and effectively guarantee the transmission requirements of low latency and low jitter.
[0104] Embodiment 6
[0105] Embodiment 6 of the present application relates to an audio transmission control system. Embodiment 6 is an improvement based on embodiment 5, and the specific improvement is that in the present embodiment, a specific implementation of the state prediction subunit is provided.
[0106] Optionally, the state prediction subunit is specifically implemented by a graph neural network model; the graph neural network model is optimized by a closed-loop self-supervised mechanism:
[0107] The predicted link performance indicators output by the graph neural network model are compared with the actual collected link performance true values after the end of the next scheduling period to generate error signals; wherein the error signals are back-propagated to the graph neural network model to drive online adjustment of the parameters of the graph neural network model, and continuous iteration of the prediction capability is realized.
[0108] Specifically, the state prediction subunit can use a GNN model to realize continuous optimization through a closed-loop self-supervised mechanism: the GNN model first outputs predicted link performance indicators, and after the end of the next scheduling period, the system compares the predicted values with the actual collected link performance true values to generate error signals; then, the error signals are fed back to the GNN model through a back-propagation mechanism to automatically drive online adjustment of the model parameters, forming a complete closed loop of prediction-comparison-correction. As can be seen, through the technical solution provided in the present embodiment, true value labels do not need to be manually labeled, and the operation and maintenance cost can be reduced; through real-time error feedback, the GNN model can be continuously self-calibrated, which can enable the GNN model to dynamically adapt to complex scenarios such as network topology changes and traffic fluctuations; after long-term iteration, the prediction accuracy can be significantly improved, and the output link performance indicators (bandwidth occupation, delay, etc.) can be closer to the actual network state.
[0109] For example, in a certain scheduling period, the GNN model predicts that the bandwidth occupation of link L1 is 50 Mbps, the delay is 200 μs, and the jitter is 10 μs based on the current graph structure feature vector. After the end of the next scheduling period, the system actually collects the bandwidth occupation of L1 as 55 Mbps, the delay as 210 μs, and the jitter as 12 μs, and generates error signals (bandwidth error 5 Mbps, delay error 10 μs, and jitter error 2 μs) through comparison. The error signals are fed back to the GNN model through a back-propagation mechanism to automatically adjust the graph convolution layer weight parameters related to link L1 in the model (such as increasing the sensitivity to the adjacent switch cache occupation rate feature). In subsequent periods, the optimized GNN model corrects the bandwidth occupation prediction of L1 to 54 Mbps and the delay to 208 μs, and the deviation from the actual value is significantly reduced, and the prediction accuracy is continuously improved through the closed loop of prediction-comparison-correction.
[0110] Optionally, in some embodiments, in the reinforcement learning framework of the scheduling generation subunit, the applicant designs the reward value calculation formula as a hybrid mechanism that balances performance improvement and optimal preservation. Specifically, the reward value calculation formula is as follows:
[0111] ;
[0112] Where ΔD represents the end-to-end delay optimization rate, This indicates the current jitter value. This indicates the current packet loss rate. Indicates the minimum value. This indicates the current critical link bandwidth utilization rate; R represents the reward value and is a configurable weighting coefficient.
[0113] Specifically, in this embodiment, the reward calculation formula is used for scheduling policy evaluation under the reinforcement learning framework, quantifying network performance through four dimensions and weighted integration:
[0114] Latency optimization (α·ΔD): ΔD represents the end-to-end latency optimization rate, i.e., the reduction in current latency compared to the baseline state. α-weighting is used to incentivize performance improvement. Higher rewards are given as long as latency can be further reduced (the larger ΔD is), driving the model to continuously optimize latency.
[0115] jitter control ( ): This is the current jitter value; the smaller the jitter, the better. The larger the β-weighted value, the more it encourages maintaining the optimal state. That is, the smaller the jitter, the larger the reciprocal. Even if the network is already in a low-jitter state (such as in a live streaming scenario), maintaining low jitter can still yield high rewards, avoiding the exploration of worse solutions due to insufficient rewards.
[0116] Packet loss suppression ( ): This is the current packet loss rate. Minimum value (e.g.) (This avoids the denominator being invalid when the packet loss rate is 0), taking the reciprocal. Furthermore, by using γ-weighting, optimal performance is also encouraged. That is, the lower the packet loss rate, the larger the reciprocal, thus adapting to scenarios sensitive to packet loss, such as file transfer.
[0117] Bandwidth utilization ( ): Represents critical link bandwidth utilization, directly through Weighted averaging encourages efficient bandwidth utilization and can balance the conflict between quality and resource utilization (such as avoiding bandwidth waste during low loads).
[0118] The above formula uses configurable weights The priority requirements of different services for delay progress, jitter maintenance, packet loss maintenance, and bandwidth utilization (such as real-time live adjustment of β and file transmission adjustment of γ) can be flexibly adapted. The problem of insufficient optimal state reward caused by only using relative optimization rate is solved. The multi-dimensional weighted guidance reinforcement learning strategy continuously optimizes network performance and stably maintains high-quality state, and maximizes the reward-driven scheduling effect.
[0119] For example, in the application scenario of the audio transmission control system, taking an actual scheduling optimization process as an example: the exchange scheduling layer needs to dynamically generate a TSI allocation scheme to guarantee audio stream transmission, the sending end module of the interface layer embeds the TSI into the AVTP frame according to the scheme, and the receiving end module parses the TSI and feeds back network performance indicators (initial delay D0=90 ms, jitter J0=12 ms, packet loss rate L0=1.2%, and key link bandwidth utilization U0=55%). After reinforcement learning optimization, the measured delay optimization rate ΔD=(90-72) / 90=20% (the delay is reduced to 72 ms), the current jitter value =8 ms, the current packet loss rate =0.6%, and the current key link bandwidth utilization =68%. The configurable weight coefficients are configured as α=0.3 (emphasizing delay optimization), β=0.25 (focusing on jitter control), γ=0.25 (valuing packet loss suppression), and =0.2 (considering bandwidth utilization). The minimum value =0.0001. Substituting the above reward value calculation formula, the delay optimization item contribution is 0.3×20%=0.06, the jitter control item contribution is 0.25×1 / 8=0.03125, the packet loss suppression item contribution is 0.25×0.6%+0.00011≈0.0416, and the bandwidth utilization item contribution is 0.2×68%=0.136. The total reward value R≈0.06+0.03125+0.0416+0.136=0.26885. The exchange scheduling layer can use this reward value in combination with the reinforcement learning framework. If the reward is high, the decision logic of the generated TSI allocation scheme is strengthened. If the network performance deteriorates (such as increased delay and increased packet loss rate, resulting in low reward), the corresponding strategy is suppressed. The continuous exploration of a more optimal TSI allocation scheme is driven through the index feedback-reward calculation-scheme optimization closed loop, to guarantee stable and efficient transmission of audio streams in the TSN network. The terminal processing layer can also adjust the audio processing algorithm start time based on the optimized TSI, accurately align the scheduling time window, and improve the overall quality of audio transmission.
[0120] Optionally, in some embodiments, the learning agent can obtain the learning agent by using an offline-online hybrid training mechanism.
[0121] In the offline stage, training data is generated by simulating different types of traffic loads in a digital twin simulation environment, and initial weights of the learning agent are determined according to the training data;
[0122] In the online stage, online learning is performed based on the initial weights.
[0123] Specifically, in the offline stage, a large number of and diversified traffic load scenarios (such as different audio stream concurrency, burst traffic impact, etc.) can be simulated in the digital twin simulation environment to generate training data covering various network states, so as to train and determine the initial weights of the learning agent; in the online stage, the learning agent can continuously learn online based on the initial weights obtained through offline training, in combination with the interactive data (such as link performance changes, reward feedback of scheduling strategies, etc.) generated in real time in actual network operation, and continuously optimize model parameters through dynamic iteration. In this way, the learning can master the basic strategies for coping with complex scenarios in the early stage of deployment by taking advantage of the simulation advantages of the digital twin environment, and can adapt to the dynamic characteristics of the real network through online learning, which is conducive to the continuous evolution of decision-making ability.
[0124] Optionally, in some embodiments, the updated TSI allocation scheme output by the scheme output unit includes two structured configuration files:
[0125] a GCL file for defining the opening and closing time of each port queue gate of each TSN switch;
[0126] a TSI mapping table for defining the association between each audio stream and the corresponding TSI identifier;
[0127] The configuration files are sent to all TSN switches and terminal nodes to ensure that the new configuration is enabled synchronously in the whole network to realize closed-loop control.
[0128] Specifically, the GCL file can explicitly define the opening and closing time of each port queue gate of each TSN switch, and can guarantee the transmission determinacy by precisely controlling the forwarding time of data frames; the TSI mapping table establishes the association between each audio stream and the corresponding TSI identifier to ensure the precise binding of streams and time slice resources. After the two configuration files are obtained by the switching and scheduling layer through the control channel, they are synchronously sent to all TSN switches and terminal nodes to realize the atomic update and network-wide enablement of the configuration. In this way, the generated optimization strategy can be quickly converted into specific configurations executable by devices, forming a complete closed loop from strategy output to network-wide execution, which is conducive to guaranteeing the dynamic adaptation of TSN networks to audio stream transmission requirements.
[0129] It should be noted that the present embodiment can also be an improvement based on any one or more of embodiments 1 to 4. It should be noted that the present embodiment can also be an improvement based on any one or more of embodiments 1 to 4.
[0130] It can be found that, in the embodiments of the present application, the state prediction subunit can adopt a graph neural network model and be optimized through a closed-loop self-supervised mechanism. Since the predicted link performance indicators output by the model can be compared with the actual true values in the next scheduling period to generate error signals, and the error is automatically adjusted by back propagation of the model parameters, this can enable the model to realize dynamic iteration of parameters without manual annotation, i.e., dynamic scenarios such as network topology changes and traffic fluctuations can be converted into correction signals in real time to continuously adapt the prediction logic to the actual environment; the continuous improvement of prediction accuracy can provide more accurate link performance input for the scheduling generation subunit, avoiding invalidation of the scheduling strategy due to prediction deviation; further, the obtained optimization scheme can more effectively match the audio stream transmission demand, guarantee low-latency and low-jitter transmission of audio data in the TSN network, and form a virtuous cycle of accurate prediction → scheduling optimization → reliable transmission.
[0131] Embodiment 7
[0132] Embodiment 7 of the present application relates to an audio transmission control system. Embodiment 7 is an improvement based on embodiment 1, and the specific improvement lies in that in the present embodiment, a specific implementation manner of a receiving end module is provided.
[0133] Specifically, the receiving end module can include:
[0134] a frame analysis module, configured to analyze TSI and a global timestamp in the AVTP frame;
[0135] a performance calculation module, configured to calculate end-to-end transmission delay, network jitter value and packet loss rate based on frame arrival time and the global timestamp;
[0136] an indicator feedback module, configured to feed back network performance indicators to the exchange scheduling layer.
[0137] Specifically, the frame analysis module can locate a custom extension field of the AVTP frame payload header, extract the TSI field therefrom, and analyze the PTP timestamp, which records the global time of the sending time of the sending end module.
[0138] Specifically, the performance calculation module can calculate the single-frame transmission delay by comparing the frame arrival time (obtained from the local PTP clock) with the global timestamp carried in the frame. With the aid of a sliding window algorithm, the change in the delay is continuously tracked, and a delay distribution histogram is generated in real time, and the extreme value and network jitter are calculated. For the calculation of the packet loss rate, a serial number comparison mechanism can be used to support the recovery of continuous serial numbers across network fragments, for example, a statistical report containing 1000 frames of performance data is generated every 20 ms.
[0139] The extension of the TSI field and the read-write mechanism can be implemented in the following manner: 16-bit space is reserved in a payload header self-defined extension field of an AVTP frame (complying with the IEEE 1722 standard) to store the TSI, after the sending end module receives the GCL and TSI mapping relationship issued by the exchange scheduling layer, the corresponding TSI is written into the extension field when the AVTP frame is encapsulated; the receiving end module locates and reads the field in the decapsulation process, and simultaneously associates with the PTP timestamp in the frame to provide a time slot alignment basis for the terminal processing layer, the whole process does not change the main structure of the AVTP protocol, and the binding of the TSI and the frame data and the end-to-end transmission are only realized through the self-defined extension.
[0140] Specifically, the index feedback module can feed back the performance calculation result to the exchange scheduling layer through a third interface (interface C). The feedback data is packaged in a TLV format and contains: end-to-end delay distribution (frame proportion of 10 intervals), jitter statistical extreme value (maximum value, minimum value, average value), packet loss rate (subdivided according to service type), and bit error rate (based on the AVTP frame CRC check result). The double-buffering mechanism can be used to realize the decoupling of data acquisition and sending, and to support lossless index acquisition under burst traffic. The feedback period can be configured (default 100 ms), and can be adaptively adjusted based on network load. When the network is congested, the feedback priority of key indicators (such as audio stream delay) is automatically improved to ensure the timeliness of scheduling optimization.
[0141] For example, when an AVTP frame containing audio data arrives at the terminal processing layer of the receiving end, the frame analysis module can first analyze the frame, extract the TSI identifier (such as "TSI_001") and the global timestamp (such as 10:00:00.123456) carried in the frame; then, the performance calculation module can compare the global timestamp with the local time (such as 10:00:00.123656) when the frame actually arrives at the receiving end, calculate the end-to-end transmission delay as 200 μs, calculate the jitter as 15 μs through the arrival time deviation of 10 consecutive frames, and determine the packet loss rate as 10% by detecting the missing frame (such as the 5th frame not arriving); the index feedback module can package the network performance indicators including "packet loss rate 10%, delay 200 μs, jitter 15 μs", and feed back to the exchange scheduling layer through the internal interface.
[0142] Optionally, in some embodiments, the terminal processing layer can be specifically used for:
[0143] Based on the parsed TSI and the corresponding global timestamp, dynamically adjusting the start time of the audio processing algorithm to align the start time with the scheduling time window mapped by the TSI;
[0144] The dynamic adjustment includes at least one of the following operations: triggering an audio processing algorithm according to a TSI mapped scheduling time window boundary; adjusting an audio buffer depth according to the end-to-end transmission delay; dynamically switching an audio redundancy error correction strategy according to the packet loss rate; and dynamically modifying an anti-jitter parameter of the audio processing algorithm according to the network jitter value.
[0145] Specifically, the terminal processing layer receives the AVTP frame with TSI from the adaptation layer through the fourth interface (interface D), analyzes the TSI and the PTP timestamp, compares with the local PTP clock to calculate the frame arrival time deviation and the time slot position, and dynamically adjusts the parameters of the audio processing algorithm (such as echo cancellation, dynamic EQ, noise reduction, etc.) on this basis, to realize deep coordination with network scheduling. Specifically, the algorithm execution can be triggered according to the scheduling time window boundary mapped by TSI, to ensure that the echo cancellation, noise reduction and other processes are started in the period when the data transmission is most stable, and the processing efficiency is improved; the audio buffer depth can be optimized through a delay prediction model according to the end-to-end transmission delay distribution, to avoid stuttering or data overflow caused by delay fluctuation; the audio redundancy error correction strategy can be dynamically switched according to the change of the packet loss rate, to enhance the data fault tolerance capability; the filter coefficient and other anti-jitter parameters can be adaptively adjusted according to the network jitter value, to improve the adaptability of the algorithm to the unstable network. Finally, through periodic reception of TSI and timestamp updates, a closed-loop dynamic adjustment mechanism of algorithm parameters is formed, to ensure that the terminal processing process is strictly aligned with the network scheduling time slot, which is conducive to fully guaranteeing the coordination and stability of audio processing and network transmission.
[0146] It should be noted that the present embodiment can also be an improvement based on any one or more of embodiments 2 to 6.
[0147] It can be understood that in the related art, the audio transmission system does not construct a systematic performance feedback link, lacking both a real-time collection mechanism for key indicators such as delay and jitter in the end-to-end transmission process, and an effective path for orderly feedback of these indicators to the scheduling core. This results in that the network manager cannot obtain quality data of the audio stream in the transmission process in real time, it is difficult to accurately evaluate the transmission stability, and it is impossible to dynamically adjust the scheduling strategy according to the actual performance, so that the system cannot quickly identify the transmission quality problem when facing network load changes or sudden conditions, and thus cannot guarantee the determinacy and reliability of audio transmission.
[0148] It can be found that, in the embodiment of the application, the receiving end module parses the TSI at the same time, generates network performance indicators in combination with the frame arrival time and the global timestamp, and feeds back to the exchange scheduling layer, so that the TSI can become an anchor point for performance testing, so that indicators such as delay and jitter can be accurately associated to specific scheduling time slots, a complete link of transmission-> analysis-> timing-> feedback is constructed, data support is provided for scheduling optimization, and the technical problems of quantifiable and traceable transmission quality and the blank of the end-to-end performance testing mechanism of related technologies are solved.
[0149] Embodiment 8
[0150] Embodiment 8 of the application relates to an audio transmission control system. Embodiment 8 is an improvement based on embodiment 1, and the specific improvement is that in the embodiment, the system further includes a protocol bridging layer deployed in the audio sending device and the audio receiving device, so that the audio transmission control system is specifically a five-layer distributed architecture as shown in Figure 3
[0151] In the embodiment, by logically dividing the audio transmission control system into five functional layers, the explicit interface and data flow realize efficient cooperation of each layer, and at the same time, the separation of the data plane and the control plane is realized, which ensures high cohesion and low coupling of each module, and is convenient for expansion and maintenance. The control plane is mainly composed of the exchange scheduling layer, which undertakes the responsibilities of global network perception, intelligent scheduling and centralized management; the data plane covers the time synchronization layer, the connection layer, the terminal processing layer and the protocol bridging layer, which are respectively responsible for the execution of specific tasks such as time synchronization establishment, data encapsulation and decapsulation, transmission and audio processing, and guarantee the stable operation of the system through layered cooperation.
[0152] Optionally, the protocol bridging layer can include:
[0153] The protocol receiving module is configured to receive a data frame of a non-TSN bus protocol;
[0154] The format conversion module is configured to parse and re-encapsulate the data frame into an AVTP frame compatible with the TSN;
[0155] The time window binding module is configured to bind the re-encapsulated audio stream to a TSN scheduling time window according to a TSI distribution scheme issued by the exchange scheduling layer;
[0156] The protocol restoration module is configured to restore the encapsulated data from the TSN backbone network into a data frame in the original non-TSN protocol format.
[0157] Specifically, the non-TSN bus protocol can include, but is not limited to, CAN, A2B, FlexRay, etc. The protocol receiving module can ensure that various non-TSN data can smoothly enter the protocol bridging layer for subsequent processing by adapting the physical interfaces and communication protocols of different heterogeneous buses.
[0158] Specifically, the format conversion module can parse the non-TSN data frame transmitted by the protocol receiving module, extract the valid data and control information therein, and then re-encapsulate these information into TSN-compatible AVTP frames according to the specifications of the TSN network, so as to realize the format conversion of non-TSN protocol data to TSN protocol data, and make the data originally only transmitted on the heterogeneous bus transparently transmitted on the TSN backbone network.
[0159] Specifically, in order to make the converted audio stream meet the scheduling requirements of the TSN network, the time window binding module can bind the re-encapsulated AVTP frame to the corresponding TSN scheduling time window according to the TSI allocation scheme issued by the exchange scheduling layer. Through this binding, it is ensured that these data frames can be transmitted within the predetermined time slot of the TSN network, which is consistent with the overall scheduling mechanism of the TSN network, and guarantees the real-time and determinacy of data transmission.
[0160] Specifically, the protocol restoration module can parse the AVTP frame from the TSN backbone network, extract the original non-TSN protocol data information, and then restore it to the data frame in the corresponding non-TSN protocol format and send it out through the corresponding heterogeneous bus interface. In this way, the reverse conversion of TSN protocol data to non-TSN protocol data can be realized, ensuring that the data can be normally transmitted in the heterogeneous bus network, and further improving the interconnection and intercommunication between the TSN network and the heterogeneous bus network.
[0161] It should be noted that the present embodiment can also be an improvement based on any one or more of embodiments 2 to 7.
[0162] It can be found that, in the present embodiment, the protocol bridging layer can realize the bidirectional transmission and protocol conversion of data between the TSN network and the heterogeneous bus network through the cooperative work of the above-mentioned modules, with the help of the TSN backbone network interface and the multiple heterogeneous bus interfaces, so as to guarantee the compatibility of the system to multiple vehicle-mounted or industrial buses.
[0163] It is worth mentioning that each module involved in the embodiment is a logical module. In actual application, one logical unit can be one physical unit, or a part of one physical unit, or realized in combination of multiple physical units. In addition, in order to highlight the innovative part of the present application, units not closely related to solving the technical problems proposed in the present application are not introduced in the embodiment, but this does not mean that there are no other units in the embodiment.
[0164] Embodiment 9
[0165] Embodiment 9 of the present application relates to an audio transmission control method. The method can be applied to the system as described in any one or more of embodiments 1 to 8.
[0166] As shown in Figure 4 , the method can at least include the following steps:
[0167] Step S101, generating a basic TSI allocation scheme containing the mapping relationship between GCL and TSI and issuing, so that the audio stream is bound to the TSN scheduling time window;
[0168] Step S102, embedding TSI in the AVTP frame extension field according to the TSI allocation scheme, and sending to the TSN network in the mapped scheduling time window;
[0169] Step S103, parsing the TSI and global timestamp in the received AVTP frame, generating network performance indicators and feeding back;
[0170] Step S104, dynamically aligning the start time of the audio processing algorithm with the scheduling time window based on the parsed TSI;
[0171] Step S105, dynamically optimizing the TSI allocation scheme of the next period by using the feedback network performance indicators;
[0172] Wherein, the above steps are continuously and cyclically run to realize adaptive scheduling.
[0173] The above steps are described in detail as follows.
[0174] For step S101, for example, after the system is started, the time synchronization layer can establish a unified time reference to complete time synchronization through the PTP protocol; then, the exchange scheduling layer can generate a basic TSI allocation scheme containing the mapping relationship between GCL and TSI according to the preset configuration such as the quality of service requirement of the audio stream, and issue it to all sending modules and receiving modules in the network through the second interface (interface B), so that the audio stream is bound to the corresponding TSN scheduling time window.
[0175] For step S102, the sending module can receive and encapsulate the audio data according to the distributed basic TSI allocation scheme, embedding the corresponding TSI in the custom extension field of the AVTP frame. After encapsulation, the adaptation layer can send the AVTP frame through the TSN network according to the precise time slot specified by the GCL, ensuring that the audio stream is transmitted within the predetermined time window.
[0176] For step S103, the receiving module receives the AVTP frame from the TSN network, parses it, extracts the TSI and global timestamp, and transmits them to the terminal processing layer. Meanwhile, the adaptation layer calculates the end-to-end transmission delay, network jitter value, and other network performance indicators based on the frame arrival time and global timestamp, and calculates the audio packet loss rate. The network performance indicators are fed back to the exchange scheduling layer through the third interface (interface C).
[0177] For step S104, the terminal processing layer receives the TSI and corresponding global timestamp from the adaptation layer, and can dynamically adjust the start time of the audio processing algorithm (such as echo cancellation, noise reduction, etc.) based on this information, so that the start time is accurately aligned with the scheduling time window mapped by the TSI, achieving deep synchronization between the terminal audio processing algorithm and the network state.
[0178] For step S105, the intelligent prediction and optimization module of the exchange scheduling layer can collect and analyze the network performance indicators fed back from the receiving module, learn and predict based on historical data, and determine whether the current scheduling scheme can meet the quality requirements of audio transmission. According to the analysis and prediction results, the intelligent prediction and optimization module dynamically recalculates and generates an optimized next-period TSI allocation scheme (including updated GCL and TSI mapping relationship), and then distributes it to each node through the second interface (interface B).
[0179] The above steps are continuously cycled to form a self-adaptive, intelligent closed-loop control system. By continuously generating schemes, transmitting data, analyzing feedback, and optimizing schemes, the system can respond to network state changes in real time, continuously ensure high-quality audio transmission, and achieve adaptive scheduling.
[0180] The above steps are only for clear description, and can be combined into one step or split into multiple steps during implementation, as long as the same logical relationship is included, and all are within the protection scope of the present application. Adding insignificant modifications or introducing insignificant designs in the algorithm or process, but not changing the core design of the algorithm and process, are within the protection scope of the present application.
[0181] It can be found that the embodiment is a method embodiment corresponding to the embodiment 1, and the embodiment can be implemented in cooperation with the embodiment 1. The related technical details mentioned in the embodiment 1 are still valid in the embodiment, and are not described here again in order to reduce repetition. Accordingly, the related technical details mentioned in the embodiment can also be applied in the embodiment 1.
[0182] Embodiment 10
[0183] The embodiment 10 of the present application relates to a specific application example of an audio transmission control system and method.
[0184] Exemplarily, the method and system can be applied to a vehicle-mounted surround sound system. The deployment and operation process of the vehicle-mounted surround sound system is as follows: the system deploys PTP master / standby clock and local crystal oscillator to constitute a time synchronization layer on a vehicle-mounted main switch and an auxiliary switch; an exchange scheduling layer sends an initial GCL / TSI to a vehicle-mounted AVTP sending module through a management port; in the adaptation layer, the sending module embeds TSI in a PCM audio stream, and a receiving module is responsible for alignment and playback; distributed audio station (DAS) nodes of each sound channel in the cabin serve as a terminal processing layer to dynamically run noise reduction and three-dimensional surround algorithms, without a protocol bridging layer. After the system is started, a unified time base is established by initializing PTP synchronization, and then GCL and TSI mapping are sent through the management port, and then the AVTP sending module encapsulates and sends an audio frame, the receiving module unpacks and calibrates the playback and feeds back timing data, the exchange scheduling layer updates GCL / TSI based on the feedback, and an iterative and adaptive scheduling process is formed.
[0185] Exemplarily, the method and system can be applied to an industrial plant broadcasting system. The deployment architecture of the industrial plant broadcasting system is as follows: a time synchronization layer deploys multi-source PTP on a plant core switch and a ring network switch to ensure that the time base of the whole network is unified; an exchange scheduling layer triggers dynamic updating of GCL based on real-time traffic monitoring indicators to realize flexible scheduling; a sending module of the adaptation layer completes AVTP broadcast frame encapsulation on a ring network node; each loudspeaker node of a terminal processing layer performs echo cancellation and reverberation compensation and other audio processing; and a protocol bridging layer can integrate an industrial monitoring system according to requirements to realize interfacing with an industrial bus. The operation process is basically the same as that of the vehicle-mounted surround sound system, and only the industrial bus interfacing requirements are adapted in the protocol interaction link to ensure the real-time performance and stability of the broadcast audio through cooperation of each layer.
[0186] Exemplarily, the method and system can be applied to intelligent building multi-room multimedia. The deployment architecture of the intelligent building multi-room multimedia is as follows: a time synchronization layer realizes clock synchronization of a building center machine room and each floor switch, providing a uniform time base for the system; a switching and scheduling layer allocates GCL and TSI according to floors, realizing hierarchical accurate scheduling; a connection layer is connected with a smart sound box through a floor AVTP sending end module / receiving end module, and is responsible for encapsulation and analysis of audio and video streams; a terminal processing layer of the smart sound box node performs dynamic EQ and scene recognition, optimizing audio output effect; and a protocol bridge layer integrates a building management system, realizing linkage with other systems of the building. The operation process focuses on multi-room synchronous playback and interactive control, and through cooperation of each layer, time consistency of audio playback in different rooms is ensured, and interactive operation of a user on multimedia content is supported, improving multimedia experience of the intelligent building.
[0187] The present application has at least the following beneficial effects: first, end-to-end synchronization accuracy is significantly improved, the system controls synchronization error of AVTP frames and GCL time slots to sub-microsecond level through deep binding of TSI and PTP timestamps; second, modular expansion capability is enhanced, each layer is independently designed and interacts through standardized interfaces, supporting single-level technology upgrade (such as upgrading time synchronization layer from PTPv2 to IEEE802.1AS-Rev) without affecting other layers; third, system maintenance complexity is reduced, standardized interface definition enables remote configuration management and fault diagnosis, combined with performance optimization suggestions of intelligent prediction and optimization modules, greatly improving operation and maintenance efficiency and functional expandability.
[0188] Embodiment 11
[0189] Embodiment 11 of the present application relates to an audio device. The audio device can be various forms of digital computers, such as a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and the like. The audio device can also be various forms of mobile devices, such as a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices.
[0190] The audio device includes one or more processors and a memory having stored computer program instructions that, when executed, cause the processor to perform the steps of the method provided by any one or more of the above embodiments. Figure 5An exemplary structural diagram of the audio device is disclosed. The audio device includes one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are interconnected by different buses, and can be mounted on a common main board or otherwise installed as needed. The processor can process instructions executed within the audio device, including instructions stored in the memory or on the memory to display a GUI on an external input / output device, such as a display device coupled to the interface, with graphic information. In some other embodiments, multiple processors and / or buses can be used with multiple memories and multiple memory, if necessary. Also, multiple audio devices can be connected, each providing part of the necessary operations. Among them, the components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or claimed herein.
[0191] The audio device can also include an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected by a bus or otherwise, and are connected by a bus in the figure.
[0192] The input device 1103 can receive input digital or character information, and generate key signal input related to user settings and function control of the audio device, such as touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (e.g., LED), and a tactile feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device can be a touch screen.
[0193] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. For example, it can be implemented by a special integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by the processor to implement the above steps or functions. Also, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a soft disk and the like. In addition, some steps or functions of the present application can be implemented by hardware, such as a circuit cooperating with the processor to perform the steps or functions.
[0194] The computer program product of the present application can be a computer program including a plurality of program instructions that control at least one processing unit of a computer to implement the method according to the embodiments of the present application. The computer program product of the present application can be stored in a computer readable storage medium (media), including but not limited to: optical read only memory (ROM), RAM, soft disc, and the like, which is coupled to the computer system bus via a coupling means, which enables the computer to read information in the computer program product, so as to implement the methods or achieve the device according to the embodiments of the present application.
[0195] The scope of the present application is defined by the appended claims rather than the description preceding it, therefore, all changes and modifications that come within the meaning and range of equivalency of the claims are to be embraced by the application. No feature of the application should be considered critical unless expressly stated in the claims. In addition, no single feature or combination of features should be considered essential unless expressly stated in the claims. Further, the mention herein of certain terms in the description and / or the claims should not be considered limiting, unless expressly stated to that effect. Specifically, any use of the terms "include", "have", "with", or the like, is not meant to exclude the presence of other elements or steps. It should be further noted that the terms "comprise", "comprising", "have", "has", "including", "including", "containing", "including" or any other variation thereof, are not meant to exclude other elements or steps. The terms "first", "second", "third", "fourth", "fifth" or the like are used merely to distinguish one element from another, and do not require a particular order or sequence. The use of the terms "first", "second", "third", "fourth", "fifth" or the like does not indicate relative importance.
[0196] The above description is merely illustrative of the application, and the scope of the application is not limited to the specific embodiments described herein. Any modification or variation of the present application that comes within the scope of the claims is intended to be included in the present application. Therefore, the scope of the present application should be determined not by the embodiments described above, but by the appended claims and their equivalents. The above embodiments are merely exemplary, and are not intended to limit the present application.
Claims
1. An audio transmission control system, characterized by, The system comprises: an exchange scheduling layer, configured to dynamically generate a time slot index (TSI) allocation scheme based on network performance indicators and to issue the TSI allocation scheme to a link layer; the link layer, comprising a sending end module and a receiving end module; the sending end module is deployed on an audio sending device and is configured to embed a TSI into an AVTP frame extension field according to the TSI allocation scheme and to send the AVTP frame to a network; the receiving end module is deployed on an audio receiving device and is configured to parse the TSI in an AVTP frame from the network, to generate network performance indicators, and to feed back the network performance indicators to the exchange scheduling layer; a terminal processing layer, deployed on the audio receiving device, and configured to dynamically adjust a starting time of an audio processing algorithm according to the parsed TSI, so that the starting time is aligned with a scheduling time window mapped by the TSI; wherein the exchange scheduling layer dynamically updates a TSI allocation scheme of a next period by using the fed-back network performance indicators, to form a closed-loop control.
2. The system of claim 1, wherein, The system further comprises a time synchronization layer, which comprises: a clock source module, configured to generate a clock synchronization signal and to perform a redundant switching of a master clock source, to provide a uniform global time reference for the system; a protocol execution module, configured to deliver the clock synchronization signal to a boundary clock, to synchronize TSN exchange devices, audio sending devices and audio receiving devices in the network through the boundary clock, and to eliminate clock deviations among the devices; an output module, configured to output the clock synchronization signal to the exchange scheduling layer and the terminal processing layer, to ensure time alignment of the TSI allocation scheme and reference synchronization of the starting time of the audio processing algorithm.
3. The system of claim 1, wherein, The exchange scheduling layer comprises: a configuration management unit, configured to manage network topology and configuration parameters of TSN switches; a scheduling management unit, configured to generate GCLs according to quality of service requirements of audio streams, to establish a mapping relationship between the GCLs and TSIs, and to generate a basic TSI allocation scheme; a scheme issuing unit, configured to issue the basic TSI allocation scheme to the link layer as a basis for execution of a first scheduling period.
4. The system of claim 3, wherein, The exchange scheduling layer further comprises an intelligent prediction and optimization module, which comprises: an indicator receiving unit, configured to receive network performance indicators from the receiving end module; a load prediction unit, configured to predict network load of a next period; a scheme optimization unit, configured to dynamically optimize the mapping relationship between the GCLs and the TSIs in the basic TSI allocation scheme according to the network performance indicators and the predicted network load; a scheme output unit, configured to output the updated TSI allocation scheme to the scheduling management unit.
5. The system of claim 4, wherein, The scheme optimization unit specifically comprises a state snapshot generation subunit and a feature conversion subunit; the state snapshot generation subunit is configured to periodically collect four types of data streams to generate a global network state snapshot; the four types of data streams comprise end-to-end performance streams, network device state streams, traffic statistics streams and application layer state streams; The feature conversion subunit is configured to determine a graph structure feature vector according to the whole-network state snapshot, wherein the graph structure feature vector is used to represent a whole-network topology state as an input feature for dynamically optimizing a GCL and TSI mapping relationship in the basic TSI allocation scheme.
6. The system of claim 5, wherein, The scheme optimization unit further includes a state prediction subunit, a scheduling generation subunit, and a strategy execution subunit. The state prediction subunit is configured to take the graph structure feature vector as an input to predict bandwidth occupation, time delay, network jitter value, and packet loss rate of each link in a future scheduling period, wherein the bandwidth occupation, time delay, network jitter value, and packet loss rate form predicted link performance indicators. The scheduling generation subunit is implemented based on a reinforcement learning framework, including receiving the predicted link performance indicators output by the state prediction subunit, generating an updated mapping relationship between the GCL and the TSI by a reinforcement learning agent, and calculating a reward value according to a multi-objective optimization formula. The strategy execution subunit is configured to train the reinforcement learning agent by the reward value, and send an updated TSI allocation scheme output by the learning agent to the scheme output unit.
7. The system of claim 6, wherein, The state prediction subunit is specifically implemented by a graph neural network model, and the graph neural network model is optimized by a closed-loop self-supervised mechanism. The predicted link performance indicators output by the graph neural network model are compared with actual collected link performance true values after the end of the next scheduling period to generate an error signal, wherein the error signal is back-propagated to the graph neural network model to drive online adjustment of parameters of the graph neural network model, thereby realizing continuous iteration of prediction capability.
8. The system of claim 6, wherein, In the reinforcement learning framework of the scheduling generation subunit, the calculation formula of the reward value is as follows: R = a-AD + b- + g- + d- + y- + s- ; wherein, ΔD represents an end-to-end delay optimization rate, represents a current jitter value, represents a current packet loss rate, represents a minimum value, represents a current key link bandwidth utilization rate; α, β, γ, δ are configurable weight coefficients, and R represents the reward value.
9. The system of claim 6, wherein, The learning agent is specifically obtained by an offline-online hybrid training mechanism: In the offline phase, training data is generated by simulating different types of traffic loads in a digital twin simulation environment, and initial weights of the learning agent are determined according to the training data; In the online phase, online learning is performed based on the initial weights.
10. The system of any one of claims 4 to 9, wherein, The updated TSI allocation scheme output by the scheme output unit includes two structured configuration files: A GCL file used to define opening and closing time points of each port queue gate of each TSN switch; A TSI mapping table used to define an association relationship between each audio stream and a corresponding TSI identifier; The configuration files are sent to all TSN switches and terminal nodes to ensure that the whole network synchronously enables the new configuration to realize closed-loop control.
11. The system of claim 1, wherein, The receiving end module includes: A frame analysis module configured to analyze TSI and global timestamps in an AVTP frame; A performance calculation module configured to calculate end-to-end transmission delay, network jitter value, and packet loss rate based on frame arrival time and the global timestamps; An indicator feedback module configured to feed back network performance indicators to the exchange scheduling layer.
12. The system of claim 11, wherein, The terminal processing layer is specifically configured to: Based on the analyzed TSI and corresponding global timestamps, dynamically adjust a start time of an audio processing algorithm, so that the start time is aligned with a scheduling time window of TSI mapping; The dynamic adjustment includes at least one of the following operations: triggering an audio processing algorithm according to a scheduling time window boundary of a TSI mapping; adjusting an audio buffer depth according to the end-to-end transmission delay; dynamically switching an audio redundancy error correction strategy according to a packet loss rate; and dynamically modifying an anti-jitter parameter of the audio processing algorithm according to the network jitter value.
13. The system of claim 1, wherein, The system further includes a protocol bridge layer deployed on an audio sending device and an audio receiving device, and includes: a protocol receiving module configured to receive a data frame of a non-TSN bus protocol; a format conversion module configured to parse and repackage the data frame into a TSN-compatible AVTP frame; a time window binding module configured to bind the repackaged audio stream to a TSN scheduling time window according to a TSI allocation scheme issued by an exchange scheduling layer; a protocol restoration module configured to restore the packaged data from the TSN backbone network into a data frame in an original non-TSN protocol format.
14. An audio transmission control method characterized by, The method is applied to the system according to any one of claims 1 to 13, and the method at least includes: generating a basic TSI allocation scheme containing a GCL and a TSI mapping relationship and issuing the basic TSI allocation scheme, so as to bind an audio stream to a TSN scheduling time window; embedding a TSI in an AVTP frame extension field according to the TSI allocation scheme and sending the TSI to a TSN network in a mapped scheduling time window; parsing the TSI and a global timestamp in a received AVTP frame, generating a network performance index, and feeding back the network performance index; dynamically aligning a starting time of an audio processing algorithm with a scheduling time window based on the parsed TSI; dynamically optimizing a TSI allocation scheme of a next period by using the fed-back network performance index.
15. An audio device, comprising: The audio device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the method of claim 14.
Citation Information
Patent Citations
Centralized network configuration entity and time sensitive network control system comprising same
CN114830611A
Real-time interactive digital human system supporting high concurrency and implementation method thereof
CN120179081A