Neural network-based network state prediction method and electronic equipment

By applying the TCN neural network architecture to the RTC system, the network status can be predicted in real time, solving the problem of unstable service quality in complex network environments and enabling accurate prediction and advance response to future network conditions.

CN121808326APending Publication Date: 2026-04-07ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing RTC systems struggle to maintain stable service quality in complex network environments, primarily due to lagging predictions and poor resilience to weak networks caused by network state lag and volatility.

Method used

A TCN-based neural network architecture is adopted to extract temporal features from real-time communication data packets through causal dilated convolution operations, construct the first temporal feature sequence, and use the prediction network to predict the network state of future time steps in real time, avoiding reliance on post-event statistics.

Benefits of technology

It enables real-time communication systems to maintain stable service quality in complex network environments, can anticipate network fluctuations, and improve service quality and bandwidth utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808326A_ABST
    Figure CN121808326A_ABST
Patent Text Reader

Abstract

The invention provides a neural network based on a TCN, and the neural network comprises a time sequence feature extraction network which is constructed based on the TCN and is used for carrying out the causal expansion convolution operation related to the TCN for an input first time sequence feature sequence, and generating a second time sequence feature sequence; wherein the first time sequence feature sequence is a first time sequence feature sequence corresponding to a real-time communication data packet sequence, which is constructed based on time sequence features extracted from transmission state data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence; the real-time communication data packet sequence is a sequence formed by real-time communication data packets generated in the real-time communication process of a sending end and a receiving end in the real-time communication system; and the prediction network is used for predicting the network state of the real-time communication system on at least one future time step based on the input second time sequence feature sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a network state prediction method and electronic device based on neural networks. Background Technology

[0002] In RTC (Real-Time Communication) systems, the unreliable UDP protocol is typically used for data transmission. During UDP data transmission, accurate network conditions are crucial input for downstream critical modules such as bandwidth estimation, bandwidth allocation, hybrid automatic repeat requests, and encoders. The accuracy of the network condition directly determines the effectiveness of Quality of Service (QoS) guarantees and the efficiency of bandwidth utilization.

[0003] However, in related technologies, the network state of RTC systems mainly relies on post-hoc feature statistics, which presents two major problems: First, the lag: the characteristics based on historical statistics cannot reflect changes in network status in a timely manner, resulting in the obtained network status lagging behind the actual network situation; Secondly, volatility: the statistical characteristics are greatly affected by instantaneous network jitter, and the values ​​fluctuate wildly, resulting in poor resistance of the QoS control system to weak networks.

[0004] These issues make it difficult for existing RTC systems to maintain stable service quality in complex network environments, often resulting in problems such as stuttering, degraded image quality, and sudden increases in latency. Summary of the Invention

[0005] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a network state prediction method based on a neural network is proposed, applied to a real-time communication system; the neural network includes a temporal feature extraction network constructed based on TCN and a prediction network; the method includes: The transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence is obtained, time-series features are extracted from the transmission status data, and a first time-series feature sequence corresponding to the real-time communication data packet sequence is constructed based on the extracted time-series features; wherein, the real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending end and receiving end in the real-time communication system during real-time communication. The first temporal feature sequence is input into the feature extraction network, so that the feature extraction network performs a causal dilated convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence. The second time-series feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence.

[0006] Optionally, the timing features include performance indicators calculated based on the transmission status data that can describe the network status of the real-time communication system; Extracting temporal features from the transmission status data and constructing an initial first temporal feature sequence based on the extracted temporal features, including: The performance index set corresponding to each real-time communication data packet contained in the real-time communication data packet sequence is calculated based on the transmission status data; wherein, the performance index set includes performance indexes of multiple dimensions calculated based on the transmission status data; A multi-dimensional time-series feature vector corresponding to each real-time communication data packet is constructed based on the set of performance indicators corresponding to each real-time communication data packet contained in the real-time communication data packet sequence; Based on the multi-dimensional temporal feature vectors corresponding to each real-time communication data packet within a preset time window contained in the real-time communication data packet sequence, a temporal feature matrix corresponding to the time window is further constructed as the first temporal feature sequence; wherein, the time window is a time window that slides according to a preset time step length, with the latest time step generated by the real-time communication as the endpoint. The second time-series feature sequence is further input into the prediction network so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence, including: The second time-series feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at the next or multiple time steps of the time window based on the second time-series feature sequence.

[0007] Optionally, the timing features may also include enhanced performance metrics obtained by further calculating the performance metrics.

[0008] Optionally, the performance metrics calculated based on the transmission status data include: A first type of performance indicator that can reflect the network transmission characteristics of the real-time communication system; A second type of performance indicator that can reflect the network latency characteristics of the real-time communication system; A third type of performance indicator that can reflect the network packet loss characteristics of the real-time communication system; A fourth type of performance indicator that can reflect the network stability characteristics of the real-time communication system; The enhanced performance index, calculated based on the performance index, includes: A fifth type of performance index is obtained by further calculation of multiple of the aforementioned performance indicators; A sixth type of performance index that can reflect the time-series characteristics of any of the aforementioned performance indices through further calculations. A seventh type of performance index is obtained by further calculation of any of the aforementioned performance indices, which can reflect the statistical characteristics of the performance indices.

[0009] Optionally, the temporal feature extraction network includes a TCN network and an attention network; The first temporal feature sequence is input into the feature extraction network, so that the feature extraction network performs convolution operations related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence. The first temporal feature sequence is input into the TCN network, so that the TCN network performs a causal dilation convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence; The second temporal feature sequence is further input into the attention network, so that the attention network performs attention calculations on each multidimensional temporal feature vector contained in the first temporal feature sequence and other multidimensional temporal feature vectors contained in the first temporal feature sequence to obtain hidden feature vectors corresponding to each multidimensional temporal feature vector; and, based on the hidden feature vectors corresponding to each multidimensional temporal feature vector, a hidden feature sequence is further generated.

[0010] Optionally, the second temporal feature sequence is further input into the prediction network so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the second temporal feature sequence, including: The hidden feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the hidden feature sequence.

[0011] Optionally, the neural network may further include a feature aggregation layer; The hidden feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the hidden feature sequence, including: The hidden feature sequence is input into the feature aggregation layer, so that the feature aggregation layer aggregates the hidden feature vectors contained in the hidden feature sequence into a feature vector of fixed length. The fixed-length feature vector is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector.

[0012] Optionally, the prediction network includes at least one fully connected layer and at least one output layer; The fixed-length feature vector is further input into the prediction network so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector, including: The fixed-length feature vector is further input into the at least one fully connected layer, so that the at least one fully connected layer can predict the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector, including: The fixed-length feature vector is further input into the at least one fully connected layer, so that the at least one fully connected layer performs a nonlinear transformation on the fixed-length feature vector to generate a target feature vector. The target feature vector is further input to the at least one output layer, so that the at least one output layer can predict the network state of the real-time communication system at at least one future time step based on the target feature vector, and output the network state.

[0013] Optionally, the network state includes multiple dimensions of network state indicators that can reflect the network state of the real-time communication system; the prediction network includes multiple output layers that correspond one-to-one with the multiple dimensions of network state indicators. The target feature vector is further input into the at least one output layer, so that the at least one output layer predicts the network state of the real-time communication system at at least one future time step based on the target feature vector, and outputs the network state, including: The target feature vector is input to the plurality of output layers respectively, so that the plurality of output layers predict the network state indicators of the real-time communication system at at least one future time step based on the target feature vector, and output the network state indicators to obtain the network state indicators of the real-time communication system in multiple dimensions at at least one future time step.

[0014] Optionally, the network state metrics of the multiple dimensions include: A first network status index is used to represent the congestion state of the real-time communication system; wherein, the first network status index includes a first probability value for representing network congestion in the real-time communication system. A second network status index is used to indicate the packet loss type of the real-time communication system; wherein the second network status index includes a second probability value indicating that the real-time communication system has no packet loss, a third probability value indicating that the real-time communication system has random packet loss, and a fourth probability value indicating that the real-time communication system has packet loss. A third network status indicator for representing the packet loss rate of the real-time communication system; wherein the third network status indicator includes a first packet loss rate for representing random packet loss in the real-time communication system, and a second packet loss rate for representing network congestion packet loss in the real-time communication system.

[0015] Optionally, the neural network is deployed at the transmitting end of the real-time communication system.

[0016] Optionally, the real-time communication system also deploys a tag generator; wherein the tag generator is used to add network status tags to the first time-series feature sequence based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, and generate a confidence score corresponding to the network status tags. The method further includes: The first time-series feature sequence is further input into the tag generator, so that the tag generator adds network status tags to the first time-series feature sequence based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, and generates a confidence score corresponding to the network status tags.

[0017] Optionally, based on the transmission status data corresponding to the real-time communication data packets generated in the next one or more time steps of the time window, a network status label is added to the first time-series feature sequence, and a confidence score corresponding to the network status label is generated, including: Based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, multiple sets of dynamic reference values ​​are calculated to determine the real network status of the real-time communication system in the next or multiple time steps of the time window. The multiple sets of dynamic reference values ​​are matched sequentially with multiple sets of preset marking rules; Based on the matching results of the multiple sets of dynamic benchmark values ​​and the multiple sets of labeling rules, network state labels are added to the first time-series feature sequence; Generate a confidence score corresponding to the network state label, including: Count the number of marking rules that match the multiple sets of dynamic benchmark values; Based on the number of labeling rules that match the multiple sets of dynamic benchmark values, the confidence score corresponding to the network state label is determined.

[0018] Optionally, the real-time communication system also deploys a network evaluator for assessing the prediction accuracy of the neural network; The method further includes: A time-series feature sequence sample is constructed based on the first time-series feature sequence, the network state label added to the first time-series feature sequence, and the confidence score corresponding to the network state label; The temporal feature sequence samples and the network state predicted by the prediction network are input into the network evaluator, so that the network evaluator can evaluate the prediction accuracy of the network state predicted by the prediction network by matching the network state label contained in the temporal feature sequence samples with the network state predicted by the prediction network.

[0019] Optionally, the real-time communication system also deploys a training manager for incremental training of the neural network; The method further includes: In response to the fulfillment of the triggering condition for incremental training of the neural network, a set of time-series feature sequence samples containing a confidence score greater than a preset threshold is selected from the time-series feature sequence samples input to the network evaluator. The neural network is incrementally trained based on the time-series feature sequence sample set; wherein the triggering condition includes the decrease in the prediction accuracy predicted by the network evaluator reaching a preset threshold.

[0020] According to a second aspect of one or more embodiments of this specification, a neural network is also provided, comprising: A temporal feature extraction network based on TCN is used to perform causal dilated convolution operations related to the TCN on the input first temporal feature sequence to generate a second temporal feature sequence. The first temporal feature sequence is a first temporal feature sequence corresponding to the real-time communication data packet sequence, constructed based on temporal features extracted from the transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence. The real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending and receiving ends in the real-time communication system during real-time communication. A prediction network is used to predict the network state of the real-time communication system at at least one future time step based on the second temporal feature sequence as input.

[0021] According to a third aspect of one or more embodiments of this specification, an electronic device is also provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor performs the executable instructions to implement the steps of the method as described in any of the first aspects above.

[0022] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is also provided, having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as described in any of the first aspects above.

[0023] According to a fifth aspect of one or more embodiments of this specification, a computer program product is also provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any of the first aspects above.

[0024] In the above embodiments, by applying the TCN network to the real-time prediction of network status in a real-time communication system, the special causal dilated convolution operation of the TCN network can be used to extract the temporal feature sequence for predicting the network status of the real-time communication system at future time steps from the transmission status data corresponding to the real-time communication data packets generated during the real-time communication process. Based on this temporal feature sequence, the network status of the real-time communication system at future time steps can be predicted in real time, eliminating the need to rely on post-event statistics. This allows the real-time communication system to respond in advance to possible network fluctuations based on the network status at future time steps, thereby maintaining stable service quality even in complex network environments. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the architecture of a real-time communication system provided in an exemplary embodiment; Figure 2 This is a flowchart illustrating a network state prediction method based on a neural network, provided in an exemplary embodiment. Figure 3 This is a network architecture diagram of an RTC system provided in an exemplary embodiment; Figure 4 This is a network architecture diagram of a TCN-based neural network provided in an exemplary embodiment; Figure 5 This is a network architecture diagram of another TCN-based neural network provided in an exemplary embodiment; Figure 6 This is a flowchart of an exemplary embodiment of performing a causal dilation convolution operation on a first temporal feature sequence; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment; Figure 8 This is a block diagram of a network state prediction device based on a neural network, provided in an exemplary embodiment. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0028] In related technologies, the mainstream solution for predicting the network status of RTC systems mainly adopts a post-hoc statistical approach. However, in RTC systems with high real-time requirements, this post-hoc statistical solution usually cannot predict the network status of the RTC system in future time steps in real time, making it difficult for the RTC system to maintain stable service quality in complex network environments.

[0029] For example, in RTC systems, the unreliable UDP protocol is typically used for data transmission. During UDP data transmission, accurate network conditions need to be predicted as input to downstream critical modules such as bandwidth estimation, bandwidth allocation, hybrid automatic repeat requests, and encoders. If the network conditions at future time steps cannot be predicted in a timely manner, these downstream critical modules cannot respond to potential network fluctuations in advance, making it difficult for the RTC system to maintain stable service quality in complex network environments.

[0030] Based on this, this specification proposes a novel neural network architecture based on TCN (Temporal Convolutional Network) and applies this neural network architecture to the RTC system to provide a solution for real-time prediction of the network state of the RTC system in future time steps.

[0031] In this solution, the aforementioned neural network may specifically include a temporal feature extraction network based on TCN and a prediction network.

[0032] After deploying the neural network to the RTC system, transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence can be obtained. Temporal features can be extracted from the transmission status data, and a first temporal feature sequence corresponding to the real-time communication data packet sequence can be constructed based on the extracted temporal features. The real-time communication data packet sequence is a sequence of real-time communication data packets generated by the sending end and receiving end in the RTC system during real-time communication. Furthermore, the first temporal feature sequence can be input into the feature extraction network, so that the feature extraction network can perform causal dilation convolution operation related to TCN on the first temporal feature sequence to generate a second temporal feature sequence for predicting the network state. Then, the second time-series feature sequence can be further input into the prediction network so that the prediction network can predict the network state of the RTC system at at least one future time step based on the second time-series feature sequence.

[0033] In the above solution, by applying the TCN network to the real-time prediction of network status in the real-time communication system, the special causal dilated convolution operation of the TCN network can be used to extract the temporal feature sequence for predicting the network status of the real-time communication system at future time steps from the transmission status data corresponding to the real-time communication data packets generated during the real-time communication process. Based on this temporal feature sequence, the network status of the real-time communication system at future time steps can be predicted in real time, eliminating the need for post-event statistics. This allows the real-time communication system to respond in advance to possible network fluctuations based on the network status at future time steps, thereby maintaining stable service quality even in complex network environments.

[0034] Figure 1 This is a schematic diagram of the architecture of a real-time communication system provided in an exemplary embodiment.

[0035] like Figure 1 As shown, the system may include a server 11, a network 12, and several electronic devices, such as a PC (Personal Computer) 13, a mobile phone 14, etc.

[0036] Server 11 can be a physical server containing an independent host, or it can be a virtual server hosted in a host cluster. During operation, server 11 can run a server-side program to implement functions related to network state prediction based on neural networks. For example, when server 11 runs the program, it can act as the server side of the real-time communication system.

[0037] PC23 and mobile phone 14 are just some of the types of electronic devices that users can use. In reality, users can obviously also use electronic devices such as tablets, laptops, PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smartwatches, etc.), etc., and one or more embodiments in this specification do not limit this. During operation, the electronic device can also run a client-side program to implement functions related to network state prediction based on neural networks. For example, when the electronic device runs this program, it can act as a client of the real-time communication system.

[0038] The client-side application of the aforementioned real-time communication system can be launched and run on an electronic device. This client-side application can be a native application installed on the electronic device, or it can be a mini-program, quick app, or other similar form. Alternatively, when using web technologies such as HTML5, the relevant functions can be implemented through a browser-displayed page. This browser can be a standalone browser application or a browser module embedded within some applications.

[0039] As for the network 12 that enables interaction between electronic devices such as PC13 and mobile phone 14 and server 11, communication can be achieved using either wired or wireless networks, depending on the communication methods supported by the respective electronic devices. This specification does not impose any restrictions on this. For example, PC13 can support both wired and wireless communication, so it can use either wired or wireless networks as needed. Mobile phone 14 typically only supports wireless communication, so it can use a wireless network for communication.

[0040] The technical solution of this specification will be described in detail below with reference to the accompanying drawings.

[0041] Please see Figure 2 , Figure 2 This document presents a flowchart of a network state prediction method based on a neural network, applied to an RTC system; the method includes the following execution process: Step 202: Obtain transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence, extract timing features from the transmission status data, and construct a first timing feature sequence corresponding to the real-time communication data packet sequence based on the extracted timing features; wherein, the real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending end and receiving end in the real-time communication system during real-time communication. The aforementioned RTC system may specifically include any form of communication system capable of enabling low-latency real-time data interaction over a network.

[0042] For example, in practical applications, the aforementioned RTC system may specifically include multimodal AI real-time interactive systems, real-time audio and video communication systems, real-time game communication systems, etc., and will not be specifically limited in this specification.

[0043] In some embodiments, to enable the RTC system to predict the network state at future time steps in real time, a TCN-based neural network can be deployed in the RTC system. This neural network can utilize the special causal dilated convolutions operation of TCN to capture long-term dependencies from the historical real-time communication sequences generated by the RTC system, and generate time-series feature sequences rich in high-level features for real-time prediction of the network state of the RTC system at future time steps.

[0044] It should be noted that an RTC system typically includes both a transmitter and a receiver; when applying the aforementioned neural network to an RTC system, the neural network can be deployed at either the transmitter or the receiver.

[0045] For example, see Figure 3 , Figure 3 This specification illustrates a network architecture diagram of an RTC system; such as Figure 3 As shown, in some embodiments, the neural network described above can be specifically deployed on the transmitting side of the RTC system.

[0046] By deploying the aforementioned neural network on the transmitting end of the RTC system, the network state predicted by the neural network for future time steps can be directly and instantly fed into the downstream key modules of the transmitting end; for example, such as... Figure 3As shown, these downstream key modules can specifically include core modules such as bandwidth estimation, bit rate allocator, and hARQ (Hybrid Automatic Repeat reQuest); thus forming a local high-speed closed loop of "perception-decision-execution", which can avoid the additional network round-trip delay caused by feeding the prediction results back from the receiving end to the sending end.

[0047] Please see Figure 4 , Figure 4 This specification illustrates a network architecture diagram of a TCN-based neural network.

[0048] like Figure 4 As shown, in some embodiments, the neural network described above may specifically include a temporal feature extraction network based on TCN, and may also include a prediction network. The aforementioned temporal feature extraction network based on TCN can be used to perform TCN-related causal dilation convolution operations on the first input temporal feature sequence to generate a second temporal feature sequence for predicting the network state. Specifically, the aforementioned prediction network can be used to predict the network state of a real-time communication system at at least one future time step based on the second temporal feature sequence of the input.

[0049] It should be noted that the above Figure 4 The network structure shown is a basic network structure of a TCN-based neural network. In practical applications, it can be adapted to meet specific needs. Figure 4 The network structure shown can be flexibly adjusted.

[0050] For example, both the aforementioned temporal feature extraction network and the aforementioned prediction network can be multi-layered neural network structures. In practical applications, the number of layers in the aforementioned temporal feature extraction network and the aforementioned prediction network can be reduced or expanded based on specific needs.

[0051] In some embodiments, during real-time communication, the transmitter and receiver in an RTC system can acquire the sequence of real-time communication data packets generated during the real-time communication process.

[0052] Specifically, the real-time communication data packet sequence can be a sequence of real-time communication data packets generated at each time step during the real-time communication process between the sending end and the receiving end.

[0053] For example, the sequence of real-time communication data packets can be denoted as [T1, T2, T3…Tn], where Tn can represent the real-time communication data packet generated at the nth time step.

[0054] After obtaining the sequence of real-time communication data packets generated during the real-time communication between the sending and receiving ends, the transmission status data corresponding to each real-time communication data packet contained in the sequence can be further obtained.

[0055] Specifically, the aforementioned transmission status data can include any form of data that reflects the transmission status of real-time communication data packets generated during real-time communication between the sending and receiving ends. For example, in practical applications, the aforementioned transmission status data can specifically include data attributes that reflect the transmission status of real-time communication data packets, observed or statistically obtained indicators, and so on.

[0056] In some embodiments, a packet statistics module and a feedback packet receiver are typically deployed at the RTC transmitter. In practical applications, the packet statistics module and the feedback packet receiver can be used as the data source for the aforementioned transmission status data, and the aforementioned transmission status data can be obtained from the packet statistics module and the feedback packet receiver.

[0057] In this case, the aforementioned transmission status data may include not only the transmission status data of the sending end obtained from the packet statistics module, but also the transmission status data of the receiving end obtained from the feedback packet receiver.

[0058] For example, the transmission status data of the sending end may specifically include data packet size, sending timestamp, sending rate, number of retransmissions, etc.; the transmission status data of the receiving end may specifically include receiving timestamp, packet loss flag, arrival interval, receiving rate, etc., which will not be listed one by one in this specification.

[0059] After obtaining the transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence, time-series features can be further extracted from the transmission status data, and a first time-series feature sequence corresponding to the real-time communication data packet sequence can be constructed based on the extracted time-series features.

[0060] The aforementioned timing characteristics can specifically include any form of characteristics that can describe the network state of the RTC system. This specification will not impose specific limitations on them, and in practical applications, they can be designed flexibly.

[0061] In one embodiment shown, the aforementioned timing characteristics may specifically include performance indicators that describe the network state of the RTC, calculated based on the aforementioned transmission state data.

[0062] For example, in some embodiments, the performance metrics calculated based on the aforementioned transmission status data may specifically include the following categories: The first category is the first type of performance indicators that can reflect the network transmission characteristics of the RTC system; For example, in practical applications, the first type of indicator mentioned above may specifically include: Send / receive rate, packet size distribution (e.g., average, standard deviation, range), send / receive packet ratio, etc.

[0063] The second type is a performance index that can reflect the network latency characteristics of the RTC system. For example, in practical applications, the second type of indicators mentioned above may specifically include: RTT (Round-Trip Time) value, RTT fluctuation value, minimum RTT value, etc.

[0064] The third type is a performance index that can reflect the network packet loss characteristics of the RTC system. For example, in practical applications, the third type of indicator mentioned above may specifically include: Packet loss rate, number of consecutive packet losses, packet loss interval, retransmission success rate, etc.

[0065] The fourth category is a performance index that reflects the network stability characteristics of an RTC system. For example, in practical applications, the fourth category of indicators mentioned above may specifically include: Jitter value, rate fluctuation coefficient, RTT value and correlation with transmission rate, etc.

[0066] In some embodiments, to enrich the types of the above-mentioned timing features, in practical applications, the above-mentioned timing features may also include enhanced performance indicators obtained by further calculating the above-mentioned performance indicators.

[0067] For example, in some embodiments, the enhanced performance metrics calculated based on the above performance metrics may specifically include the following categories: The first category is the fifth category of performance indicators, which are obtained by further calculation of multiple of the above-mentioned performance indicators; that is, the fifth category of performance indicators can be new indicators generated by further interaction of multiple performance indicators.

[0068] For example, in practical applications, the fifth type of performance index mentioned above may specifically include the transmission pressure index; the transmission pressure index may specifically refer to the index obtained by multiplying the transmission rate and the average packet size mentioned above.

[0069] The second category is a sixth type of performance index, which is obtained by further calculation of any of the above-mentioned performance indices and can reflect the time-series characteristics of that performance index; that is, the sixth type of performance index can be a new index that reflects the time-series characteristics of a certain performance index, which is obtained by further calculation of a certain of the above-mentioned performance indices.

[0070] For example, in practical applications, the sixth type of performance index mentioned above may specifically include the RTT change rate obtained by further calculating the RTT value, the rate change trend obtained by further calculating the transmit / receive rate, and so on.

[0071] The third category is a seventh category of performance indicators that can reflect the statistical characteristics of any of the above-mentioned performance indicators through further calculations; that is, the seventh category of performance indicators can specifically be a new indicator that can reflect the statistical characteristics of a certain of the above-mentioned performance indicators through further calculations.

[0072] For example, in practical applications, the sixth type of performance indicators mentioned above may specifically include statistical quantities such as mean, variance, and maximum value obtained by further statistical calculations of all the indicators mentioned above.

[0073] In some embodiments, when extracting time-series features from the above-mentioned transmission status data and constructing an initial first time-series feature sequence based on the extracted time-series features, the performance index set corresponding to each real-time communication data packet contained in the above-mentioned real-time communication data packet sequence can be calculated first based on the above-mentioned transmission status data. Specifically, this set of performance metrics may include multiple dimensions of performance metrics that can describe the network status of the RTC, calculated based on the aforementioned transmission status data.

[0074] For example, the performance metric set can include the four categories of performance metrics mentioned above, totaling 18 dimensions, and the three categories of enhanced performance metrics mentioned above, totaling 3 dimensions. In this case, the performance metric set can contain a total of 21 dimensions of performance metrics.

[0075] Furthermore, a multi-dimensional time-series feature vector corresponding to each real-time communication data packet in the above-mentioned real-time communication data packet sequence can be constructed based on the set of performance indicators corresponding to each real-time communication data packet in the above-mentioned real-time communication data packet sequence; then, based on the multi-dimensional time-series feature vector corresponding to each real-time communication data packet in the above-mentioned real-time communication data packet sequence, a time-series feature matrix can be further constructed as the above-mentioned first time-series feature sequence.

[0076] For example, assuming the above real-time communication data packet sequence contains 32 real-time communication data packets, and the feature dimension of the above multi-dimensional time-series feature vector is 21-dimensional, then the above time-series feature matrix can specifically be a 32×21 tensor.

[0077] In some embodiments, to prevent the temporal features from fading over time, a dynamic sliding window mechanism can be used when constructing the aforementioned temporal feature matrix.

[0078] In this case, a time-series feature matrix corresponding to the time window can be further constructed based on the multi-dimensional time-series feature vectors corresponding to each real-time communication data packet within a preset time window contained in the above-mentioned real-time communication data packet sequence, so as to serve as the first time-series feature sequence; wherein, the time window can specifically be a time window that slides according to a preset time step length, with the latest time step generated by the sending end and the receiving end during real-time communication as the endpoint.

[0079] For example, taking the above real-time communication data packet sequence as [T1, T2, T3…Tn] as an example, assuming the length of the time window is 10, the feature dimension of the above multi-dimensional time series feature vector is 18, and the time step length used when the time window is sliding is 4.

[0080] In this scenario, if the latest time step is 32, the real-time communication data packets within the sliding time window contained in the aforementioned real-time communication data packet sequence can specifically include 10 real-time communication data packets, T23-T32, which can be denoted as [T23, T24, T25…T31, T32]. At this point, the time-series feature matrix constructed based on the 18-dimensional temporal feature vectors of these 10 real-time communication data packets can specifically be a 10×18 tensor constructed at time step 32. This tensor refers to the temporal feature sequence constructed at time step 32.

[0081] Since the time window uses a time step length of 4 when sliding, it will slide again at the latest time step of 36, and the same steps can be repeated. At this point, the real-time communication data packets contained in the above real-time communication data packet sequence within the sliding time window can specifically include 10 real-time communication data packets T27-T36 (i.e., slid 4 time steps to the right), which can be denoted as [T27, T28, T29…T35, T36]. Based on the 18-dimensional temporal feature vectors of these 10 real-time communication data packets, the constructed temporal feature matrix can be a 10×18 tensor constructed at time step 36. This tensor refers to the temporal feature sequence constructed at time step 36.

[0082] Similarly, since the length of the set time step is usually much smaller than the length of the time window according to this sliding window mechanism, it takes a certain amount of time to slide from the first time step to the last time step of a time window according to the set step length. At this time, for the time series feature corresponding to a certain time step in this time window, it is equivalent to caching the same amount of time, thereby achieving the effect of caching the time series feature and preventing the time series feature from fading away over time.

[0083] For example, assuming the time window length is 36, and the sliding window uses a time step length of 4, with each time step lasting 0.4 seconds, then sliding from the first time step to the last time step in the time window with a time step length of 4 requires a total of 36 / 4 slides. Since each time step lasts 0.4 seconds, the maximum buffer time for the corresponding timing feature at any time step is 36 / 4 × 0.4 = 3.6 seconds.

[0084] Step 204: Input the first temporal feature sequence into the feature extraction network, so that the feature extraction network performs causal dilation convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence; After constructing the first time-series feature sequence based on the multi-dimensional time-series feature vectors corresponding to each real-time communication data packet contained in the above-mentioned real-time communication data packet sequence, the first time-series feature sequence can be further input into the feature extraction network based on TCN contained in the above-mentioned neural network, so that the feature extraction network can perform causal dilation convolution operation related to TCN on the first time-series feature sequence to generate a second time-series feature sequence for predicting the network state.

[0085] Please see Figure 5 , Figure 5 This is a diagram illustrating another TCN-based neural network architecture shown in this specification.

[0086] In some embodiments, such as Figure 5 As shown, the aforementioned temporal feature extraction network may specifically include a TCN network and an attention network; for example, in practical applications, the aforementioned TCN network and attention network may each correspond to an independent neural network layer.

[0087] In this case, the first temporal feature sequence can be input into the TCN network, and the TCN network can perform a TCN-related causal dilation convolution operation on the first temporal feature sequence to generate a second temporal feature sequence for predicting the network state.

[0088] The aforementioned TCN network can typically be a multi-layered network composed of multiple layers of residual blocks. When performing TCN-related causal dilation convolution operations on the first temporal feature sequence, it is necessary to perform multi-level causal convolution operations on the first temporal feature sequence through the multiple layers of residual blocks according to different dilation rates.

[0089] Please see Figure 6 , Figure 6 This is a flowchart illustrating a causal dilation convolution operation on a first temporal feature sequence as shown in this specification.

[0090] like Figure 6 As shown, in this example, taking the first temporal feature sequence as a 32×21 tensor, the TCN network can specifically include a 1D convolutional layer for initial feature mapping of the first temporal feature sequence and a stacked four-layer residual block.

[0091] In this case, the first temporal feature sequence (a 32×21 tensor) can be input into a 1D convolutional layer for initial feature mapping.

[0092] For example, such as Figure 6 As shown, this 1D convolutional layer can use 32 filters and set the kernel size to 1. This convolutional layer can map 21-dimensional temporal features to a 32-dimensional feature space, preserving the original temporal relationship while compressing redundant information. The activation function uses the ReLU function to introduce a nonlinear transformation.

[0093] After the initial feature mapping is completed, the first temporal feature sequence (32×32 tensor) after the initial feature mapping can be input into the stacked four-layer residual block, and causal dilation convolution calculation is performed sequentially through residual block 1 to residual block 4.

[0094] For example, such as Figure 6 As shown, the size of the convolution kernels in the stacked 4 residual blocks can be uniformly set to 3, and the dilation rates can be set to 1, 2, 4, and 8 respectively. The first temporal feature sequence after the initial feature mapping first passes through residual block 1 and performs causal convolution calculation according to dilation rate 1. The output of residual block 1 is then further input into residual block 2 and continues to perform causal convolution calculation according to dilation rate 2. The output of residual block 2 is then further input into residual block 3 and continues to perform causal convolution calculation according to dilation rate 4. The output of residual block 3 is then further input into residual block 4 and continues to perform causal convolution calculation according to dilation rate 8. Finally, the output of residual block 4 is the second temporal feature sequence used to predict the network state.

[0095] By stacking in this way, the receptive field can be increased exponentially (for example, after causal convolution calculation of 4 layers of residual blocks, it can cover 31 time steps) to capture long-term dependencies (for example, it can capture periodic congestion precursors in RTC systems).

[0096] The specific process of causal convolution operations performed on each residual block will not be detailed in this specification. For example, causal convolution operations usually use left padding to ensure that the model only depends on the temporal features of the past and the current time step, avoiding the leakage of future temporal features and conforming to the temporal logic of real-time prediction.

[0097] Furthermore, after inputting the first temporal feature sequence into the TCN network, and having the TCN network perform a TCN-related causal dilation convolution operation on the first temporal feature sequence to generate a second temporal feature sequence for predicting the network state, the second temporal feature sequence can be further input into the attention network. The attention network then performs attention calculations on each multidimensional temporal feature vector contained in the first temporal feature sequence and other multidimensional temporal feature vectors contained in the first temporal feature sequence to obtain hidden feature vectors corresponding to each of the multidimensional temporal feature vectors. Based on the hidden feature vectors corresponding to each of the multidimensional temporal feature vectors, a hidden feature sequence for predicting the network state is further generated.

[0098] The process of performing attention calculations on each multidimensional temporal feature vector contained in the first temporal feature sequence and other multidimensional temporal feature vectors contained in the first temporal feature sequence to obtain the hidden feature vectors corresponding to each of the aforementioned multidimensional temporal feature vectors typically includes: The attention weights between the multidimensional time-series feature vectors and other multidimensional time-series feature vectors contained in the second time-series feature sequence are calculated using the query vectors corresponding to each multidimensional time-series feature vector and the key vectors and value vectors corresponding to other multidimensional time-series feature vectors contained in the second time-series feature sequence as calculation parameters. Then, based on attention weights, the Value vectors corresponding to each multidimensional temporal feature vector contained in the second temporal feature sequence are weighted and calculated to obtain the hidden vectors corresponding to each multidimensional temporal feature contained in the second temporal feature sequence. The specific calculation process and formulas will not be described in detail in this specification; those skilled in the art can refer to relevant technologies.

[0099] By adding an attention network after the TCN network, the weight coefficients of the temporal features at different time steps can be calculated, thereby enhancing the model's sensitivity to key network events (such as sudden congestion and instantaneous packet loss).

[0100] For example, in practical applications, not all temporal features in a long sequence of time-series features are equally important. The model needs a mechanism to "focus" on the time steps most critical for predicting the network state. Adding an attention network after the TCN network can achieve the following effect: First, it can automatically focus on key time steps: When analyzing temporal feature sequences, attention networks no longer treat all temporal features equally, but automatically learn and assign higher weights to key time steps that are more important for predicting the network state (such as sudden packet loss or the moment when RTT begins to rise sharply). For example, if an RTC system experiences a sudden, instantaneous packet loss, the attention network can increase the weight of features associated with moments of high packet loss. Secondly, it can enhance the capture of weak signals: making the model more sensitive to early, weak but persistent trends that indicate changes in network state (such as the slow rise of RTT), which helps to achieve earlier and more accurate predictions.

[0101] In short, the attention layer adds the ability of "selective attention" to the model, enabling it to intelligently filter out the key time steps that are most valuable for predicting the network state from the continuous historical data stream, and make more reliable judgments based on the temporal features corresponding to these key time steps.

[0102] Step 206: The second time-series feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence.

[0103] After the first temporal feature sequence is further input into the TCN-based feature extraction network included in the neural network, and the feature extraction network performs a TCN-related causal dilation convolution operation on the first temporal feature sequence to generate a second temporal feature sequence for predicting the network state, the second temporal feature sequence can be further input into the prediction network so that the prediction network can predict the network state of the RTC system at at least one future time step based on the second temporal feature sequence.

[0104] In one embodiment shown, please continue to see Figure 5In the case where the aforementioned temporal feature extraction network may specifically include a TCN network and an attention network, the hidden feature sequence output by the attention network can be further input into the prediction network so that the prediction network can predict the network state of the RTC system at at least one future time step based on the hidden feature sequence.

[0105] Specifically, the aforementioned at least one future time step refers to the next one or more time steps following the last time step contained in the aforementioned first time-series feature sequence and the aforementioned second time-series feature sequence (these two sequences contain the same time steps).

[0106] For example, when constructing the aforementioned temporal feature matrix, if the dynamic sliding window mechanism mentioned above is used, then the aforementioned at least one future time step can specifically refer to the next one or more time steps of that time window. In this case, the second temporal feature sequence can be further input into the prediction network, which then predicts the network state of the RTC system at the next one or more time steps of that time window based on the second temporal feature sequence.

[0107] In some embodiments, please continue to see Figure 5 The aforementioned neural network may further include a feature aggregation layer; this feature aggregation layer may be located between the aforementioned TCN network and the aforementioned prediction network. This feature aggregation layer is specifically used to perform global feature aggregation on the input feature sequence, aggregating the input feature sequence into a fixed-length feature vector.

[0108] In this scenario, when further inputting the aforementioned hidden feature sequence into the prediction network, and having the prediction network predict the network state of the RTC system at at least one future time step based on the hidden feature sequence, the hidden feature sequence can first be input into a feature aggregation layer. This layer aggregates the individual hidden feature vectors contained in the hidden feature sequence into a fixed-length feature vector, thereby eliminating the time-step dimension of the hidden feature sequence and transforming it into a data format suitable for processing by the prediction network. Then, this fixed-length feature vector is further input into the prediction network, which then predicts the network state of the RTC system at at least one future time step based on this fixed-length feature vector.

[0109] It should be noted that the specific form of the above feature aggregation layer is not specifically limited in this specification; for example, in one embodiment, the above feature aggregation layer may be a pooling layer; the pooling layer may specifically adopt a global average pooling method to aggregate the various hidden feature vectors contained in the hidden feature sequence into a feature vector of fixed length (i.e., global temporal features).

[0110] In some embodiments, please continue to see Figure 5 The aforementioned prediction network may specifically include at least one fully connected layer and at least one output layer.

[0111] The fully connected layer can be located between the feature aggregation layer and at least one output layer, and is used to integrate features from a fixed-length input vector.

[0112] For example, a fully connected layer can combine and interact with the fixed-length feature vectors output by the pooling layer through a nonlinear transformation (ReLU) to learn a more discriminative abstract feature representation for the downstream output layer, providing a unified and efficient feature representation for subsequent prediction tasks.

[0113] It should be noted that in practical applications, the number of fully connected layers in the prediction network described above can be flexibly designed based on specific needs. For example, as Figure 5 As shown, the aforementioned prediction network can specifically include a first fully connected layer with 32 neurons and a second fully connected layer with 16 neurons. In this way, the second fully connected layer can compress the feature dimension of the feature representation output by the first fully connected layer from 32 to 16, thereby refining the information, retaining the most essential discriminative information, and reducing the number of parameters to prevent overfitting and improve the model's generalization ability.

[0114] In this case, the fixed-length feature vector can be further input into at least one fully connected layer to perform a nonlinear transformation on the fixed-length feature vector to generate a target feature vector for predicting the network state. Then, the target feature vector can be further input into at least one output layer to predict the network state of the RTC system at at least one future time step based on the target feature vector, and output the predicted network state.

[0115] In some embodiments, please continue to see Figure 5 The network state predicted by the prediction network can specifically include network state indicators in multiple dimensions that can reflect the network state of the RTC system; correspondingly, the prediction network can specifically include multiple output layers that correspond one-to-one with the above-mentioned multiple dimensions of network state indicators.

[0116] In this case, the target feature vector can be input into the multiple output layers respectively, so that the multiple output layers can predict the network state index of the RTC system at at least one future time step based on the target feature vector, and output the network state index to obtain the network state index of the RTC system in multiple dimensions at at least one future time step.

[0117] It should be noted that the network status indicators mentioned above can include any form of indicator that can reflect the network status of the RTC system, and will not be specifically limited in this specification.

[0118] In some embodiments, the network state indicators of the aforementioned multiple dimensions, which serve as the final prediction result of the network state, may specifically include: A first network state indicator for representing the congestion state of an RTC system; wherein, the first network state indicator may specifically include a first probability value for representing network congestion in the RTC system.

[0119] For example, such as Figure 5 As shown, the first output layer among the above multiple output layers can specifically be a sigmoid output layer with 1 neuron. The first network state index can be a congestion probability (with a value between 0 and 1) output by the first output layer. In practical applications, if the probability is greater than a threshold (e.g., 0.5), it can be determined that there is congestion; otherwise, it can be determined that there is no congestion.

[0120] A second network status indicator is used to represent the packet loss type of the RTC system; wherein, the second network status indicator may specifically include a second probability value for indicating that the RTC system does not have packet loss, a third probability value for indicating that the RTC system has random packet loss, and a fourth probability value for indicating that the RTC system has packet loss.

[0121] For example, such as Figure 5 As shown, the second output layer among the above multiple output layers can specifically be a softmax output layer with 3 neurons. The state index of the second network can be the three probability values ​​(between 0 and 1) output by the second output layer, representing "no packet loss", "random packet loss" and "congestion packet loss". In practical applications, the packet loss type corresponding to the largest probability value among these three types can be taken as the final packet loss type.

[0122] A third network status indicator used to represent the packet loss rate of an RTC system; wherein, the third network status indicator may specifically include a first packet loss rate used to represent random packet loss in the RTC system, and a second packet loss rate used to represent network congestion packet loss in the RTC system.

[0123] For example, such as Figure 5 As shown, the third output layer among the above multiple output layers can specifically be a sigmoid output layer with 2 neurons. The third network state index can be the two packet loss rate values ​​output by the third output layer (both normalized to the range of 0-100%), representing the random packet loss rate and the network congestion packet loss rate, respectively.

[0124] Once the aforementioned multiple output layers have predicted the congestion state, packet loss type, and packet loss rate of the RTC system at at least one future time step, these parameters can be further fed back in real-time to the downstream key modules of the sending end; for example, Figure 3 The core modules, such as bandwidth estimation, rate allocator, and hARQ (Hybrid Automatic Repeat reQuest), enable these downstream critical modules to anticipate potential network fluctuations, thereby ensuring that the RTC system can maintain stable service quality even in complex network environments.

[0125] In some embodiments, a tag generator may also be deployed in the RTC system described above; wherein, the tag generator may be used to add network state tags to the first time-series feature sequence input to the TCN network based on the transmission state data corresponding to the real-time communication data packets generated at least one future time step (e.g., the next one or more time steps of the time window), and generate a confidence score corresponding to the network state tag.

[0126] In this case, the first time-series feature sequence can be further input into the tag generator, so that the tag generator can add network status tags to the first time-series feature sequence based on the transmission status data corresponding to the real-time communication data packets generated at at least one future time step, and generate a confidence score corresponding to the network status tags.

[0127] The aforementioned label generator is the logic for automatically labeling the first temporal feature sequence. In practical applications, it can be flexibly designed based on specific needs.

[0128] In some embodiments, when the tag generator automatically tags the first time-series feature sequence, it may specifically calculate multiple sets of dynamic benchmark values ​​for determining the real network state of the RTC system at the at least one future time step based on the transmission status data corresponding to the real-time communication data packets generated at the at least one future time step; the multiple sets of dynamic benchmark values ​​may be benchmark values ​​that are calculated and updated in real time by the tag generator.

[0129] Then, these multiple sets of dynamic baseline values ​​can be matched with multiple preset marking rules in sequence; Finally, network state labels can be added to the first time-series feature sequence based on the matching results of the multiple sets of dynamic benchmark values ​​and the multiple sets of labeling rules.

[0130] For example, in some embodiments, taking the network state of the RTC system predicted by the prediction network at at least one future time step as including the congestion state, packet loss type, and packet loss rate of the RTC system at at least one future time step as an example, the label generator needs to automatically add three types of labels corresponding to the congestion state, packet loss type, and packet loss rate to the above-mentioned first time-series feature sequence.

[0131] At this point, the multiple sets of dynamic baseline values ​​calculated by the aforementioned tag generator can specifically include: First set of dynamic benchmark values: P75 quantile of RTT deviation; For example, for each data packet within the aforementioned time window, the difference between the RTT of that data packet and the minimum RTT within the window (i.e., the basic RTT deviation, where the minimum RTT represents the basic link delay and the deviation only reflects the additional delay caused by congestion) can be calculated. Then, the P75 quantile of the RTT deviation values ​​of each data packet within the time window can be taken (to filter out extreme fluctuations caused by random packet loss and focus on the delay trend of most data packets).

[0132] The second set of dynamic benchmark values: RTT deviation coefficient of variation (CV); For example, based on the aforementioned "basic RTT deviation", the ratio of the standard deviation of the RTT deviation value within the time window to the P50 quantile (the P50 quantile replaces the mean to avoid interference from extreme values) can be calculated to reflect the difference in queuing delay of different data packets, i.e., transmission stability.

[0133] The third set of dynamic baseline values: the proportion of time-series anomalies; For example, the impact of transmission interval fluctuations can be calibrated by using "theoretical reception time = (current packet sequence number - first packet sequence number) × average transmission interval within the window", calculating the proportion of data packets with "(actual reception time - theoretical reception time) / average transmission interval > 0.5", and using this indicator to reflect the timing disorder caused by congestion (regardless of packet loss type).

[0134] The third set of dynamic baseline values: the coefficient of variation of the packet loss interval.

[0135] For example, the coefficient of variation of the packet loss interval can be the ratio of the standard deviation of the packet loss interval sequence to the mean of the packet loss interval sequence. This feature quantifies the stability of the timing of packet loss events; that is, it focuses on the fluctuation characteristics of when packet loss occurs, rather than the packet loss rate itself.

[0136] On the one hand, the labeling rules corresponding to the congestion state labels can include: Rule 1 (Delay Cumulative): RTT deviation P75 quantile > 30% of the smallest RTT within the time window; Rule 2 (Transmission Instability): RTT deviation coefficient of variation (CV) > 0.3; Rule 3 (Temporal Disorder): The proportion of temporal anomalies > 10%; At this point, based on the matching results of the multiple sets of dynamic benchmark values ​​and the multiple sets of labeling rules, the final determination logic for adding congestion state labels to the first time-series feature sequence may include: If the multiple sets of dynamic baseline values ​​match at least two of the multiple sets of labeling rules, add a first label (e.g., 1) to the first time-series feature sequence to indicate the congestion state.

[0137] If the multiple sets of dynamic baseline values ​​match less than two of the multiple sets of labeling rules, a second label (e.g., 0) representing a non-congestion state is added to the first time-series feature sequence.

[0138] Secondly, the tagging rules corresponding to packet loss type labels may include: Rule 4: If the packet loss rate is ≤0.1%, add a third tag to indicate no packet loss.

[0139] Rule 5: If the first label has already been added, and the packet loss rate is >3%; and the number of consecutive packet losses is ≥2; directly add the fourth label indicating congestion and packet loss. Rule 6: If a second tag has already been added, and the jitter is greater than 15ms, and the number of consecutive packet losses is less than or equal to 1, then directly add a fifth tag indicating the presence of random packet loss.

[0140] Thirdly, the labeling rules corresponding to the packet loss rate label may include: Rule 7: If a fourth label has already been added, take the actual packet loss rate as the sixth label corresponding to the congestion packet loss rate.

[0141] Rule 8: If a fifth tag has already been added, take the actual packet loss rate as the seventh tag corresponding to the random packet loss rate.

[0142] It should be emphasized that the rules and labeling logic shown above are only illustrative. In practical applications, the rules and labeling logic can be flexibly defined based on specific needs.

[0143] In some embodiments, when generating a confidence score corresponding to a network state label added to the first temporal feature sequence, the label generator may specifically count the number of labeling rules that match the multiple sets of dynamic benchmark values; and then determine the confidence score corresponding to the network state label based on the counted number of labeling rules that match the multiple sets of dynamic benchmark values.

[0144] For example, in practical applications, certain quantification rules can be used to quantify the number of labeling rules that match the aforementioned multiple sets of dynamic benchmark values ​​into confidence scores. Specific quantification rules will not be detailed in this specification.

[0145] In some embodiments, a network evaluator for evaluating the prediction accuracy of the neural network may also be deployed in the real-time communication system described above. In this case, a time-series feature sequence sample can be constructed based on the first time-series feature sequence, the network state label added to the first time-series feature sequence, and the confidence score corresponding to the network state label. The time-series feature sequence sample and the network state predicted by the prediction network are then input into the network evaluator. The network evaluator evaluates the prediction accuracy of the network state predicted by the prediction network by matching the network state label contained in the time-series feature sequence sample with the network state predicted by the prediction network.

[0146] The network evaluator evaluates the accuracy of the predicted network state by matching the network state labels contained in the temporal feature sequence samples with the network state predicted by the prediction network. The specific evaluation logic is not described in detail in this specification. In practical applications, it can be flexibly designed based on the requirements.

[0147] For example, metrics for evaluating prediction accuracy can be defined based on requirements, and the accuracy of the prediction network's predicted network state can be evaluated based on the specific values ​​of these metrics.

[0148] In some embodiments, the RTC described above can also deploy a training manager for incremental training of the neural network. Specifically, when the triggering conditions for incremental training of the neural network are met, the training manager can select high-quality temporal feature sequence samples from the temporal feature sequence samples input to the network evaluator as training samples for incremental training of the neural network. Incremental training refers to the ability of the neural network, after initial training and deployment, to continuously and safely update its parameters using newly generated training samples without interrupting service.

[0149] In this case, when the triggering condition for incremental training of the neural network is met, the training manager can, in response to the triggering condition, select a set of time-series feature sequence samples from all time-series feature sequence samples input to the network evaluator that contain the aforementioned confidence score greater than a preset threshold; of course, in practical applications, other methods can also be used to select high-quality time-series feature sequence samples.

[0150] Then, the neural network can be incrementally trained based on this set of time-series feature sequence samples.

[0151] In some embodiments, the triggering condition may specifically include the decrease in the prediction accuracy of the predicted network predicted by the network evaluator reaching a preset threshold.

[0152] In this way, a closed loop can be formed between the training manager and the network evaluator. When the prediction accuracy of the network predicted by the network evaluator drops to a certain level, the training manager is automatically triggered to select a high-quality set of time-series feature sequence samples from all the time-series feature sequence samples input to the network evaluator and continue to train the neural network, thus forming a positive cycle.

[0153] Of course, in practical applications, the triggering conditions mentioned above can include not only the triggering condition that the decrease in the prediction accuracy of the prediction network predicted by the network evaluator reaches a preset threshold, but also other triggering conditions. For example, the triggering conditions may also include, for instance, the number of samples in the time-series feature sequence sample set with confidence scores greater than the preset threshold has accumulated to a certain number, or the time interval since the last incremental training has reached a certain length, and so on.

[0154] It should be noted that, whether it is the initial training or incremental training of the above-mentioned neural network, it is usually necessary to design a task loss for the task of predicting the network state mentioned above and select an appropriate loss function. In this specification, the details of how to deal with task loss and how to select an appropriate loss function will not be described in detail. Those skilled in the art can flexibly design and select based on specific needs.

[0155] For example, taking the network state of the RTC system predicted by the prediction network at at least one future time step, including the congestion state, packet loss type, and packet loss rate of the RTC system at at least one future time step, the task of predicting the congestion state and packet loss type can be understood as a classification task. For classification tasks, cross-entropy loss can be used, and the cross-entropy loss function can be used to describe the loss of the task of predicting the congestion state and packet loss type. On the other hand, the task of predicting the packet loss rate can be understood as a regression task. For regression tasks, mean squared error loss can be used, and the mean squared error loss function can be used to describe the loss of the task of predicting the packet loss rate.

[0156] Figure 7 This is a schematic structural diagram of an electronic device provided in an exemplary embodiment. For example... Figure 7As shown, device 700 mainly consists of a communication interface 702, a user interface 704, a processor 706, and a data storage 708. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 710. The communication interface 702 enables device 700 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 702 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 702 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 702 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 702 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces.

[0157] User interface 704 includes receiving user input and providing output to the user. Therefore, user interface 704 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 704 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 704 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 700 may support remote access from other devices via communication interface 702 or another physical interface (not shown). User interface 704 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 704 may also be configured as a display device for rendering or displaying text fragments.

[0158] Processor 706 may contain one or more general-purpose processors and / or special-purpose processors.

[0159] Data storage 708 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 706. Data storage 708 may include removable and non-removable components.

[0160] Processor 706 is capable of executing program instructions 718 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 708 to perform the various functions described herein. Data storage 708 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 700, enable device 700 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Processor 706 executing program instructions 718 may result in processor 706 using data 712.

[0161] For example, program instructions 718 may include an operating system 722 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 700 and one or more applications 720 (e.g., a browser, social application, or game application). Similarly, data 712 may include operating system data 716 and application data 714. Operating system data 716 is primarily accessible to the operating system 722, while application data 714 is primarily accessible to one or more applications 720. Application data 714 may reside in a file system visible or hidden from the user of device 700.

[0162] Application 720 can communicate with operating system 722 through one or more application programming interfaces (APIs). These APIs help application 720 read and / or write application data 714, transmit or receive information via communication interface 702, receive or display information on user interface 704, etc.

[0163] In some terminology, application 720 may be simply referred to as "app". Furthermore, application 720 can be downloaded to device 700 through one or more online app stores or app markets. However, applications can also be installed on device 700 in other ways, such as through a web browser or a physical interface on device 700 (e.g., a USB port).

[0164] Please refer to Figure 8 This specification also proposes a network state prediction device based on a neural network, which can be applied to applications such as... Figure 7 The device shown implements the technical solution of this specification. The neural network includes a temporal feature extraction network and a prediction network built based on TCN; the device may include: The acquisition module 801 acquires transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence, extracts timing features from the transmission status data, and constructs a first timing feature sequence corresponding to the real-time communication data packet sequence based on the extracted timing features; wherein, the real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending end and receiving end in the real-time communication system during real-time communication. The generation module 802 inputs the first temporal feature sequence into the feature extraction network, so that the feature extraction network performs a causal dilated convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence. The prediction module 803 further inputs the second time-series feature sequence into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence.

[0165] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0166] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0167] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0168] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0169] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0170] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0171] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0172] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0173] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0174] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0175] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. A network state prediction method based on a neural network, applied to a real-time communication system; the neural network includes a temporal feature extraction network constructed based on TCN and a prediction network; the method includes: The transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence is obtained, time-series features are extracted from the transmission status data, and a first time-series feature sequence corresponding to the real-time communication data packet sequence is constructed based on the extracted time-series features; wherein, the real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending end and receiving end in the real-time communication system during real-time communication. The first temporal feature sequence is input into the feature extraction network, so that the feature extraction network performs a causal dilated convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence. The second time-series feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence.

2. The method as described in claim 1, wherein the timing features include performance indicators calculated based on the transmission status data that can describe the network status of the real-time communication system; Extracting temporal features from the transmission status data, and constructing an initial first temporal feature sequence based on the extracted temporal features, including: The performance index set corresponding to each real-time communication data packet contained in the real-time communication data packet sequence is calculated based on the transmission status data; wherein, the performance index set includes performance indexes of multiple dimensions calculated based on the transmission status data; A multi-dimensional time-series feature vector corresponding to each real-time communication data packet is constructed based on the set of performance indicators corresponding to each real-time communication data packet contained in the real-time communication data packet sequence; Based on the multi-dimensional temporal feature vectors corresponding to each real-time communication data packet within a preset time window contained in the real-time communication data packet sequence, a temporal feature matrix corresponding to the time window is further constructed as the first temporal feature sequence; wherein, the time window is a time window that slides according to a preset time step length, with the latest time step generated by the real-time communication as the endpoint. The second time-series feature sequence is further input into the prediction network so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence, including: The second time-series feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at the next or multiple time steps of the time window based on the second time-series feature sequence.

3. The method of claim 2, wherein the timing features further include enhanced performance indicators obtained by further calculating the performance indicators.

4. The method as described in claim 3, wherein the performance metrics calculated based on the transmission status data include: A first type of performance indicator that can reflect the network transmission characteristics of the real-time communication system; A second type of performance indicator that can reflect the network latency characteristics of the real-time communication system; A third type of performance indicator that can reflect the network packet loss characteristics of the real-time communication system; A fourth type of performance indicator that can reflect the network stability characteristics of the real-time communication system; The enhanced performance index, calculated based on the performance index, includes: A fifth type of performance index is obtained by further calculation of multiple of the aforementioned performance indicators; A sixth type of performance index that can reflect the time-series characteristics of any of the aforementioned performance indices through further calculations. A seventh type of performance index is obtained by further calculation of any of the aforementioned performance indices, which can reflect the statistical characteristics of the performance indices.

5. The method as described in claim 2, wherein the temporal feature extraction network comprises a TCN network and an attention network; The first temporal feature sequence is input into the feature extraction network, so that the feature extraction network performs convolution operations related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence. The first temporal feature sequence is input into the TCN network, so that the TCN network performs a causal dilation convolution operation related to the TCN on the first temporal feature sequence to generate a second temporal feature sequence; The second temporal feature sequence is further input into the attention network, so that the attention network performs attention calculations on each multidimensional temporal feature vector contained in the first temporal feature sequence and other multidimensional temporal feature vectors contained in the first temporal feature sequence to obtain hidden feature vectors corresponding to each multidimensional temporal feature vector; and, based on the hidden feature vectors corresponding to each multidimensional temporal feature vector, a hidden feature sequence is further generated.

6. The method of claim 5, wherein the second time-series feature sequence is further input into the prediction network, so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the second time-series feature sequence, comprising: The hidden feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the hidden feature sequence.

7. The method of claim 6, wherein the neural network further comprises a feature aggregation layer; The hidden feature sequence is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the hidden feature sequence, including: The hidden feature sequence is input into the feature aggregation layer, so that the feature aggregation layer aggregates the hidden feature vectors contained in the hidden feature sequence into a feature vector of fixed length. The fixed-length feature vector is further input into the prediction network so that the prediction network can predict the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector.

8. The method of claim 7, wherein the prediction network comprises at least one fully connected layer and at least one output layer; The fixed-length feature vector is further input into the prediction network so that the prediction network predicts the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector, including: The fixed-length feature vector is further input into the at least one fully connected layer, so that the at least one fully connected layer can predict the network state of the real-time communication system at at least one future time step based on the fixed-length feature vector, including: The fixed-length feature vector is further input into the at least one fully connected layer, so that the at least one fully connected layer performs a nonlinear transformation on the fixed-length feature vector to generate a target feature vector. The target feature vector is further input to the at least one output layer, so that the at least one output layer can predict the network state of the real-time communication system at at least one future time step based on the target feature vector, and output the network state.

9. The method of claim 8, wherein the network state includes multiple dimensions of network state indicators that can reflect the network state of the real-time communication system; the prediction network includes multiple output layers that correspond one-to-one with the multiple dimensions of network state indicators; The target feature vector is further input into the at least one output layer, so that the at least one output layer predicts the network state of the real-time communication system at at least one future time step based on the target feature vector, and outputs the network state, including: The target feature vector is input to the plurality of output layers respectively, so that the plurality of output layers predict the network state indicators of the real-time communication system at at least one future time step based on the target feature vector, and output the network state indicators to obtain the network state indicators of the real-time communication system in multiple dimensions at at least one future time step.

10. The method of claim 9, wherein the multiple dimensions of network state indicators include: A first network status index is used to represent the congestion state of the real-time communication system; wherein, the first network status index includes a first probability value for representing network congestion in the real-time communication system. A second network status index is used to indicate the packet loss type of the real-time communication system; wherein the second network status index includes a second probability value indicating that the real-time communication system has no packet loss, a third probability value indicating that the real-time communication system has random packet loss, and a fourth probability value indicating that the real-time communication system has packet loss. A third network status indicator for representing the packet loss rate of the real-time communication system; wherein the third network status indicator includes a first packet loss rate for representing random packet loss in the real-time communication system, and a second packet loss rate for representing network congestion packet loss in the real-time communication system.

11. The method of claim 1, wherein the neural network is deployed at the transmitting end of the real-time communication system.

12. The method of claim 9, wherein a tag generator is further deployed in the real-time communication system; wherein, The tag generator is used to add network status tags to the first time-series feature sequence based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, and generate a confidence score corresponding to the network status tags. The method further includes: The first time-series feature sequence is further input into the tag generator, so that the tag generator adds network status tags to the first time-series feature sequence based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, and generates a confidence score corresponding to the network status tags.

13. The method of claim 12, wherein based on transmission status data corresponding to real-time communication data packets generated in the next or multiple time steps of the time window, a network status label is added to the first time-series feature sequence, and a confidence score corresponding to the network status label is generated, comprising: Based on the transmission status data corresponding to the real-time communication data packets generated in the next or multiple time steps of the time window, multiple sets of dynamic reference values ​​are calculated to determine the real network status of the real-time communication system in the next or multiple time steps of the time window. The multiple sets of dynamic reference values ​​are matched sequentially with multiple sets of preset marking rules; Based on the matching results of the multiple sets of dynamic benchmark values ​​and the multiple sets of labeling rules, network state labels are added to the first time-series feature sequence; Generate a confidence score corresponding to the network state label, including: Count the number of marking rules that match the multiple sets of dynamic benchmark values; Based on the number of labeling rules that match the multiple sets of dynamic benchmark values, the confidence score corresponding to the network state label is determined.

14. The method of claim 13, wherein the real-time communication system further comprises a network evaluator for evaluating the prediction accuracy of the neural network; The method further includes: A time-series feature sequence sample is constructed based on the first time-series feature sequence, the network state label added to the first time-series feature sequence, and the confidence score corresponding to the network state label; The temporal feature sequence samples and the network state predicted by the prediction network are input into the network evaluator, so that the network evaluator can evaluate the prediction accuracy of the network state predicted by the prediction network by matching the network state label contained in the temporal feature sequence samples with the network state predicted by the prediction network.

15. The method of claim 14, wherein the real-time communication system further deploys a training manager for incrementally training the neural network; The method further includes: In response to the fulfillment of the triggering condition for incremental training of the neural network, a set of time-series feature sequence samples containing a confidence score greater than a preset threshold is selected from the time-series feature sequence samples input to the network evaluator. The neural network is incrementally trained based on the time-series feature sequence sample set; wherein the triggering condition includes the decrease in the prediction accuracy predicted by the network evaluator reaching a preset threshold.

16. A neural network, comprising: A temporal feature extraction network based on TCN is used to perform causal dilated convolution operations related to the TCN on the input first temporal feature sequence to generate a second temporal feature sequence. The first temporal feature sequence is a first temporal feature sequence corresponding to the real-time communication data packet sequence, constructed based on temporal features extracted from the transmission status data corresponding to each real-time communication data packet contained in the real-time communication data packet sequence. The real-time communication data packet sequence is a sequence composed of real-time communication data packets generated by the sending and receiving ends in the real-time communication system during real-time communication. A prediction network is used to predict the network state of the real-time communication system at at least one future time step based on the second temporal feature sequence as input.

17. An electronic device, characterized in that, include: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-15 by executing the executable instructions.

18. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-15.

19. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-15.