Generative video transmission method and device

By optimizing generative video transmission through fast UDP internet connections and the QUIC/COPA algorithm, the instability of video transmission in low-bandwidth and weak network environments is solved, enabling priority transmission of critical data and efficient bandwidth utilization, thereby improving the stability and quality of video transmission.

CN122069256APending Publication Date: 2026-05-19CHINA TELECOM CORP LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TELECOM CORP LTD
Filing Date
2026-04-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In low-bandwidth, weak network environments, traditional generative video transmission suffers from poor stability, including head-of-line congestion, low retransmission efficiency, inability to perceive service priorities, and inadequate congestion control, leading to video transmission stuttering, latency, and quality degradation.

Method used

Using a fast UDP internet connection, the transmission strategy for multiple network abstraction layer units is determined based on the network bandwidth information fed back by the video data receiver. The transmission priorities of key frames, edge features, and potential features are distinguished. Combined with the QUIC protocol and COPA congestion control algorithm, the transmission rate and priority queue management are adjusted in real time to ensure that critical data is transmitted first.

Benefits of technology

It improves the stability and bandwidth utilization of video data transmission, reduces latency and packet loss rate, ensures the continuity and quality of video reconstruction, and adapts to low-bandwidth, high-latency network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069256A_ABST
    Figure CN122069256A_ABST
Patent Text Reader

Abstract

The invention discloses a generative video transmission method and device. The method comprises the steps that video data to be transmitted are acquired, the video data to be transmitted are converted into multiple pieces of network abstraction layer unit data, and the types of the network abstraction layer unit data comprise key frames, edge features, audios and potential features; determining a transmission strategy of multiple pieces of network abstraction layer unit data according to network bandwidth information fed back by a receiving end of the to-be-transmitted video data; the multiple pieces of network abstraction layer unit data are transmitted to a receiving end according to the transmission strategy in a rapid UDP internet connection mode, and the transmission priority of the network abstraction layer unit data is determined according to the types of the multiple pieces of network abstraction layer unit data. The method provided by the invention at least solves the technical problem of poor stability of generated video transmission in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video transmission technology, and more specifically, to a generative video transmission method and apparatus. Background Technology

[0002] Generative video compression technology significantly reduces the bandwidth required for video transmission by encoding video into multiple data streams, such as keyframes, edge features, and latent features, and then using large-scale generative models for video reconstruction at the decoding end. This provides a foundation for applications such as remote monitoring, real-time communication, and remote control. However, in low-bandwidth, weak network environments, network characteristics are complex. Typical low-bandwidth, weak network environments include satellite communication (high latency, Doppler effect), mountain / desert networks (insufficient coverage, aging equipment), mobile edge networks (base station handover, mobile interference), and deep-sea communication (acoustic signal transmission, ultra-long distance). These networks typically exhibit the following weak network characteristics: high latency and high jitter, high packet loss rate, drastic bandwidth fluctuations, and asymmetric bandwidth: in low-bandwidth environments, uplink and downlink bandwidths are often asymmetrical. For example, in satellite communication, the uplink bandwidth is much smaller than the downlink bandwidth; downlink is better than uplink in communication links; and uplink limitations may also occur in shared network environments. Asymmetric bandwidth affects the quality of bidirectional communication, requiring separate optimization of uplink and downlink data in the transmission scheme. Rapid network state switching is also a concern. In the aforementioned low-bandwidth, weak-network environments, traditional transmission methods suffer from the following shortcomings: **Head-of-line congestion:** In traditional transmission mechanisms, a single packet loss can block all subsequent data packets, leading to severe stuttering in video transmission under high packet loss conditions and a significant reduction in real-time performance. **Low retransmission efficiency:** Traditional transmission mechanisms rely on timeouts and repeated acknowledgments, requiring long wait times before triggering retransmissions, resulting in decreased throughput and inefficient use of limited bandwidth. **Inability to perceive service priorities:** Traditional transmission mechanisms treat all data equally, failing to differentiate between data of varying importance, such as keyframes, edge features, and latent features in generative compression. Critical data may be delayed, affecting decoding continuity and quality. **Inadequate congestion control in weak networks:** Traditional congestion control algorithms (such as Cubic) grow slowly in high-latency, low-bandwidth environments and rely on packet loss-triggered congestion assessment, resulting in delayed responses and difficulty in quickly adjusting to drastic bandwidth changes, leading to low bandwidth utilization. **Network handover causing disconnections:** Frequent handovers in mobile or satellite networks can cause network connection drops, requiring re-handshakes to establish new connections, increasing latency and causing video transmission interruptions, particularly impacting real-time monitoring and remote control applications. The aforementioned problems result in poor transmission stability of generated video using traditional transmission mechanisms, which fails to meet user needs. Summary of the Invention

[0003] This application provides a generative video transmission method and apparatus to at least solve the technical problem of poor stability in generative video transmission in related technologies.

[0004] According to one aspect of the embodiments of this application, a generative video transmission method is provided, comprising: acquiring video data to be transmitted; converting the video data to be transmitted into multiple Network Abstraction Layer (NET) unit data, wherein the types of NET unit data include: keyframes, edge features, audio, and latent features; determining a transmission strategy for the multiple NET unit data based on network bandwidth information fed back by the receiving end of the video data to be transmitted; and transmitting the multiple NET unit data to the receiving end using a Fast UDP Internet connection according to the transmission strategy, wherein the transmission priority of the NET unit data is determined according to the types of the multiple NET unit data.

[0005] Optionally, a fast UDP internet connection is used to transmit multiple network abstraction layer unit data to the receiving end according to a transmission strategy, including: obtaining the transmission priority of the multiple network abstraction layer unit data respectively, wherein the transmission priority of keyframes, audio, edge features, and latent features decreases in that order; and transmitting the multiple network abstraction layer unit data to the receiving end in that order according to different transmission priorities, wherein before sending edge features and latent features, it is determined whether the keyframes corresponding to the edge features and latent features have been transmitted, and if the keyframes corresponding to the edge features and latent features have been transmitted, the transmission of edge features and latent features is determined.

[0006] Optionally, multiple Network Abstraction Layer (NAL) unit data are transmitted to the receiving end using a fast UDP Internet connection according to a transmission strategy. This includes: obtaining the round-trip time from the data sender to the receiver at a target time, and determining the queue delay of the transmission queue containing the NAL unit data at the target time based on the round-trip time at the target time and the minimum round-trip time within a preset period; reducing the transmission rate of the multiple NAL unit data if the rate of increase of the queue delay at the target time is greater than a preset rate of change; and increasing the transmission rate of the multiple NAL unit data by a preset step size if the queue delay at the target time is lower than a first delay threshold.

[0007] Optionally, the transmission strategy for multiple network abstraction layer unit data is determined based on the network bandwidth information fed back by the receiver, including: decomposing the potential features in the multiple network abstraction layer unit data into multiple potential features, wherein the multiple potential features include: high-resolution features, medium-resolution features, and low-resolution features, with the resolution of high-resolution features, medium-resolution features, and low-resolution features decreasing sequentially; receiving the bandwidth rate at the target time from the receiver, and determining the type of potential feature currently being transmitted in the transmission strategy for the multiple network abstraction layer unit data based on the rate range in which the bandwidth rate at the target time is located.

[0008] Optionally, the method further includes: obtaining the transmission result fed back by the receiving end and the video scene corresponding to the video data at the target time, wherein the transmission result includes: the actual bitrate of the video data to be transmitted received by the receiving end and the quality score of the video data to be transmitted received by the receiving end; if the actual bitrate of the video data to be transmitted received by the receiving end is lower than the estimated bitrate, adjusting the encoding bitrate of the video data to be transmitted, wherein the estimated bitrate is determined according to the bandwidth rate of the encoding end; if the video scene corresponding to the video data at the target time belongs to a scene-switching video scene or the video image frame change rate corresponding to the video data at the target time is higher than the preset rate, interrupting the current encoding process and generating a new keyframe; if the quality score of the video data to be transmitted received by the receiving end is lower than the preset score, generating a new keyframe and increasing the encoding bitrate of the video data to be transmitted.

[0009] Optionally, the method further includes: obtaining the data packet transmission duration threshold for each type of network abstraction layer unit data; detecting whether the data packet transmission duration of each type of network abstraction layer unit data exceeds the corresponding transmission duration threshold according to a preset period; and discarding data packets that exceed the corresponding transmission duration threshold in sequence according to the priority of the data packets.

[0010] Optionally, the video data to be transmitted is converted into multiple network abstraction layer unit (NAL) data, including: when the NAL data is greater than a preset maximum transmission unit, the NAL data is fragmented to obtain multiple fragments, wherein the unique identifier of the multiple fragments is the same as the unique identifier of the NAL data, and the integrity of the multiple fragments is verified after the receiving end receives the multiple fragments.

[0011] According to another aspect of the embodiments of this application, a generative video transmission apparatus is also provided, comprising: an acquisition module, configured to acquire video data to be transmitted and convert the video data to be transmitted into multiple Network Abstraction Layer (NET) unit data, wherein the types of NET unit data include: keyframes, edge features, audio, and latent features; a determination module, configured to determine a transmission strategy for the multiple NET unit data based on network bandwidth information fed back by the receiving end of the video data to be transmitted; and a transmission module, configured to transmit the multiple NET unit data to the receiving end using a Fast UDP Internet connection according to the transmission strategy, wherein the transmission priority of the NET unit data is determined according to the types of the multiple NET unit data.

[0012] According to another aspect of the embodiments of this application, a computer device is also provided, including: a memory and a processor, wherein the memory is used to store program instructions; and the processor, connected to the memory, is used to execute the above-described generative video transmission method.

[0013] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the above-described generative video transmission method.

[0014] In this embodiment, the method involves acquiring video data to be transmitted and converting it into multiple Network Abstraction Layer (NAL) unit data. The types of NAL unit data include keyframes, edge features, audio, and latent features. A transmission strategy for the multiple NAL unit data is determined based on network bandwidth information fed back from the receiving end of the video data. The multiple NAL unit data are then transmitted to the receiving end using a Fast UDP Internet connection according to the transmission strategy. The transmission priority of the NAL unit data is determined based on its type. By using a UDP Internet connection to transmit multiple NAL unit data and determining different transmission priorities according to their types, the method effectively utilizes network bandwidth and ensures stable video data transmission. This improves the stability of video data transmission and solves the technical problem of poor stability in generative video transmission in related technologies. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0016] Figure 1 This is a hardware structure block diagram of a computer target terminal for implementing a generative video transmission method according to an embodiment of this application;

[0017] Figure 2 This is a flowchart of a generative video transmission method according to an embodiment of this application;

[0018] Figure 3 This is a schematic diagram of the structure of a generative video transmission system according to an embodiment of this application;

[0019] Figure 4 This is a schematic diagram of priority scheduling in a generative video transmission process according to an embodiment of this application;

[0020] Figure 5 This is a schematic diagram of bandwidth feedback in a generative video transmission process according to an embodiment of this application;

[0021] Figure 6 A flowchart of timeout handling in a generative video transmission process according to an embodiment of this application;

[0022] Figure 7A schematic diagram of the architecture of a network damage simulation tool according to an embodiment of this application;

[0023] Figure 8 A structural diagram of a generative video transmission device according to an embodiment of this application. Detailed Implementation

[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] The information collected in this application embodiment is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, and necessary confidentiality measures have been taken. It does not violate public order and good morals, and provides corresponding operation entry points for users to choose to authorize or reject the automated decision results. If the user chooses to reject, the process will proceed to the expert decision-making process.

[0027] The technical terms used in this application are explained as follows:

[0028] Generative video compression: a video coding technique that significantly reduces video transmission bandwidth requirements by decomposing video into structured data such as keyframes, edge features, and latent features, and then reconstructing it using a large video generation model at the decoding end.

[0029] NALU (Network Abstraction Layer Unit): In generative video compression, it is the smallest decodeable unit. Each NALU corresponds to a complete piece of data, such as a keyframe, a sequence of edge features, or a sequence of latent features. If the NALU is not delivered completely, the generative model cannot correctly reconstruct the corresponding video content.

[0030] QUIC (Quick UDP Internet Connections): A transport protocol based on UDP (User Datagram Protocol) that achieves reliable transmission at the application layer. It supports features such as multiplexing, headless blocking, fast connection establishment, and connection migration, and integrates TLS 1.3 encryption to improve transmission efficiency and security.

[0031] COPA (Congestion Control using Packet-Pair Arrival): A congestion control algorithm based on network latency variations. It predicts network congestion by monitoring round-trip time (RTT) changes in real time, making it more sensitive than packet loss-based detection and particularly suitable for low-bandwidth, high-latency environments.

[0032] Multi-scale latent features: At the encoder, the original latent features are decomposed into high, medium, and low resolution features through a multi-scale downsampling network. Based on the current network bandwidth conditions, latent features of appropriate resolution are adaptively selected for transmission to achieve a balance between quality and bandwidth.

[0033] To address the problems existing in related technologies, this application provides a generative video transmission method, which can be run on... Figure 1 The computer target terminal shown will be explained below.

[0034] The generative video transmission method embodiments provided in this application can be executed in a mobile target terminal, a computer target terminal, or a similar computing device. Figure 1 A hardware block diagram of a computer target terminal for implementing a generative video transmission method is shown. Figure 1As shown, the computer target terminal 10 may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions connected via wired and / or wireless networks. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, the computer target terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0035] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be implemented wholly or partially as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element in the computer target terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor target terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the generative video transmission method in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned generative video transmission method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer target terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0037] The transmission module 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer target terminal 10. In one example, the transmission module 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0038] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer target terminal 10.

[0039] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer target terminal shown may include hardware components (including circuitry), software components (including computer code stored on a computer-readable medium), or a combination of both hardware and software components. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computer target terminal.

[0040] In the above operating environment, this application provides an embodiment of a generative video transmission method. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0041] Figure 2 This is a flowchart of a generative video transmission method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0042] Step S202: Obtain the video data to be transmitted and convert it into multiple network abstraction layer unit data. The types of network abstraction layer unit data include: keyframes, edge features, audio, and latent features.

[0043] Step S204: Determine the transmission strategy for multiple network abstraction layer unit data based on the network bandwidth information fed back by the receiving end of the video data to be transmitted.

[0044] Step S206: Using a fast UDP Internet connection, multiple Network Abstraction Layer (NAT) unit data are transmitted to the receiving end according to a transmission strategy. The transmission priority of the NAT unit data is determined based on the types of NAT unit data.

[0045] In step S206, a fast UDP internet connection is used to transmit multiple network abstraction layer unit data to the receiving end. This fast UDP internet connection supports parallel transmission of multiple data streams within a single connection, with each stream operating independently, effectively eliminating the head-of-line blocking problem found in traditional transmission mechanisms. Simultaneously, the fast UDP internet connection reduces initial connection latency, making it particularly suitable for bandwidth-constrained and high-latency scenarios. Furthermore, the fast UDP internet connection maintains an uninterrupted connection even when the terminal's network address or port changes. This feature is highly suitable for scenarios involving mobile movement or cross-network handover, preventing transmission interruptions.

[0046] Through steps S202 to S206 above, the raw speech data to be processed is received; short-time Fourier transforms of different scales are performed on the raw speech data to obtain frequency domain signals of different scales; the frequency domain signals of different scales are processed to obtain mask matrices of different scales, and the mask matrices of different scales are fused to obtain a target mask matrix; the target mask matrix is ​​used to process the frequency domain signal corresponding to the raw speech data to obtain the processed frequency domain signal; the processed speech signal is determined based on the processed frequency domain signal. By converting the raw speech data into frequency domain signals of different scales and processing the frequency domain signals of different scales to obtain mask matrices of different scales, and using the mask matrix to process the frequency domain signal corresponding to the raw speech data to obtain the processed speech signal, the goal of quickly and accurately processing noise in the raw speech data is achieved, thereby improving the speech quality and solving the technical problem in related technologies where a large amount of noise in the received speech leads to poor speech quality. The following is a detailed explanation.

[0047] To better illustrate the generative video transmission method proposed in the embodiments of this application, a generative video transmission system is also proposed in the embodiments of this application, such as... Figure 3 As shown, it includes: a data sending end (encoding end), a transport layer, and a receiving end (decoding end). It can be understood that the encoding end includes: a video encoder and a network transmitter, and the decoding end includes: a video decoder and a network receiver.

[0048] In some embodiments of this application, the specific steps for transmitting multiple Network Abstraction Layer (NET) unit data to the receiving end using a fast UDP internet connection according to a transmission strategy include: obtaining the transmission priority of each NET data set, wherein the transmission priority of keyframes, audio, edge features, and latent features decreases sequentially; and transmitting the NET data set sequentially to the receiving end according to different transmission priorities. Specifically, before transmitting edge features and latent features, it is determined whether the keyframes corresponding to the edge features and latent features have been transmitted. If the keyframes corresponding to the edge features and latent features have been transmitted, the edge features and latent features are then transmitted. This method ensures the integrity and availability of video reconstruction. By transmitting keyframes first, followed by edge features and latent features, it ensures that the decoding end has complete keyframes as a basis for reconstruction when it receives feature data, avoiding the inability to parse feature data, video reconstruction failure, or screen tearing and stuttering due to missing keyframes. To enhance robustness in low-bandwidth weak network environments characterized by high packet loss, high latency, and bandwidth fluctuations, priority is given to ensuring the reliable arrival of core data such as key frames. Even if subsequent feature data is partially lost or delayed, the receiving end can still complete basic video reconstruction based on key frames, maintaining basic availability for communication and monitoring services. Bandwidth utilization is optimized, adapting to dynamic networks by transmitting key frames, audio, edge features, and potential features in sequence. When bandwidth is limited or suddenly drops, important data is sent first, reducing the preemption of core link resources by non-critical data and improving the efficiency of effective data transmission under limited bandwidth. Head-of-line congestion and transmission latency are reduced. Based on QUIC / Fast UDP, combined with priority scheduling and key frame pre-positioning strategies, head-of-line congestion in traditional transmission mechanisms is avoided, while ensuring priority transmission of key data. This significantly reduces end-to-end latency for services such as real-time video and remote control, improving real-time performance. After key frame transmission is complete, corresponding edge features and potential features are then sent, enabling the decoding end's generation model to reconstruct video according to the correct timing and complete dependencies, improving the continuity, clarity, and quality stability of the reconstructed image.

[0049] like Figure 4 As shown, data is prioritized according to its importance in the decoding process: keyframe data has the highest priority, followed by audio data, then edge feature data, and finally latent feature data. The transport layer maintains multiple priority queues, with all keyframe data entering the highest priority queue. The send scheduler always prioritizes sending data packets from the highest priority queue, ensuring that keyframes arrive at the decoder first. Before sending edge and latent feature data, the system checks whether the corresponding keyframe has been fully transmitted. If a keyframe has not been fully transmitted, feature data dependent on that keyframe is delayed to prevent the decoder from being blocked due to missing keyframes.

[0050] In some embodiments of this application, the specific steps for converting the video data to be transmitted into multiple network abstraction layer unit (NET) data are as follows: the video data to be transmitted is converted into multiple NET data; if the NET data is greater than a preset maximum transmission unit (MTU), the NET data is fragmented to obtain multiple fragments; wherein the unique identifier of the multiple fragments is the same as the unique identifier of the NET data; and after the receiving end receives the multiple fragments, the integrity of the multiple fragments is verified.

[0051] Specifically, each NALU (Network Abstraction Unit) data is sent with a unique identifier, sequence number, and integrity checksum. When a NALU exceeds the network MTU, it is fragmented by the transport layer. The receiver reassembles the fragments according to the sequence number and verifies their integrity, ensuring atomic delivery of the entire NALU. Data is treated differently based on its importance. Reliable transmission (e.g., acknowledgment and retransmission mechanisms) is implemented for critical frames and base layer NALUs, while a best-effort transmission strategy is used for enhancement layer and non-critical NALUs to achieve a balance between reliability and efficiency. The receiver maintains a NALU reassembly buffer and supports out-of-order reception. When a fragment or missing NALU is detected, a retransmission is immediately requested via the sequence number or a timeout retransmission mechanism is triggered to quickly recover from transmission errors.

[0052] It should be noted that the base layer NALU includes at least: keyframes and low-resolution latent features, while the enhancement layer NALU includes at least: edge features, medium-resolution latent features, and high-resolution latent features.

[0053] In some embodiments of this application, the specific steps for transmitting multiple Network Abstraction Layer (NAL) unit data to the receiving end using a fast UDP Internet connection according to a transmission strategy are as follows: The round-trip time (RTT) from the data sender to the receiver at the target time is obtained, and the queue delay of the transmission queue containing the NAL unit data at the target time is determined based on the RTT at the target time and the minimum RTT within a preset period; if the rate of increase of the queue delay at the target time is greater than a preset rate of change, the transmission rate of the multiple NAL unit data is reduced; if the queue delay at the target time is lower than a first delay threshold, the transmission rate of the multiple NAL unit data is increased by a preset step size. By obtaining the RTT in real time and calculating the queue delay in conjunction with the minimum RTT, the current network link congestion level and buffer backlog can be accurately reflected, enabling refined perception of the network status and providing a reliable basis for transmission rate adjustment. When the queue delay increases too rapidly, actively reducing the transmission rate can effectively alleviate network link congestion, avoid a large accumulation of data packets, increased forwarding delay, and increased packet loss rate, and is particularly suitable for weak network environments with drastic bandwidth fluctuations and high latency. When the network is idle, bandwidth is fully utilized to improve transmission efficiency. When the queue latency is below the threshold and the network is idle or lightly loaded, the transmission rate is gradually increased according to a preset step size. This allows for full utilization of available bandwidth without causing congestion, improving the transmission efficiency and real-time performance of video data. Smooth and stable rate adjustment avoids drastic jitter. A gradual adjustment method based on the latency change rate and a fixed step size ensures smooth changes in the transmission rate, preventing sharp increases and decreases, guaranteeing smooth video transmission and continuous decoding, and improving user experience. Adaptable to weak network, high latency, and high jitter environments, improving transmission robustness. Compared to traditional packet loss-based congestion control, the scheme provided in this application uses latency for congestion judgment, resulting in a more timely and sensitive response. It can maintain good transmission performance in scenarios with long round-trip time (RTT) and unstable bandwidth, such as satellite communication and mobile edge networks.

[0054] In practical applications, the transport layer employs a delay-based COPA algorithm for congestion control, monitoring round-trip time (RTT) changes in real time and predicting network congestion through rapid increases in queue latency. When available bandwidth rapidly decreases from sufficient to limited, COPA can detect bandwidth contraction in time by increasing RTT before packet loss occurs. When network bandwidth drops sharply, the COPA algorithm quickly converges to a new safe sending rate within several RTT cycles, avoiding continuous over-sending; when bandwidth recovers, the sending rate is smoothly increased to fully utilize available bandwidth. Low-bandwidth optimization: In low-bandwidth environments, a more conservative congestion window growth strategy is adopted to avoid exacerbating congestion; when bandwidth recovers slightly, the growth rate is gradually increased to improve utilization.

[0055] To better control the transmission rate, in this embodiment, the transmission scheduler distributes data packets evenly according to the transmission budget, avoiding congestion and packet loss caused by sudden transmissions. The transmission rate is dynamically adjusted based on the available bandwidth estimated by the congestion control algorithm, matching the transmission rate with network capacity and avoiding prolonged backlogs or idle periods. Service priority is incorporated into the transmission rate control, giving high-priority data packets more transmission opportunities within the same budget, ensuring that critical data is transmitted first.

[0056] In some embodiments of this application, the specific steps for determining the transmission strategy of multiple network abstraction layer unit data based on the network bandwidth information fed back by the receiving end are as follows: The potential features in the multiple network abstraction layer unit data are decomposed into various potential features, including: high-resolution features, medium-resolution features, and low-resolution features, with the resolution decreasing sequentially from high-resolution to medium-resolution. The bandwidth rate at the target time is received from the receiving end, and the type of potential feature currently being transmitted in the transmission strategy of the multiple network abstraction layer unit data is determined based on the rate range in which the bandwidth rate at the target time falls. This scheme achieves a dynamic balance between video quality and transmission stability by classifying potential features by resolution and adaptively selecting the transmission type based on real-time bandwidth. When bandwidth is sufficient, high-resolution features are transmitted to improve reconstruction quality; when bandwidth is insufficient, only low-resolution features are transmitted to ensure smooth transmission. This significantly improves the transmission adaptability and system robustness in weak network and bandwidth fluctuation scenarios.

[0057] Specifically, the encoder uses a multi-scale downsampling network to decompose the original latent features into three scales: high-resolution (preserving more details), medium-resolution (balancing details and semantics), and low-resolution (primarily preserving semantics), providing alternatives for subsequent adaptive transmission. Based on real-time bandwidth feedback from the receiver or transport layer, when bandwidth is sufficient (e.g., >1Mbps), high-resolution latent features are prioritized for transmission to achieve higher reconstruction quality; when bandwidth is moderate (e.g., 500kbps~1Mbps), medium-resolution features are transmitted; and when bandwidth is insufficient (e.g., <500kbps), only low-resolution features are transmitted to prioritize transmission continuity and basic image quality. When a sudden bandwidth drop is detected (e.g., from 5Mbps to 100kbps), the encoder immediately switches from high-scale to low-scale features to avoid excessive data volume causing congestion or loss; when bandwidth recovers, it quickly switches back to a higher scale to improve image quality. The switching can be completed within seconds, maintaining smooth video playback. Based on a bandwidth trend prediction mechanism, the scale is preemptively downscaled when bandwidth shows a continuous downward trend to avoid the drastic impact of sudden bandwidth drops; when bandwidth shows a clear upward trend, the scale can be preemptively increased to accelerate bandwidth utilization. During scale switching, a gradual strategy should be adopted as much as possible, such as frame-by-frame mixing or gradual enhancement, to avoid image quality jitter during the switching process. At the same time, transmission continuity should be prioritized when bandwidth changes drastically. The decoding end calls the corresponding upsampling network to reconstruct the video based on the characteristic scale of the received signal.

[0058] In one optional approach, the transmission results fed back from the receiving end and the video scene corresponding to the video data at the target time are obtained. The transmission results include: the actual bitrate of the video data to be transmitted received by the receiving end and the quality score of the video data to be transmitted received by the receiving end. If the actual bitrate of the video data to be transmitted received by the receiving end is lower than the estimated bitrate, the encoding bitrate of the video data to be transmitted is adjusted, wherein the estimated bitrate is determined based on the bandwidth rate of the encoding end. If the video scene corresponding to the video data at the target time belongs to a scene-changing video scene or the frame change rate of the video image data corresponding to the video data at the target time is higher than a preset rate, the current encoding process is interrupted, and a new keyframe is generated. If the quality score of the video data to be transmitted received by the receiving end is lower than the preset score, a new keyframe is generated, and the encoding bitrate of the video data to be transmitted is increased.

[0059] To enable the encoding end to adaptively adjust the transmission strategy based on network conditions, this invention designs various information feedback mechanisms, such as... Figure 5As shown, the receiving end determines the currently available bandwidth and notifies the encoding end via a dedicated feedback packet. The encoding end dynamically selects appropriate scaled potential features for transmission based on the feedback bandwidth estimate, achieving adaptive transmission rate. The receiving end statistically analyzes the actual received data rate and feeds it back to the encoding end for calibrating the bandwidth estimate. When the actual received bitrate deviates significantly from the estimated value, the encoding end can adjust its transmission strategy based on the feedback. When the receiving end detects a drastic change in the video frame (such as scene switching, rapid motion, etc.), it notifies the encoding end via a feedback signal. Upon receiving the signal, the encoding end immediately generates and sends a new keyframe to ensure the decoding end can correctly reconstruct the new frame content. The receiving end monitors the quality of the decoded image; if it finds severely distorted or incorrect reconstruction results, it notifies the encoding end via a feedback signal. Upon receiving the feedback, the encoding end can take measures, such as regenerating keyframes, improving the quality of subsequent frames, or requesting retransmission of key data, to restore or improve video quality.

[0060] In latency-sensitive scenarios such as real-time video transmission (e.g., video calls), strictly ensuring all data retransmissions can lead to latency accumulation, impacting real-time performance. This application proposes a real-time performance guarantee mechanism based on timeout packet loss: It acquires the data packet transmission duration threshold for each type of network abstraction layer unit (NET) data; checks at preset intervals whether the data packet transmission duration for each NET NET unit exceeds the corresponding transmission duration threshold; and discards data packets exceeding the corresponding transmission duration threshold sequentially according to their priority. For example... Figure 6 As shown, this includes setting transmission deadlines (transmission duration thresholds) for different types of data packets, such as 300ms for keyframe data packets and 200ms for feature data packets. The sender records the transmission time of each data packet and periodically checks whether unacknowledged data packets have exceeded their deadlines. Packets that have exceeded their deadlines are discarded according to priority from low to high, with priority given to retaining high-priority data such as keyframes. For discarded outdated data segments, the sender skips the time window and continues sending subsequent data, avoiding latency accumulation and ensuring timely transmission of subsequent critical data. This method, by setting transmission duration thresholds for different NALU data and selectively discarding outdated messages according to priority, avoids latency accumulation caused by retransmissions while prioritizing the transmission of critical data, balancing real-time transmission and video reconstruction availability, significantly improving transmission smoothness and system robustness in latency-sensitive scenarios.

[0061] To verify the performance of the generative video transmission system provided in this application embodiment, this application embodiment also provides a network impairment simulation method, including: an application-layer UDP proxy program running between the encoding and decoding ends, used to introduce network impairment effects during packet forwarding; the UDP proxy program provides a configuration interface, which can independently set network parameters such as round-trip time, packet loss rate, bandwidth limit, and jitter; it supports configuring different network impairment parameters for the uplink and downlink respectively to simulate asymmetric bandwidth; during forwarding, according to the current configuration, corresponding delay, dropping, or rate limiting operations are applied to the transmitted data packets, specifically, such as... Figure 7 As shown, the process includes: deploying a UDP proxy tool between the encoding and decoding ends to insert network impairment effects; configuring initial network environment parameters, such as a latency of 200ms, a packet loss rate of 5%, a bandwidth limit of 100kbps, and jitter of 50ms; setting different network impairment parameters for uplink and downlink respectively to simulate an asymmetric network; and conducting tests in the following scenarios: Scenario 1: Rapid bandwidth reduction – dynamically reducing the bandwidth from 1Mbps to 100kbps during operation and observing the response of the bitrate adaptive mechanism; Scenario 2: Bandwidth fluctuation – simulating periodic bandwidth fluctuations between 1Mbps and 5Mbps to verify the rapid adjustment of COPA congestion control to bandwidth changes; and Scenario 3: Sudden latency changes – dynamically increasing the latency from 50ms to 300ms to evaluate the system's adaptability to sudden latency changes. During the tests, video quality, transmission latency, packet loss rate, and other indicators are monitored in real time to evaluate the system's overall performance in dynamic weak network environments.

[0062] To better illustrate the generative video transmission method in the embodiments of this application, a specific embodiment will be used for explanation below.

[0063] This application also provides a generative video compression intelligent transmission method for low-bandwidth scenarios, including the following steps: Step 1, transmitting video data using the UDP-based QUIC protocol; Step 2, dividing the video data to be transmitted into blocks according to the NALU structure of generative video compression, dividing the video data into the smallest decodeable units such as keyframes, edge features, and latent features, and fragmenting, sending, and reassembling each NALU at the transport layer to ensure atomic delivery of each NALU; Step 3, using a delay-based COPA algorithm for congestion control, dynamically adjusting the transmission rate according to the real-time monitored round-trip time (RTT) changes; Step 4, processing the data according to the decoding dependencies of the generative video. The process involves several steps: Step 5: Prioritizing keyframes, audio, edge features, and latent features, assigning different priorities and prioritizing the transmission of high-priority data; Step 6: For data packets with high real-time requirements, setting a transmission deadline and implementing a timeout discard policy, discarding packets that are not acknowledged before the deadline; Step 7: Adaptively selecting latent features of different resolutions for transmission based on real-time network bandwidth information from the receiver; Step 8: Dynamically adjusting encoding parameters and transmission strategies at the encoding end based on the actual bitrate, video scene changes, and decoding quality information from the receiver; and Step 9: Utilizing the multiplexing, fast handshake, and connection migration functions of the QUIC protocol to achieve reliable transmission of video data along the end-to-end path.

[0064] In step 2 above, a unique identifier and sequence number are generated for each NALU, and fragmentation is performed when the NALU size exceeds the MTU. The receiving end reassembles the fragments according to the sequence number and verifies their integrity. If a missing or incorrect fragment is detected, a retransmission is requested. Reliable transmission (e.g., acknowledgment and retransmission) is used for keyframes and base layer NALUs, while best-effort transmission is used for enhancement layer NALUs. In step 4 above, priority scheduling includes: setting keyframe data as the highest priority; setting audio data as high priority; setting edge feature data as medium priority; setting latent feature data as low priority; always prioritizing the transmission of high-priority data during the transmission scheduling process, and checking whether the corresponding keyframe has been transmitted before sending feature data. In step 3 above, the COPA algorithm specifically involves: real-time monitoring of RTT changes and calculation of queue delay; reducing the transmission rate when a significant increase in RTT or queue delay is detected; smoothly increasing the transmission rate when RTT recovers to a low level; and dynamically adjusting the incremental parameters of the COPA algorithm to achieve a balance between latency sensitivity and throughput. The multi-resolution feature selection in step 6 above includes: the encoder generating three potential features—high-resolution, medium-resolution, and low-resolution—through a multi-scale downsampling network; based on real-time bandwidth feedback, high-resolution features are selected for transmission when bandwidth is sufficient, medium-resolution features are selected when bandwidth is moderate, and low-resolution features are selected when bandwidth is limited; and smooth switching between different resolution features occurs when bandwidth conditions change. The multi-type feedback in step 7 above includes: the receiver estimating the currently available bandwidth based on a congestion control algorithm and feeding it back to the encoder; the server calculating the actual received bitrate and feeding it back to the encoder; and the receiver sending a trigger signal to the encoder when it detects a video scene change or a decrease in decoding quality, prompting the encoder to generate new keyframes or improve the encoding quality of subsequent frames.

[0065] Figure 8 A generative video transmission apparatus is shown, the apparatus comprising:

[0066] The acquisition module 80 is used to acquire the video data to be transmitted and convert the video data to be transmitted into multiple network abstraction layer unit data. The types of network abstraction layer unit data include: keyframes, edge features, audio and latent features.

[0067] The determination module 82 is used to determine the transmission strategy of multiple network abstraction layer unit data based on the network bandwidth information fed back by the receiving end of the video data to be transmitted.

[0068] The transmission module 84 is used to transmit multiple Network Abstraction Layer (NAL) unit data to the receiving end using a fast UDP Internet connection according to a transmission strategy. The transmission priority of the NAL unit data is determined based on the types of NAL unit data.

[0069] The aforementioned generative video transmission device acquires the video data to be transmitted, converts it into multiple Network Abstraction Layer (NAL) unit data, including keyframes, edge features, audio, and latent features. It then determines a transmission strategy for these NAL unit data based on network bandwidth information from the receiving end of the video data. Finally, it transmits the NAL unit data to the receiving end using a Fast UDP Internet connection according to the transmission strategy. The transmission priority of the NAL unit data is determined based on its type. By using a UDP Internet connection to transmit multiple NAL unit data and determining different transmission priorities according to their types, the device effectively utilizes network bandwidth and ensures stable video data transmission. This improves the stability of video data transmission and solves the technical problem of poor stability in generative video transmission in related technologies.

[0070] The transmission module 84 includes: a transmission submodule, used to obtain the transmission priorities of multiple network abstraction layer unit data respectively, wherein the transmission priorities of keyframes, audio, edge features, and latent features decrease in order; and to send multiple network abstraction layer unit data to the receiving end in order according to different transmission priorities, wherein, before sending edge features and latent features, it is determined whether the keyframes corresponding to the edge features and latent features have been transmitted; if the keyframes corresponding to the edge features and latent features have been transmitted, it is determined to transmit the edge features and latent features.

[0071] The transmission module 84 includes an adjustment submodule, which is used to obtain the round-trip delay from the data sender to the receiver at the target time, and determine the queue delay of the transmission queue where the network abstraction layer unit data is located at the target time based on the round-trip delay at the target time and the minimum round-trip delay within a preset period; if the rate of increase of the queue delay at the target time is greater than the preset rate of change, the transmission rate of multiple network abstraction layer unit data is reduced; if the queue delay at the target time is lower than the first delay threshold, the transmission rate of multiple network abstraction layer unit data is increased according to a preset step size.

[0072] The determination module 82 includes: a determination submodule, used to determine the transmission strategy of multiple network abstraction layer unit data based on the network bandwidth information fed back by the receiving end, including: decomposing the potential features in the multiple network abstraction layer unit data into multiple potential features, wherein the multiple potential features include: high-resolution features, medium-resolution features and low-resolution features, with the resolution of high-resolution features, medium-resolution features and low-resolution features decreasing sequentially; receiving the bandwidth rate at the target time from the receiving end, and determining the type of potential feature currently being transmitted in the transmission strategy of the multiple network abstraction layer unit data according to the rate range in which the bandwidth rate at the target time is located.

[0073] The determination submodule includes: an evaluation unit, used to obtain the transmission results fed back from the receiving end and the video scene corresponding to the video data at the target time. The transmission results include: the actual bitrate and quality score of the video data to be transmitted received by the receiving end; if the actual bitrate of the video data to be transmitted received by the receiving end is lower than the estimated bitrate, the encoding bitrate of the video data to be transmitted is adjusted, where the estimated bitrate is determined based on the bandwidth rate of the encoding end; if the video scene corresponding to the video data at the target time belongs to a scene-changing video scene or the frame change rate of the video image data corresponding to the video data at the target time is higher than a preset rate, the current encoding process is interrupted, and a new keyframe is generated; if the quality score of the video data to be transmitted received by the receiving end is lower than the preset score, a new keyframe is generated, and the encoding bitrate of the video data to be transmitted is increased.

[0074] The aforementioned generative video transmission device further includes: a timeout processing submodule, used to obtain the data packet transmission duration threshold for each type of network abstraction layer unit data; detect whether the data packet transmission duration of each type of network abstraction layer unit data exceeds the corresponding transmission duration threshold according to a preset period; and discard data packets that exceed the corresponding transmission duration threshold in sequence according to the priority of the data packets.

[0075] The aforementioned generative video transmission device further includes: a fragmentation submodule, used to fragment the network abstraction layer unit data when the network abstraction layer unit data is greater than the preset maximum transmission unit, to obtain multiple fragments, wherein the unique identifier of the multiple fragments is the same as the unique identifier of the network abstraction layer unit data, and wherein the receiver performs integrity verification on the multiple fragments after receiving the multiple fragments.

[0076] It should be noted that, Figure 8 The generative video transmission device shown is used to perform Figure 2 The generative video transmission method shown above also applies to this generative video transmission device, and will not be repeated here.

[0077] This application also provides a computer device, including: a memory and a processor, wherein the memory is used to store program instructions; and the processor, connected to the memory, is used to execute the above-described generative video transmission method.

[0078] This application also provides a computer program product, including computer instructions that, when executed by a processor, implement the steps of the generative video transmission method in this application.

[0079] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0080] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0081] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0082] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0083] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0084] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0085] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A generative video transmission method, characterized in that, include: The video data to be transmitted is acquired and converted into multiple network abstraction layer unit data, wherein the types of network abstraction layer unit data include: keyframes, edge features, audio and latent features; The transmission strategy for the data of the multiple network abstraction layer units is determined based on the network bandwidth information fed back by the receiving end of the video data to be transmitted. The multiple network abstraction layer unit data are transmitted to the receiving end using a fast UDP internet connection according to the transmission strategy, wherein the transmission priority of the network abstraction layer unit data is determined according to the types of the multiple network abstraction layer unit data.

2. The method according to claim 1, characterized in that, The multiple network abstraction layer unit data are transmitted to the receiving end using a fast UDP internet connection according to the transmission strategy, including: The transmission priorities of the multiple network abstraction layer unit data are obtained respectively, wherein the transmission priorities of the keyframe, audio, edge features, and latent features decrease in that order; The data of the multiple network abstraction layer units are sent to the receiving end in sequence according to different transmission priorities. Before sending the edge features and the potential features, it is determined whether the key frames corresponding to the edge features and the potential features have been transmitted. If the key frames corresponding to the edge features and the potential features have been transmitted, it is determined to transmit the edge features and the potential features.

3. The method according to claim 1, characterized in that, The multiple network abstraction layer unit data are transmitted to the receiving end using a fast UDP internet connection according to the transmission strategy, including: The round-trip time from the data sender to the receiver at the target time is obtained, and the queue delay of the transmission queue where the network abstraction layer unit data is located at the target time is determined based on the round-trip time at the target time and the minimum round-trip time within a preset period. If the rate of increase of queue delay at the target time is greater than a preset rate of change, reduce the data transmission rate of the multiple network abstraction layer units. If the queue delay at the target time is lower than the first delay threshold, the transmission rate of the data of the plurality of network abstraction layer units is increased by a preset step size.

4. The method according to claim 1, characterized in that, The transmission strategy for the data of the multiple network abstraction layer units is determined based on the network bandwidth information fed back by the receiving end, including: The latent features in the data of the multiple network abstraction layer units are decomposed into multiple latent features, wherein the multiple latent features include: high-resolution features, medium-resolution features and low-resolution features, and the resolution of the high-resolution features, medium-resolution features and low-resolution features decreases in that order; The receiver receives the bandwidth rate at the target time and determines the potential feature types of the current transmission in the transmission strategy of the multiple network abstraction layer unit data based on the rate range in which the bandwidth rate at the target time is located.

5. The method according to claim 4, characterized in that, The method further includes: The transmission result fed back by the receiving end and the video scene corresponding to the video data at the target time are obtained, wherein the transmission result includes: the actual bit rate of the video data to be transmitted received by the receiving end and the quality score of the video data to be transmitted received by the receiving end. If the actual bitrate of the video data to be transmitted received at the receiving end is lower than the estimated bitrate, the encoding bitrate of the video data to be transmitted is adjusted, wherein the estimated bitrate is determined based on the bandwidth rate of the data sending end. If the video scene corresponding to the video data at the target time belongs to a scene-switching video scene or the video image frame change rate corresponding to the video data at the target time is higher than the preset rate, the current encoding process is interrupted and a new keyframe is generated. If the quality score of the video data to be transmitted received at the receiving end is lower than the preset score, a new keyframe is generated and the encoding bitrate of the video data to be transmitted is increased.

6. The method according to claim 2, characterized in that, The method further includes: Obtain the data packet transmission duration threshold for each type of network abstraction layer unit data; The data packet transmission time of each network abstraction layer unit is checked at a preset period to see if it exceeds the corresponding transmission time threshold. Data packets that exceed the corresponding transmission time threshold will be discarded in order of their priority.

7. The method according to claim 1, characterized in that, The video data to be transmitted is converted into multiple network abstraction layer unit data, including: When the network abstraction layer unit data is greater than a preset maximum transmission unit, the network abstraction layer unit data is fragmented to obtain multiple fragments. The unique identifier of the multiple fragments is the same as the unique identifier of the network abstraction layer unit data. After the receiving end receives the multiple fragments, it performs integrity verification on the multiple fragments.

8. A generative video transmission device, characterized in that, include: The acquisition module is used to acquire the video data to be transmitted and convert the video data to be transmitted into multiple network abstraction layer unit data, wherein the types of network abstraction layer unit data include: keyframes, edge features, audio and latent features; The determination module is used to determine the transmission strategy of the multiple network abstraction layer unit data based on the network bandwidth information fed back by the receiving end of the video data to be transmitted. The transmission module is used to transmit the plurality of network abstraction layer unit data to the receiving end using a fast UDP Internet connection according to the transmission strategy, wherein the transmission priority of the network abstraction layer unit data is determined according to the types of the plurality of network abstraction layer unit data.

9. A computer device, characterized in that, include: A memory and a processor, wherein the memory is used to store program instructions; The processor, connected to the memory, is used to execute the generative video transmission method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the generative video transmission method according to any one of claims 1 to 7.