Audio and video signal fusion transmission method and device, equipment and medium

By combining cross-modal attention mechanism and quantum random number generator with dynamic encryption strategy, the problems of low transmission efficiency, weak anti-interference ability and low security in existing audio and video transmission technology are solved, realizing high-efficiency, low-latency and quantum attack resistant audio and video transmission, which can meet the needs of different scenarios.

CN120956967APending Publication Date: 2025-11-14JIE XUN TECH (GUANGZHOU) CO LTD

Patent Information

Application Number
CN202511139459.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing audio and video transmission technologies suffer from low transmission efficiency, weak anti-interference capabilities, low security, and poor adaptability to various scenarios, failing to meet the demands of high efficiency, low latency, and resistance to quantum attacks in the 5G/8K era.

Method used

Semantic features of audio and video are extracted through a cross-modal attention mechanism to generate a joint feature matrix. Resources are dynamically allocated in combination with scene coefficients. A quantum random number generator is used to generate a physically unclonable session master key. The key is encapsulated using the Kyber-768 algorithm, and the encryption strategy is dynamically adjusted. The transmission is carried out in combination with GPMI Type-B and QUIC protocols.

Benefits of technology

It improves transmission efficiency, enhances anti-interference capabilities and security, optimizes scene adaptability, and meets the needs of real-time interactive scenarios such as remote surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956967A_ABST
    Figure CN120956967A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of audio and video fusion transmission, and discloses an audio and video signal fusion transmission method and device, equipment and a medium, and the method comprises the steps: extracting 512-dimensional semantic feature vectors of an audio and a video, and generating a joint feature matrix through a cross-modal attention mechanism; calculating a fusion feature flow based on a coefficient alpha output by the scene classification model; a quantum random number generator is used for generating a session master key Kmain, the session master key Kmain is packaged into an encrypted session key Kenc through a Kyber-768 algorithm, and frame-level dynamic encryption is carried out on the fusion feature flow; and selecting a GPMI Type-B wired protocol or wireless transmission based on QUIC according to the transmission environment. According to the invention, the transmission efficiency, the anti-interference capability and the anti-quantum-attack security are improved, and multi-scene requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio and video fusion transmission, and specifically relates to a method, apparatus, device and medium for audio and video signal fusion transmission. Background Technology

[0002] Existing audio and video transmission technologies mainly rely on physical layer multiplexing, such as Time Division Multiplexing (TDM) in HDMI (High-Definition Multimedia Interface) and micro-packet transmission or independent stream compression encapsulation (such as H.265 video encoding + AAC audio encoding) in DisplayPort. These technologies suffer from the following key technical bottlenecks: Low transmission efficiency: Physical layer multiplexing only realizes signal superposition and does not consider the correlation between audio and video content, resulting in bandwidth utilization generally being less than 70%. In weak network environments (such as packet loss rate > 5%), key information (such as voice and action frames) is easily lost, causing audio and video desynchronization. Limited anti-interference capabilities: It relies on a single transmission protocol (wired such as HDMI 2.1 supports 48Gbps, wireless such as Wi-Fi 6E supports 9.6Gbps), cannot dynamically adapt to link status, and the wireless transmission latency is usually >50ms, making it difficult to meet the needs of real-time interactive scenarios (such as remote surgery and drone control). Security protection is lagging behind: When using traditional encryption algorithms such as AES-256, the key cracking time can be shortened to minutes when facing attacks by the quantum computing Shor algorithm, and the key management relies on a centralized CA mechanism, which poses a risk of man-in-the-middle attacks. Lack of scenario adaptability: Fixed resource allocation strategies cannot match differentiated needs—for example, meeting scenarios require priority to ensure voice clarity (audio weight > video), while film and television scenarios require optimization of picture details (video weight > audio), resulting in resource waste or quality degradation.

[0003] Although existing technologies attempt to optimize performance through link aggregation (such as IEEE 802.3ad) or dynamic coding (such as AV1 adaptive bitrate), they have not made breakthroughs in semantic layer fusion and quantum-level security, and still cannot meet the transmission requirements of "high efficiency, low latency, and resistance to quantum attacks" in the 5G / 8K era. Summary of the Invention

[0004] This invention provides a method, apparatus, device, and medium for audio and video signal fusion transmission, aiming to solve the technical problems of low transmission efficiency, weak anti-interference ability, low security, and poor scene adaptability caused by the reliance on physical layer multiplexing in audio and video transmission.

[0005] To achieve the above-mentioned objective, the first aspect of the present invention provides a method for audio and video signal fusion transmission, comprising the following steps: The input audio signal undergoes Mel-frequency spectrum transformation and MFCC (Mel-Frequency Cepstral Coefficients) feature extraction. Temporal dependencies are captured using an LSTM (Long Short-Term Memory) network, generating a 512-dimensional audio semantic feature vector. Keyframe sampling is performed on the input video signal, and a visual Transformer model is used to extract scene, action, and facial features, generating a 512-dimensional video semantic feature vector. The similarity weight between the audio semantic feature vector and the video semantic feature vector is calculated through a cross-modal attention mechanism. Based on the similarity weight, the audio semantic feature vector and the video semantic feature vector are mapped to a shared semantic space to generate a joint audio-video feature matrix. The cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer, and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type. The scene classification model outputs scene coefficients α, and the audio and video joint feature matrix is ​​adjusted based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; A physically unclonable session master key K_main is generated by a quantum random number generator, which is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. The session master key K_main is encapsulated using the Kyber-768 algorithm to generate the encrypted session key K_enc; Frame-level dynamic encryption is performed on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission environment is determined as follows: If the transmission link is wired, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth ≥192Gbps. If the transmission link is wireless, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC (Quick UDP Internet Connections) protocol. Multi-link aggregation is achieved through SD-ARC technology, with an end-to-end latency ≤20ms. When the packet loss rate is within 20%, data integrity is restored through forward error correction.

[0006] Furthermore, the step of outputting scene coefficients α through the scene classification model includes: Acquire the speech activity detection results of the audio signal and the motion vector amplitude of the keyframes of the video signal; The speech activity detection results and the keyframe motion vector magnitudes are input into the scene classification model to obtain scene coefficients α; wherein, the scene classification model is trained based on a support vector machine classifier.

[0007] Furthermore, the similarity calculation layer employs a multi-head attention mechanism with 8 heads, each attention head having a feature dimension of 64. The similarity weight calculation formula is as follows:

[0008] Let be the similarity weight between the i-th audio feature query vector and the j-th video feature key vector; This is the vector of the i-th row of the audio feature query matrix; Let j be the vector of the j-th row of the video feature key matrix; =64 represents the feature dimension of a single attention head; Represents the dot product of vectors. This is a scaling factor to avoid gradient vanishing.

[0009] Furthermore, the GPMI (General Purpose Multimedia Interface) Type-B interface protocol uses four-channel differential signal transmission with a rate of 48Gbps per channel, is compatible with the USB Type-C physical interface, and supports reverse power supply.

[0010] Furthermore, the coding redundancy of the forward error correction is dynamically adjusted according to the packet loss rate, including: redundancy = 10% when packet loss rate ≤ 5%, redundancy = 20% when 5% < packet loss rate ≤ 15%, and redundancy = 30% when packet loss rate > 15%.

[0011] Furthermore, the dynamic weight allocation layer calculates the audio feature weights using the following nonlinear fusion formula. and video feature weights : in: For scene coefficients (as defined in claim 2); For audio signal-to-noise ratio, Peak signal-to-noise ratio (PSNR) of the video; For link packet loss rate, Video block error rate; For network jitter, For transmission delay, =30ms =100ms is the threshold; =0.05、 =0.03 is the signal quality gain coefficient. =0.1、 =0.08 is the link impairment attenuation coefficient. =0.02、 =0.015 is the real-time penalty coefficient.

[0012] Furthermore, the video semantic feature vector is extracted using the improved visual Transformer model, and its attention mask matrix M satisfies:

[0013] in: x, y are the image patch indices, and k=3 is the size of the local attention window; For image blocks and The intersection and union ratio; A mask value of 1 indicates strong attention, 0.2 indicates weak attention, and 0 indicates no attention.

[0014] A second aspect of the present invention provides an audio and video signal fusion transmission device, comprising: The audio semantic feature extraction module performs Mel-frequency conversion and MFCC feature extraction on the input audio signal, captures temporal dependencies through an LSTM network, and generates a 512-dimensional audio semantic feature vector; and, The video semantic feature extraction module is used to sample keyframes of the input video signal, and uses a visual Transformer model to extract scene, action and expression features, generating a 512-dimensional video semantic feature vector. A cross-modal joint module is used to calculate the similarity weight between the audio semantic feature vector and the video semantic feature vector through a cross-modal attention mechanism, and to map the audio semantic feature vector and the video semantic feature vector to a shared semantic space based on the similarity weight to generate an audio-video joint feature matrix; the cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type; The scene fusion module is used to output scene coefficients α through the scene classification model, and adjust the audio and video joint feature matrix based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; The generation module is used to generate a physically unclonable session master key K_main through a quantum random number generator. The quantum random number generator is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. The encapsulation module is used to encapsulate the session master key K_main using the Kyber-768 algorithm to generate the encrypted session key K_enc. The encryption module is used to perform frame-level dynamic encryption on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission determination module is used to determine the transmission environment. If the transmission link is a wired environment, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth ≥192Gbps. If the transmission link is a wireless environment, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, with an end-to-end latency ≤20ms. When the packet loss rate is within 20%, data integrity is restored through forward error correction.

[0015] A third aspect of the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any of the preceding claims.

[0017] Beneficial effects: The audio and video signal fusion transmission method, apparatus, device, and medium of the present invention have the following beneficial effects: Improving transmission efficiency: Existing technologies rely on physical layer multiplexing, resulting in bandwidth utilization rates generally below 70%, and critical information is easily lost in weak networks. This application achieves audio-video semantic layer fusion through a cross-modal attention mechanism, and dynamically allocates resources based on the scene coefficient α (e.g., prioritizing audio in conference scenarios and video in film and television scenarios), avoiding unnecessary bandwidth occupation and increasing bandwidth utilization to over 90%. Simultaneously, wireless transmission employs SD-ARC multi-link aggregation and dynamic FEC, enabling complete data recovery even with a packet loss rate of less than 20%, thus resolving the audio-visual desynchronization problem in weak networks.

[0018] Enhanced anti-interference capability: Existing technologies rely on a single protocol, with wireless latency typically exceeding 50ms, making it difficult to meet the needs of real-time scenarios. This application adopts a GPMI Type-B interface (total bandwidth ≥192Gbps) for wired environments and a QUIC-based BBRv3 algorithm for wireless environments, achieving end-to-end latency ≤20ms and supporting dynamic switching between wired and wireless modes to adapt to different link states, thus meeting the real-time interaction requirements of remote surgery, drone control, and other applications.

[0019] Enhanced security: Existing technologies employ traditional encryption methods such as AES-256, which are vulnerable to quantum computing attacks, and key management suffers from man-in-the-middle risks. This application generates a physically unclonable session master key (entropy value ≥ 128 bits) using QRNG based on laser phase noise, and encapsulates the key using a Kyber-768 post-quantum algorithm, resisting Shor's algorithm attacks. The key cracking time is extended to an order of magnitude unattainable by quantum computers, while avoiding the security vulnerabilities of centralized CA mechanisms.

[0020] Optimizing scenario adaptability: Existing technologies with fixed resource allocation cannot match differentiated needs. This application uses a scenario classification model to output a coefficient α, dynamically adjusting the audio-video fusion ratio to ensure that core information (such as conference audio and video footage) is transmitted preferentially in different scenarios, reducing resource waste and improving transmission quality. Attached Figure Description

[0021] Figure 1 A flowchart illustrating an embodiment of an audio / video signal fusion transmission method of the invention; Figure 2 This is a schematic block diagram of an audio and video signal fusion transmission device according to an embodiment of the invention; Figure 3 This is a schematic block diagram of a computer device according to an embodiment of the invention.

[0022] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0024] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, modules, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any modules and all combinations of one or more associated listed items.

[0025] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0026] Reference Figure 1 This invention provides an audio and video signal fusion transmission method, comprising the following steps: S1. Perform Mel-spectrum transformation and MFCC feature extraction on the input audio signal, capture the temporal dependence through an LSTM network, and generate an audio semantic feature vector with a dimension of 512.

[0027] This step is the process of extracting audio semantic features. This embodiment takes a remote surgical scenario as an example. The input audio is the doctor's real-time instruction (e.g., "Move the hemostat 3 mm to the left"). First, preprocessing is performed: the speech is converted into a digital signal using a 16kHz sampling rate. A Mel filter bank (40 filters) is used to generate a Mel spectrum (time frame × frequency bin = 128 × 64). Then, the MFCC algorithm extracts 13-dimensional basic features, which are expanded to 39-dimensional features by combining first-order and second-order differences to capture the time-varying characteristics of the speech (e.g., the tone change of "to the left"). The processed features are input into a 3-layer LSTM network (256 neurons per layer, Dropout ratio 0.2). The network filters background noise from surgical instruments through a forget gate and reinforces the temporal dependencies of key instructions such as "move" and "3 mm" through an input gate. Finally, a 512-dimensional audio semantic feature vector is output, accurately representing the semantic intent of the instruction.

[0028] S2. Perform keyframe sampling on the input video signal (sampling frequency 15fps), and use the visual Transformer model to extract scene, action and expression features to generate a 512-dimensional video semantic feature vector.

[0029] This step involves extracting semantic features from the video. Taking a remote surgical scene as an example, the input video is a shot of an abdominal surgery (8K resolution, 30fps). Keyframes are sampled at 15fps (1 frame every 66ms) to retain key information such as scalpel movement and tissue state. Each frame is processed using an improved Visual Transformer (ViT) model: the image is divided into 16×16 pixel blocks (32×32=1024 blocks in total). Features are extracted using a self-attention mechanism; for example, strong attention is given to the block containing the scalpel and adjacent tissue blocks, while distant background blocks are assigned a value of 0 to reduce unnecessary computation. Finally, a 512-dimensional video semantic feature vector is generated, containing the surgical scene (abdomen), actions (cutting), and subtle facial expressions (tissue traction and deformation).

[0030] S3. Calculate the similarity weight between the audio semantic feature vector and the video semantic feature vector through a cross-modal attention mechanism, and map the audio semantic feature vector and the video semantic feature vector to a shared semantic space based on the similarity weight to generate a joint audio-video feature matrix; the cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type.

[0031] This step is the cross-modal feature fusion process. Taking a remote surgical scenario as an example, the cross-modal attention mechanism is processed in three layers: Feature Alignment Layer: Audio commands are matched with corresponding video frames using timestamps (e.g., aligning the "move" voice with the frame where the scalpel begins to move), and the similarity sharpness is adjusted according to the adaptive temperature coefficient formula in the following embodiment (when audio and video synchronization is high, the temperature coefficient T=0.3, enhancing high similarity features). Similarity Calculation Layer: An 8-head attention mechanism (64 dimensions per head) is used to calculate the similarity weight between the audio query vector and the video key vector according to the similarity weight formula in subsequent embodiments (e.g., the weight W=0.85 for the "move" voice and the scalpel movement). Dynamic Weight Allocation Layer: Audio feature weights are calculated using the non-linear fusion formula in the following embodiment. and video feature weights For example, in the case of audio during surgery =30dB, video Calculate audio weights when the image quality is 25dB (slightly blurry). =0.52, Video Weight =0.48, generating a 1024×512 audio-video joint feature matrix. In this embodiment, audio and video belong to different modalities (auditory and visual), and their original feature vectors are in different semantic spaces (audio focuses on the temporal features of sound signals, while video focuses on the spatial and motion features of images). The shared semantic space is a unified high-dimensional space that can simultaneously represent the semantic information of audio and video, making cross-modal features comparable. Based on the above similarity weights, the audio semantic feature vectors and video semantic feature vectors are "projected" into the shared semantic space through a feature alignment layer and a dynamic weight allocation layer of the cross-modal attention mechanism. Specifically, high similarity weights (such as the weights of the audio "coughing sound" and the video "patient coughing action") strengthen the association representation of corresponding features in the shared space, while low similarity weights (such as background noise and static scenes) weaken the association, achieving semantic alignment of features from different modalities. After mapping through the shared semantic space, the semantic features of audio and video are integrated into a matrix (i.e., the audio-video joint feature matrix). The rows of this matrix can correspond to time steps or feature dimensions, while the columns contain the fused audio and video joint semantic information. It preserves the temporal dependencies of the audio (such as the order of voice commands) and the scene / action features of the video (such as the trajectory of the equipment), and reflects the semantic relationship between the two through similarity weights (such as "the synchronicity between voice commands and corresponding actions"), providing a foundation for subsequent scene adaptive fusion (based on scene coefficient α).

[0032] S4. Output scene coefficients α through the scene classification model, and adjust the audio and video joint feature matrix based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene.

[0033] This step is the core of achieving "scene-adaptive fusion" of audio and video features. It dynamically adjusts the fusion weights of audio and video features through scene classification, ensuring the fusion result matches the priority requirements of different scenes for audio and video. Specifically, it can be divided into two parts: "scene coefficient α generation" and "fusion feature stream calculation," which are explained below with reference to specific scenarios: Generation of scene coefficient α: Scene type recognition based on audio and video features The essence of the scene coefficient α is the weight ratio of the audio semantic feature vector in the fusion ((1-α) is the weight of the video semantic feature vector). Its value is automatically determined by the scene classification model based on the key features of audio and video. The core logic is to "infer the scene type through the objective features of audio and video, and then match the preset α range".

[0034] Input features: The input to the scene classification model includes two types of features: Audio characteristics: Voice activity detection (VAD) results, such as the percentage of speech signal (non-silent portion) within a continuous 500ms time window. For example, in a meeting scenario where participants speak continuously, the VAD result may reach 60%~80%; in a film or television scenario where speech is mostly intermittent dialogue, the VAD result may only be 20%~30%; in a monitoring scenario (such as remote surgery), where speech is discontinuous commands, the VAD result is approximately 40%~50%.

[0035] Video characteristics: Motion vector amplitude of keyframes, which is the average distance a pixel moves in a video frame (reflecting the dynamic range of the image). For example, in film and television scenes with frequent scene changes and intense camera movement, the motion vector amplitude may reach 20-30 pixels / frame; in meeting scenes, where the characters are mostly in a static sitting posture or with slight head movements, the motion vector amplitude is about 5-10 pixels / frame; in surveillance scenes (such as surgery), where the movement of instruments is moderate, the motion vector amplitude is about 15-20 pixels / frame.

[0036] Scene classification model: The model is trained based on Support Vector Machine (SVM). By learning the combined features of the VAD results and motion vector magnitudes, it categorizes scenes into types such as "meeting," "movie," and "surveillance," and outputs the corresponding α range. For example: When the input VAD=70% (high speech percentage) and motion vector=8 pixels / frame (low dynamic range), the model identifies it as a "conference scene" and outputs α=0.6~0.7; When the input VAD=25% (low speech percentage) and motion vector=25 pixels / frame (high dynamic range), the model identifies it as a "film and television scene" and outputs α=0.2~0.3; When the input VAD=45% (medium voice percentage) and motion vector=18 pixels / frame (medium dynamic range), the model identifies it as a "monitoring scene" and outputs α=0.5.

[0037] Calculation of fused feature streams: Alpha-based weighted fusion of audio and video features The fused feature flow is the result of a weighted sum of audio and video semantic feature vectors with weights of α. The core of the formula "fused feature flow = α•audio semantic feature vector + (1-α)•video semantic feature vector" is to "give higher-priority modal features a greater weight in the fusion result". The specific effect varies depending on the scenario, for example: Meeting scenario (α=0.6~0.7): The core requirement of a meeting is to clearly convey audio information (such as the speaker's viewpoints and instructions), with video serving only as an auxiliary tool (such as observing facial expressions). In this case, a larger value for α results in a higher weight for the audio semantic feature vector. For example, in a video conference, the audio semantic feature vector contains the audio semantics of "project deadlines postponed to Friday," while the video semantic feature vector contains the action features of participants nodding. During fusion, the audio weight accounts for 60% to 70%, ensuring that audio information is preferentially preserved during transmission. Even if the video is slightly compressed due to bandwidth limitations, it does not affect the transmission of core information.

[0038] Film and television scenes (α=0.2~0.3): The core requirement of film and television is to present high-quality visuals (such as high-definition picture quality and smooth action), with audio (such as background music and dialogue) serving the visuals. In this case, the value of α is relatively small, and the video semantic feature vector has a higher weight. For example, in a 4K movie, the video semantic feature vector contains the lighting and dynamic details of the explosion scene, while the audio semantic feature vector contains the explosion sound effects. During fusion, the video weight accounts for 70% to 80%, ensuring that the visual details are preserved first during transmission, and even if the audio is slightly lost, the viewing experience can still be guaranteed.

[0039] Monitoring scenario (α=0.5): In monitoring scenarios (such as remote surgery and security monitoring), audio and video information must be equally important (e.g., surgeon's instructions and instrument movements during surgery, and abnormal sounds and images in security monitoring). In this case, α=0.5, ensuring balanced audio and video weights. For example, in remote surgery, the audio semantic feature vector contains the instruction "move the hemostat to the left," while the video semantic feature vector contains the real-time position of the hemostat. During fusion, each accounts for 50%, ensuring synchronized transmission of instructions and actions and avoiding errors caused by the loss of information in one modality.

[0040] Existing technologies use a fixed ratio to blend audio and video (e.g., 50% audio + 50% video), which cannot match the differentiated needs of different scenarios (e.g., bandwidth is wasted on video in meetings, and bandwidth is wasted on audio in movies). This step, however, dynamically outputs α through a scenario classification model to achieve "weight allocation on demand," ensuring the quality of high-priority modalities while reducing the ineffective bandwidth usage of low-priority modalities, ultimately improving transmission efficiency and scenario adaptability.

[0041] S5. Generate a physically unclonable session master key K_main using a quantum random number generator. The quantum random number generator is based on the principle of laser phase noise and outputs an entropy value ≥128 bits.

[0042] S6. Use the Kyber-768 algorithm to encapsulate the session master key K_main to generate the encrypted session key K_enc; Steps S5 and S6 above describe the quantum key generation and encapsulation process. A quantum random number generator (QRNG) emits laser light through a helium-neon laser, collects the laser phase noise signal using a photon detector, converts it into a raw random sequence via a 14-bit ADC, and then processes it using a Toeplitz matrix hash function to generate a 128-bit session master key K_main (whose randomness is verified by NIST SP 800-22 testing). The Kyber-768 algorithm is used to encrypt K_main using a public key, generating an encrypted session key K_enc, which resists Shor's algorithm attacks from quantum computers. The quantum randomness of the laser phase noise solves the "unpredictability" of the key, while the Kyber-768 algorithm solves the "resistance to quantum attacks" in key transmission. The combination of these two approaches forms a complete security chain from key generation to transmission. In scenarios with extremely high security requirements, such as remote surgery and financial transactions, this scheme not only meets encryption needs but also reserves security redundancy for the future quantum computing era, making it a typical implementation of "quantum-secure communication" in the 5G / 8K era.

[0043] S7. Perform frame-level dynamic encryption on the fused feature stream: Subkey K_frame=HMAC-SHA256(K_enc,frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Encrypt each frame of the fused feature stream using the AES-256-GCM algorithm to obtain the encrypted fused feature stream.

[0044] This step involves frame-level dynamic encryption. The fused feature stream is encrypted frame by frame. For the t-th frame, the subkey K_frame is generated using the HMAC-SHA256 algorithm, with inputs including K_enc, the frame number t, and a millisecond-level timestamp (e.g., 1620000000000), ensuring the key is unique for each frame. The frame data is encrypted using the AES-256-GCM algorithm, and a 128-bit authentication tag is generated simultaneously to prevent data tampering (e.g., hackers forging a "stop surgery" instruction).

[0045] S8. Determine the transmission environment. If the transmission link is a wired environment, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth ≥192Gbps. If the transmission link is a wireless environment, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, with an end-to-end latency ≤20ms. When the packet loss rate is within 20%, data integrity is restored through forward error correction.

[0046] This step is an adaptive transmission process. If the operating room uses a wired connection (such as fiber optic), the GPMI Type-B interface protocol is enabled: 4-channel differential signal transmission (48Gbps per channel) achieves a total bandwidth of 192Gbps, simultaneously transmitting fused feature streams (surgical audio and video), control signal streams (such as robotic arm feedback signals), and quantum key streams (real-time key updates) to meet the high bandwidth requirements of 8K surgical video. If a sudden wired failure switches to a 5G wireless link, the QUIC-based BBRv3 algorithm is used: 3 5G sub-links (main link, backup link, and emergency link) are aggregated using SD-ARC technology, dynamically adjusting the transmission ratio of each link. When the packet loss rate is 10%, the forward error correction (FEC) redundancy is automatically adjusted to 20% (e.g., adding 20 redundant frames for every 100 data frames transmitted), ensuring smooth surgical footage and end-to-end latency within 18ms, meeting the real-time requirements of surgical operations.

[0047] The audio and video signal fusion transmission method of this embodiment has the following advantages compared with the prior art: High transmission efficiency: By replacing physical layer multiplexing with semantic layer fusion (audio and video joint feature matrix), bandwidth utilization is increased to over 90%, solving the bandwidth waste problem of traditional TDM technology; Low latency and interference resistance: Wireless transmission latency ≤20ms, FEC dynamic adjustment mechanism can still ensure data integrity under 20% packet loss rate, meeting the needs of real-time scenarios such as remote surgery; Quantum-resistant security: An encryption system based on quantum random numbers and the Kyber-768 algorithm extends key cracking time to an order of magnitude unattainable by quantum computers, overcoming the vulnerability of traditional AES-256 to quantum attacks. Simultaneously, scene-adaptive dynamic weight allocation (such as balancing the importance of audio and video in remote surgery) avoids resource waste and adapts to the diverse transmission needs of the 5G / 8K era.

[0048] In one implementation, the scene coefficient α output by the scene classification model includes: Acquire the speech activity detection (VAD) results of the audio signal (speech percentage within 500ms) and the keyframe motion vector magnitude of the video signal (average pixel movement distance). The voice activity detection (VAD) results and the keyframe motion vector magnitudes are input into the scene classification model to obtain scene coefficients α; wherein, the scene classification model is trained based on a support vector machine (SVM) classifier.

[0049] Specifically, taking the remote surgery scenario as an example, we will compare it with the conference scenario: VAD results: In remote surgery, doctors spoke for an average of 200ms every 500ms (40% speech); in video conferencing scenarios, participants spoke for an average of 350ms every 500ms (70% speech).

[0050] Motion vector amplitude: In the surgical scene, the average pixel distance of the scalpel movement is 15 pixels / frame; in the meeting scene, the average pixel distance of the participant's head rotation is 5 pixels / frame.

[0051] SVM classifier training and inference: Training Phase: Collect 100,000 sets of samples (including meeting, video, and surveillance scenes). Each set of samples includes VAD value (0~100%), motion vector amplitude (0~50 pixels), and corresponding α label (0.6~0.7 for meetings, 0.2~0.3 for videos, and 0.5 for surveillance). SVM maps the samples to a high-dimensional space using a kernel function (such as the RBF kernel), constructs a classification hyperplane, and learns the combined features of VAD and motion vector amplitude (e.g., high VAD + low motion vector corresponds to meeting scenes).

[0052] Inference phase: Input the surgical scene with VAD=40% and motion vector=15 pixels into the model, and the model outputs α=0.5; input the meeting scene with VAD=70% and motion vector=5 pixels into the model, and the model outputs α=0.65, achieving accurate matching of scene coefficients.

[0053] In this embodiment, by combining the features of VAD and motion vectors, the SVM classifier can quickly distinguish scene types (accuracy ≥ 95%). Compared with a fixed allocation strategy, the dynamic output of scene coefficient α makes the fused feature stream more in line with scene requirements (such as prioritizing voice in a meeting and balancing audio and video in surgery), reducing invalid data transmission and further improving transmission efficiency.

[0054] In one embodiment, the above similarity calculation layer employs a multi-head attention mechanism with 8 attention heads, each with a feature dimension of 64. The similarity weight calculation formula is as follows:

[0055] Let be the similarity weight between the i-th audio feature query vector and the j-th video feature key vector; The i-th row vector (query vector) of the audio feature query matrix; Let be the j-th row vector (key vector) of the video feature key matrix. =64 is the feature dimension of a single attention head (total feature dimensions 512 = number of heads 8 × ); Represents the vector dot product. This is a scaling factor to avoid gradient vanishing.

[0056] Taking the scenario of "matching doctor's instructions with instrument movements" in remote surgery as an example, the implementation of the similarity calculation layer is explained: Multi-head attention mechanism decomposition: The 512-dimensional audio semantic feature vector (such as the speech features of "clamping the hemostat") and video semantic feature vector (such as the action features of the hemostat closing) are decomposed into 8 attention heads with 64 dimensions per head. Each head independently calculates local similarity. For example, the first head focuses on the speech spectrum features of "clamping" and the motion trajectory features of the hemostat "closing"; the second head focuses on the correlation between speech rhythm (such as the speed of instruction) and the speed of the instrument action, achieving multi-dimensional semantic matching.

[0057] Query (Q) and Key (K) Vector Generation: Audio Feature Query Vector Generated from audio semantic features through linear transformation, representing the semantic query target of the audio (e.g., "finding the action corresponding to 'clamp'"); video feature key vector. Generated by linear transformation of video semantic features, representing the semantic attributes of the video (e.g., "state of the hemostat: closed / open"). Used in remote surgery. This may correspond to the pronunciation characteristics of "clamping". The pixel change features corresponding to the closure of the hemostat.

[0058] Similarity weight calculation: Calculation and The inner product (such as the feature matching value between the "clamping" speech and the closing action), divided by =8 (scaling factor) to avoid excessively large inner product values ​​due to excessively high feature dimensions, which could lead to gradient vanishing in the softmax function (i.e., weights concentrate on a very small number of vectors, ignoring secondary correlations). After normalization using the softmax function... Quantify the semantic similarity between the two (e.g., the sound of "clamping" versus the action of closing). =0.8, compared to the opening action =0.1).

[0059] Multi-head result aggregation: The weight results of the 8 attention heads are concatenated and integrated into a 512-dimensional global similarity weight matrix through linear transformation, which is finally used for cross-modal mapping of audio and video features (such as prioritizing the fusion of highly similar "clamping" speech and closing action features).

[0060] In this embodiment, compared with single-head attention, the 8-head mechanism can simultaneously capture multi-dimensional associations of audio and video (such as semantics, rhythm, and intensity), improving the similarity calculation accuracy by more than 30%. The scaling factor effectively avoids gradient vanishing, making the softmax weight distribution more balanced, ensuring that secondary associated features (such as background noise and slight instrument shaking) can also be included in the fusion, and improving the robustness of complex scenarios (such as sudden noise and unexpected movements during surgery).

[0061] In one implementation, the aforementioned GPMI Type-B interface protocol uses four-channel differential signal transmission with a rate of 48Gbps per channel, is compatible with the USB Type-C physical interface, and supports reverse power supply (up to 100W).

[0062] Taking the scenario of "wired connection between surgical robot and main control console" in remote surgery as an example, the application of GPMI Type-B interface is explained: Four-channel differential transmission: The interface contains four sets of differential signal pairs (two sets each of Tx+ / - and Rx+ / -), with each channel supporting a rate of 48Gbps (using PAM4 modulation technology, with each symbol carrying 2 bits of information), for a total bandwidth of 4 × 48Gbps = 192Gbps. In remote surgery, channel 1 transmits the fused feature stream of 8K surgical video (approximately 120Gbps), channel 2 transmits the control signal stream of the doctor's operation instructions (approximately 5Gbps), channel 3 transmits the quantum key stream (approximately 2Gbps), and channel 4 is reserved as a backup link (to cope with sudden increases in data volume), ensuring that multiple streams run concurrently without conflict.

[0063] USB Type-C Compatibility: The physical interface adopts the USB Type-C form factor (24-pin design), compatible with the Type-C ports of existing medical devices (such as the data interface of surgical robots and the display interface of the main control console), eliminating the need for additional wiring. For example, the surgical robot can be directly connected to the main control console via a Type-C cable, simultaneously transmitting audio and video data and power, simplifying the wiring complexity of the operating room.

[0064] Reverse power supply function: The interface supports the PD3.1 protocol and provides up to 100W of power (20V / 5A), which can directly power low-power devices in surgical scenarios, such as high-definition endoscope cameras (power consumption 30W) and voice acquisition microphones (power consumption 5W). In the event of a sudden power failure, the main control panel can provide reverse power to the camera through the interface to avoid interruption of the surgical video.

[0065] In this embodiment, the total bandwidth of 192Gbps meets the transmission requirements of 8K / 60fps video + 32-channel audio, which is 4 times higher than HDMI 2.1 (48Gbps); USB Type-C compatibility reduces the difficulty of device modification, and reverse power supply reduces power cables, making it suitable for space-constrained operating room scenarios; differential signal transmission improves electromagnetic interference resistance (EMI attenuation ≥40dB), avoiding the impact of electromagnetic noise from surgical equipment (such as electrosurgical units) on data transmission, etc.

[0066] In one implementation, the coding redundancy of the aforementioned forward error correction is dynamically adjusted according to the packet loss rate, including: redundancy = 10% when packet loss rate ≤ 5%, redundancy = 20% when 5% < packet loss rate ≤ 15%, and redundancy = 30% when packet loss rate > 15%.

[0067] Taking the scenario of "sudden packet loss in 5G wireless transmission" during remote surgery as an example, the FEC dynamic adjustment mechanism is explained: Real-time packet loss rate monitoring: Through the ACK mechanism of the QUIC protocol, the packet loss rate (number of lost data packets / total number of transmitted data packets) is calculated every 100ms. For example, in the early stage of surgery, when the wireless link is stable, the packet loss rate is 3% (≤5%); during sudden interference, the packet loss rate rises to 12% (5% < 12% ≤ 15%); when the interference intensifies, the packet loss rate reaches 18% (> 15%).

[0068] Redundancy is dynamically adjusted: When the packet loss rate is 3%, the redundancy is 10%: for every 100 data frames transmitted, 10 redundant frames (using RS(110,100) encoding) are added. The amount of redundant data is small, avoiding bandwidth waste. At this time, the surgical screen is smooth and there is no obvious lag.

[0069] When the packet loss rate is 12% and the redundancy is 20%, switch to RS(120,100) encoding, add 20 redundant frames every 100 frames, and recover up to 20 lost frames. Even if some data is lost, the surgical scene remains intact through reconstruction using redundant frames (latency increase ≤5ms).

[0070] When the packet loss rate is 18% and the redundancy is 30%, RS(130,100) encoding is used, increasing the number of redundant frames to 30, ensuring that packet loss within 20% can be recovered. Although bandwidth usage increases at this point, no critical surgical frames (such as the moment of hemostasis) are lost, meeting the operational safety requirements.

[0071] In this embodiment, compared to a fixed redundancy (e.g., 20%), the dynamic adjustment mechanism saves 30% of bandwidth under low packet loss conditions and improves recovery success rate under high packet loss conditions, balancing transmission efficiency and reliability. In remote data transmission, it can avoid screen tearing or command delays caused by packet loss, ensuring the security of remote control (such as remote surgery).

[0072] In one implementation, the original random sequence of the aforementioned quantum random number generator is generated by acquiring laser phase noise signals through a photon detector, performing a 14-bit analog-to-digital conversion, and then post-processing it using a Toeplitz matrix hash function. The processed sequence is then tested and verified using NIST SP 800-22.

[0073] Taking the "session key generation" scenario in remote surgery as an example, the workflow of QRNG is explained: Laser phase noise acquisition: A 1550nm wavelength distributed feedback laser (DFB) is used, whose output laser phase fluctuates randomly due to spontaneous emission (phase noise). This noise signal is acquired by an avalanche photodiode (APD) detector and converted into an electrical signal (analog quantity) to capture the unpredictable quantum randomness in the surgical environment (non-pseudo-random sequence, which cannot be predicted by the algorithm).

[0074] 14-bit ADC Conversion: Analog electrical signals are converted into digital sequences via a 14-bit ADC (sampling rate 1 GSps), generating a raw random sequence (14 bits per sample, outputting 14 Gbit of raw data per second). The 14-bit precision ensures that subtle fluctuations in the sequence are preserved (1.8 times more randomness compared to a 12-bit ADC), avoiding loss of randomness due to quantization errors.

[0075] Toeplitz matrix hashing post-processing: The original sequence is hashed using a Toeplitz matrix (a special type of strip matrix) to eliminate biases in the sequence (such as mean shifts caused by detector dark currents), outputting a random sequence that conforms to a uniform distribution. For example, 14 Gbit of raw data is processed into 128-bit / group subsequences as candidate values ​​for the session master key K_main.

[0076] NISTSP800-22 Validation: The processed sequence undergoes 20 statistical tests (such as frequency test, run-length test, and linear complexity test) to ensure that all metrics are met (e.g., a p-value > 0.01 for the frequency test, indicating no significant bias). In remote surgery, only validated sequences are used to generate K_main, guaranteeing the unpredictability of the key.

[0077] In this embodiment, compared to traditional pseudo-random number generators (such as CTR_DRBG based on AES), the random sequence of QRNG originates from quantum physical processes, improving its anti-predictability by 100% (it cannot be derived in reverse by algorithms); Toeplitz hashing and NIST testing ensure sequence uniformity, and resists key cracking attacks by quantum computers from the root through quantum randomness, providing "quantum-level" security for audio and video transmission in scenarios such as remote surgery.

[0078] In one embodiment, the LSTM network described above contains 3 hidden layers, each with 256 neurons. Dropout technology (scale 0.2) is used to prevent overfitting, and the input Mel spectrum dimension is 128×64 (time frame × frequency bin).

[0079] Taking the scenario of "feature extraction of doctor's voice commands" in remote surgery as an example, the workings of the LSTM network are explained: Input Mel spectrum construction: The doctor's voice command (such as "move the hemostatine 3 cm to the left") is preprocessed (sampling rate 16kHz, mono) and converted into a Mel spectrum: the horizontal axis is the time frame (25ms per frame, step size 10ms), and the vertical axis is 40 Mel frequency bins (covering the 100Hz~8kHz speech band). It is extended to a fixed dimension of 128×64 (128 frames × 64 bins) by zero padding to adapt to LSTM input requirements.

[0080] Calculation of 3-layer LSTM hidden layers: Layer 1: 256 neurons, learn short-term features of speech (such as the pronunciation of the syllable "to the left"), and filter background noise (such as the friction sound of surgical instruments) through gating mechanisms (input gate, forget gate, output gate).

[0081] Layer 2: 256 neurons, capturing mid-term temporal dependencies (such as the logical association between "move" and "3 centimeters") and distinguishing similar pronunciations (such as the spectral differences between "left" and "right").

[0082] Layer 3: 256 neurons, extracting long-term semantic features (such as the intent of a whole instruction: adjust the position of the device), and outputting a 512-dimensional audio semantic feature vector (each neuron corresponds to 1 semantic dimension).

[0083] Dropout prevents overfitting: Each layer is configured with a 20% Dropout ratio, randomly dropping 20% ​​of neural connections during training (but not during inference) to prevent the network from overfitting to specific speech patterns during surgery (such as a doctor's accent). For example, if the training samples contain the speech of 100 doctors, Dropout ensures that the network can still accurately extract features from instructions from unknown doctors (reducing generalization error by 25%).

[0084] In this embodiment, the 3-layer LSTM structure can effectively capture short, medium, and long-term features of speech, improving the accuracy of semantic feature extraction compared to a single-layer LSTM. The 128×64 Mel-spectrum input preserves the details of speech (such as tone changes), and the combination of 256 neurons and Dropout balances the model's complexity and generalization ability, ensuring that even ambiguous instructions from doctors during remote surgery (such as "hurry up" with background noise) can be accurately represented, thus improving the reliability of audio semantic features.

[0085] In one implementation, the aforementioned dynamic weight allocation layer calculates audio feature weights using the following nonlinear fusion formula. and video feature weights : in: For scene coefficients (as defined in claim 2); Audio signal-to-noise ratio (dB) Peak signal-to-noise ratio (dB) for video. The packet loss rate (%) is the link loss rate. Video block error rate (%); Network jitter (ms) Transmission delay (ms) =30ms =100ms is the threshold; =0.05、 =0.03 is the signal quality gain coefficient. =0.1、 =0.08 is the link impairment attenuation coefficient. =0.02、 =0.015 is the real-time penalty coefficient.

[0086] Taking the scenario of "sudden network jitter and video blurring" in remote surgery as an example (α=0.5), the weight adjustment process is calculated as follows: Parameter initialization: Scenario coefficient α = 0.5 (remote surgery), initial state: =30dB (clear speech) =40dB (clear picture), PLR=2%, BLER=1%, Jitter=10ms, Delay=15ms.

[0087] Initial weight calculation: Substitute the above parameters into and ,get ≈0.52, ≈0.48. At this point, the audio weight is slightly higher because the speech clarity is better than the video (but the difference is small).

[0088] Emergency Scenario Adjustment: If a sudden network jitter occurs (Jitter = 25ms, close to the threshold of 30ms), the video will be affected by the interference. =25dB (image blurry): At this point, ≈0.51 (Jitter has a relatively small impact). ≈0.49 (Video quality decreased, weight slightly decreased).

[0089] If the video BLER rises to 10% (high block error rate): ≈0.24, at this time ≈0.76, automatically increasing audio weight and prioritizing the retention of doctor's instructions (voice is more crucial when the picture is blurry).

[0090] In this embodiment, the nonlinear formula achieves adaptive adjustment by dynamically weighting signal quality (SNR / PSNR), link status (PLR / BLER), and real-time performance (Jitter / Delay), resulting in "higher weight for better quality and lower weight for poorer link performance." Compared with fixed-ratio fusion, the retention rate of key information (such as voice commands in blurred images) in remote surgery is improved, enhancing robustness in weak network environments.

[0091] In one implementation, the feature alignment layer of the aforementioned cross-modal attention mechanism dynamically adjusts the sharpness of the similarity calculation using the following adaptive temperature coefficient formula:

[0092] in: =1.0 is the initial temperature coefficient. =0.8 is the adjustment factor; Cosine similarity of the global feature vectors of the audio / video (value range [-1, 1]); =0.3、 =0.2 represents the mean and standard deviation of the similarity distribution (obtained through statistics from the pre-trained dataset).

[0093] Taking the scenario of "audio-video synchronization fluctuations" in remote surgery as an example, the role of the temperature coefficient T is explained: Global feature cosine similarity calculation: For global audio features (such as the average vector of “cut-off” speech). This represents the global features of the video (such as the average vector of a scalpel cutting action). The overall correlation between the two is quantified (the value range is [-1,1]). When the synchronization is high, the similarity is high (e.g., the similarity between the voice of "removal" and the cutting action = 0.6), and when the synchronization is low, the similarity is low (e.g., the similarity between the voice of "removal" and the hemostasis action = 0.1).

[0094] Dynamic adjustment of temperature coefficient T: When the similarity is high (e.g., 0.6): T = 1.0 × exp(-0.8 × (0.6 - 0.3) / 0.2) = 1.0 × exp(-0.8 × 1.5) = 1.0 × 0.301 ≈ 0.301 (Decrease T, increase sharpness). At this time, the softmax function enhances the discriminative power of the weights (the weights of high similarity vectors are more concentrated, such as the weight of 0.6 similarity being amplified and the weight of 0.5 similarity being compressed), and prioritizes strengthening strongly correlated features (such as the fusion of "removal" and cutting actions).

[0095] When the similarity is moderate (e.g., 0.3): T = 1.0 × exp(-0.8 × (0.3 - 0.3) / 0.2) = 1.0 × 1 = 1.0 (default sharpness), with balanced weight distribution, taking into account both strong and weak correlations.

[0096] When the similarity is low (e.g., 0.1): T = 1.0 × exp(-0.8 × (0.1 - 0.3) / 0.2) = 1.0 × exp(-0.8 × (-1)) = 1.0 × 2.225 ≈ 2.225 (Increasing T decreases sharpness). At this time, the softmax function weight distribution is more gradual (the difference in weight between 0.1 similarity and 0.2 similarity is reduced), avoiding the neglect of potential associations due to weak overall association (such as the indirect association between the "removal" speech and the adjustment of instrument position).

[0097] Impact on similarity calculation: Temperature coefficient T is adjusted by the above. The input to the softmax function controls the "sharpness" of the weight distribution: the smaller T is, the more concentrated the weights are on highly similar vectors (high sharpness); the larger T is, the more dispersed the weights are (low sharpness).

[0098] In this embodiment, compared to a fixed temperature coefficient (e.g., T=1.0), the adaptive T can dynamically adjust the fusion sharpness according to the global correlation of audio and video: when the synchronization is high, the key correlation is strengthened (improving the fusion accuracy by 15%), and when the synchronization is low, the secondary correlation is retained (avoiding information loss), effectively solving the problem of audio and video asynchrony caused by device delay in remote surgery (e.g., when there is a 50ms delay between voice commands and actions, the correlation can still be captured through low-sharpness fusion).

[0099] In one embodiment, the aforementioned frame-level dynamic encryption employs a "quantum key pool-dynamic window" update mechanism, with the subkey... Generate using the following sliding window function:

[0100] in: For a quantum key pool (capacity N = 1024, storing 128-bit subkeys generated by QRNG). The current frame number. The message authentication code for the previous frame (to prevent replay attacks); After every N / 2 frames transmitted, the key pool is updated via the post-quantum KEM protocol (CRYSTALS-Kyber). .

[0101] Taking the scenario of "8K surgical video encryption" in remote surgery as an example (each frame encryption requires a 128-bit subkey), the key update mechanism is explained as follows: Quantum key pool initialization: 1024 128-bit subkeys are generated from QRNG. ~ These subkeys are stored in an encryption chip (such as an HSM hardware security module) and form the initial key pool. Each subkey has quantum randomness and is stored only locally, not transmitted over a network.

[0102] Subkey generate: When transmitting frame t=100, 1027 = 100, take... (128 bits); Calculate HMAC-SHA384 (Hash Message Authentication Code): Input is the session master key. The current frame number 100 and the previous frame (t=99) (64-bit), outputting a 384-bit hash value, with the first 128 bits truncated as a dynamic factor; By XOR operation By fusing with dynamic factors, a unique subkey is generated. This is used to encrypt the 100th frame of data.

[0103] Anti-replay attack mechanism: The message authentication code (generated by AES-GCM) for the previous frame contains the frame number and data digest. If a hacker replays an old frame (e.g., frame t=50), the current frame t=100... With the old frame Mismatch leads to incorrect subkey generation and decryption failure, effectively preventing repeated attacks (such as repeatedly playing the old "stop surgery" command).

[0104] Key pool updates dynamically: Every 512 frames (approximately 34 seconds, 15fps) are transmitted, and 1024 new subkeys are generated via the CRYSTALS-Kyber protocol (post-quantum KEM), overwriting the old ones. During the update process, the sender encrypts the new key pool using the receiver's public key, and the receiver decrypts it using their private key, ensuring the security of the update process (resistant to quantum computing interception).

[0105] In this embodiment, compared to static keys (such as a single session key), the dynamic window mechanism makes the subkey "one-time pad" (the key is different for each frame), so even if the key of a certain frame is leaked, it is impossible to deduce the keys of other frames; the key pool is updated once every 512 frames, which achieves a balance between security (short key lifespan) and efficiency (low update overhead). Combined with the unpredictability of the quantum key pool, a "dynamic + quantum" dual security barrier is built for audio and video transmission.

[0106] In one embodiment, the link switching decision for the hybrid protocol transmission described above is achieved through the following multi-objective optimization formula:

[0107] in: B represents the current link bandwidth. =200Gbps is the theoretical maximum bandwidth; =0.4 (bandwidth weight) =0.35 (delay weight) =0.25 (packet loss rate weight), determined using the Analytic Hierarchy Process (AHP); When wired link When the value is less than 0.6, the system automatically switches to the wireless aggregation link (SD-ARC technology).

[0108] Taking the scenario of "sudden wired link failure" in remote surgery as an example, this illustrates the link switching decision: Parameter definition: =200Gbps (theoretical maximum bandwidth). =200ms (maximum tolerable latency), weight =0.4、 =0.35、 =0.25, switching threshold 0.6.

[0109] Normal wired link rating: Wired link status: B=192Gbps (GPMI Type-B interface), Delay=10ms, PLR=0.1%.

[0110] ≈0.966 (>0.6, maintaining wired transmission).

[0111] Wired fault scenario rating: Sudden fiber optic cable breakage, wired link speed reduced to B=20Gbps, Delay=150ms, PLR=5%. ≈0.366 (<0.6, triggers switching).

[0112] Wireless aggregation link switching: Switching to 3-link aggregation (SD-ARC technology), status: B=5Gbps×3=15Gbps (although lower than wired, it meets the minimum requirements for 8K video), Delay=18ms, PLR=3%. ≈≈0.591 (close to the threshold, maintaining wireless transmission).

[0113] In this embodiment, the multi-objective optimization formula integrates the weights of bandwidth, latency, and packet loss rate to avoid misjudgment based on a single indicator (such as prioritizing links with high bandwidth but high latency); the switching threshold of 0.6 ensures a smooth transition between "available and unavailable" links, with a link switching time of <100ms during remote surgery and no obvious stuttering, thus solving the latency problem of traditional "either / or" switching strategies.

[0114] In one embodiment, the aforementioned video semantic feature vector is extracted using the improved visual Transformer model, whose attention mask matrix M satisfies:

[0115] Wherein: x and y are image patch indices, and k = 3 is the local attention window size; is the image patch and 's intersection over union (object detection box overlap); The mask value 1 represents strong attention, 0.2 represents weak attention, and 0 represents no attention.

[0116] Taking the scenario of "image feature extraction of surgical field" in remote surgery as an example, the attention mechanism of the improved Vision Transformer model is illustrated: Image patch segmentation and indexing: The surgical field image (e.g., 512×512 pixels) is segmented into image patches of 32×32 pixels (a total of 16×16 = 256 patches), and the indices x, y = 0~255 (sorted by row first). For example is an image patch containing a hemostatic forceps, is an adjacent tissue patch.

[0117] Local attention window (k = 3): When |x - y| ≤ 3 (e.g., x = 10, y = 8~12) = 1, indicating that strong attention interaction (capturing local associations) is allowed between these two patches. For example, the hemostatic forceps patch and the adjacent 3 tissue patches (y = 11, 12, 13)'s = 1, preferentially learning the spatial relationship between the instrument and the surrounding tissues.

[0118] Intersection over union (IoU) judgment: is the proportion of the overlapping area of two image patches (e.g., and overlap due to the extended part of the hemostatic forceps, = 0.4): > 0.3 (e.g., 0.4): = 1, strong attention (such as the association between the hemostatic forceps and its extended part); 0.1 < IoU ≤ 0.3 (e.g., 0.2): = 0.2, weak attention (such as the slight overlap between the hemostatic forceps and a distant gauze); IoU ≤ 0.1: = 0, no attention (such as the patch of the hemostatic forceps and a background instrument).

[0119] Attention weight adjustment: The mask matrix M is multiplied by the original attention weights to strengthen the association between local and highly overlapping blocks (weight × 1), weaken the association between low-overlapping blocks (weight × 0.2), and ignore irrelevant blocks (weight × 0). For example, feature extraction of hemostats focuses more on local tissue and direct contact areas, reducing background interference. Compared to the standard ViT (Global Attention), the improved ViT in this embodiment reduces invalid attention calculations (such as associations between background blocks) by 80% through the mask matrix, improving feature extraction speed by 2 times. At the same time, local window and IoU judgment ensure that features of key surgical areas (such as instrument-tissue contact points) are captured first, improving feature extraction accuracy by 25%, providing a more accurate foundation for video semantic understanding in remote surgery.

[0120] Reference Figure 2 This invention also provides an audio-visual signal fusion transmission apparatus for implementing the audio-visual signal fusion transmission method of any of the above embodiments, comprising: The audio semantic feature extraction module 10 is used to perform Mel-spectrum transformation and MFCC feature extraction on the input audio signal, capture temporal dependencies through an LSTM network, and generate a 512-dimensional audio semantic feature vector; and, The video semantic feature extraction module 20 is used to sample keyframes of the input video signal, extract scene, action and expression features using a visual Transformer model, and generate a 512-dimensional video semantic feature vector. The cross-modal joint module 30 is used to calculate the similarity weight between the audio semantic feature vector and the video semantic feature vector through a cross-modal attention mechanism, and to map the audio semantic feature vector and the video semantic feature vector to a shared semantic space based on the similarity weight to generate an audio-video joint feature matrix; the cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type; The scene fusion module 40 is used to output scene coefficients α through the scene classification model, and adjust the audio and video joint feature matrix based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; The generation module 50 is used to generate a physically unclonable session master key K_main through a quantum random number generator. The quantum random number generator is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. Encapsulation module 60 is used to encapsulate the session master key K_main using the Kyber-768 algorithm to generate an encrypted session key K_enc; The encryption module 70 is used to perform frame-level dynamic encryption on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission judgment module 80 is used to determine the transmission environment. If the transmission link is a wired environment, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth of ≥192Gbps. If the transmission link is a wireless environment, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, with an end-to-end latency of ≤20ms. When the packet loss rate is within 20%, data integrity is restored through forward error correction.

[0121] Reference Figure 3 The present invention also provides a computer device, the internal structure of which can be as follows: Figure 3 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor is designed to provide computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores operating devices, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores audio and video data, etc. The network interface is used to communicate with external terminals via a network connection. Furthermore, the computer device may also include input devices and a display screen. When the computer program is executed by the processor, it implements an audio and video signal fusion transmission method, including the following steps: The input audio signal undergoes Mel-spectrum transformation and MFCC feature extraction. Temporal dependencies are captured using an LSTM network, generating a 512-dimensional audio semantic feature vector. Keyframe sampling is performed on the input video signal, and a visual Transformer model is used to extract scene, action, and facial features, generating a 512-dimensional video semantic feature vector. The similarity weight between the audio semantic feature vector and the video semantic feature vector is calculated through a cross-modal attention mechanism. Based on the similarity weight, the audio semantic feature vector and the video semantic feature vector are mapped to a shared semantic space to generate a joint audio-video feature matrix. The cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer, and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type. The scene classification model outputs scene coefficients α, and the audio and video joint feature matrix is ​​adjusted based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; A physically unclonable session master key K_main is generated by a quantum random number generator, which is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. The session master key K_main is encapsulated using the Kyber-768 algorithm to generate the encrypted session key K_enc; Frame-level dynamic encryption is performed on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission environment is determined. If the transmission link is wired, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth of ≥192Gbps. If the transmission link is wireless, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, and data integrity is restored through forward error correction when the packet loss rate is within 20%.

[0122] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.

[0123] One embodiment of this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements an audio-video signal fusion transmission method, including the following steps: performing Mel-frequency conversion and MFCC feature extraction on the input audio signal, capturing temporal dependencies through an LSTM network, and generating a 512-dimensional audio semantic feature vector; and sampling keyframes of the input video signal, extracting scene, action, and facial expression features using a visual Transformer model, and generating a 512-dimensional video semantic feature vector; calculating the similarity weight between the audio semantic feature vector and the video semantic feature vector through a cross-modal attention mechanism, and mapping the audio semantic feature vector and the video semantic feature vector to a shared semantic space based on the similarity weight, generating an audio-video joint feature matrix; the cross-modal attention mechanism includes... The system consists of a feature alignment layer, a similarity calculation layer, and a dynamic weight allocation layer. The dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type. A scene classification model outputs a scene coefficient α, and the audio-video joint feature matrix is ​​adjusted based on α to obtain the fused feature stream: fused feature stream = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scenes, film and television scenes, and surveillance scenes. A physically unclonable session master key K_main is generated using a quantum random number generator based on the laser phase noise principle, with an output entropy value ≥128 bits. The session master key K_main is encapsulated using the Kyber-768 algorithm to generate an encrypted session key K_enc. The fused feature stream is then dynamically encrypted at the frame level: subkey K_frame = HMAC-SHA256(K_enc, The data is processed using the `frame_index || timestamp` method, where `frame_index` is the frame sequence number and `timestamp` is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission environment is determined: if the transmission link is wired, the encrypted fused feature stream is transmitted using the GPMI Type-B interface protocol supporting bidirectional multi-stream transmission. This protocol includes three parallel sub-streams: the fused feature stream, the control signal stream, and the quantum key stream, with a total bandwidth ≥ 192 Gbps. If the transmission link is wireless, the encrypted fused feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, and data integrity is restored through forward error correction when the packet loss rate is within 20%. It is understood that the computer-readable storage medium in this embodiment can be either a volatile readable storage medium or a non-volatile readable storage medium.

[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0125] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0126] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for audio and video signal fusion transmission, characterized in that, Includes the following steps: The input audio signal is subjected to Mel spectrum transformation and MFCC feature extraction. Temporal dependencies are captured by an LSTM network to generate a 512-dimensional audio semantic feature vector. as well as, Keyframe sampling is performed on the input video signal, and a visual Transformer model is used to extract scene, action, and facial features, generating a 512-dimensional video semantic feature vector. The similarity weight between the audio semantic feature vector and the video semantic feature vector is calculated through a cross-modal attention mechanism. Based on the similarity weight, the audio semantic feature vector and the video semantic feature vector are mapped to a shared semantic space to generate a joint audio-video feature matrix. The cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer, and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type. The scene classification model outputs scene coefficients α, and the audio and video joint feature matrix is ​​adjusted based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; A physically unclonable session master key K_main is generated by a quantum random number generator, which is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. The session master key K_main is encapsulated using the Kyber-768 algorithm to generate the encrypted session key K_enc; Frame-level dynamic encryption is performed on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index|| timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission environment is determined. If the transmission link is wired, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth of ≥192Gbps. If the transmission link is wireless, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, and data integrity is restored through forward error correction when the packet loss rate is within 20%.

2. The audio and video signal fusion transmission method according to claim 1, characterized in that, The step of outputting scene coefficient α through the scene classification model includes: Acquire the speech activity detection results of the audio signal and the motion vector amplitude of the keyframes of the video signal; The speech activity detection results and the keyframe motion vector magnitudes are input into the scene classification model to obtain scene coefficients α; wherein, the scene classification model is trained based on a support vector machine classifier.

3. The audio and video signal fusion transmission method according to claim 1, characterized in that, The similarity calculation layer employs a multi-head attention mechanism with 8 attention heads, each with a feature dimension of 64. The similarity weight calculation formula is as follows: Let be the similarity weight between the i-th audio feature query vector and the j-th video feature key vector; This is the vector of the i-th row of the audio feature query matrix; Let j be the vector of the j-th row of the video feature key matrix; =64 represents the feature dimension of a single attention head; Represents the dot product of vectors. This is a scaling factor to avoid gradient vanishing.

4. The audio and video signal fusion transmission method according to claim 1, characterized in that, The GPMI Type-B interface protocol uses four-channel differential signal transmission with a rate of 48Gbps per channel, is compatible with the USB Type-C physical interface, and supports reverse power supply.

5. The audio and video signal fusion transmission method according to claim 1, characterized in that, The coding redundancy of the forward error correction is dynamically adjusted according to the packet loss rate, including: redundancy = 10% when packet loss rate ≤ 5%, redundancy = 20% when 5% < packet loss rate ≤ 15%, and redundancy = 30% when packet loss rate > 15%.

6. The audio and video signal fusion transmission method according to claim 1, characterized in that, The dynamic weight allocation layer calculates audio feature weights using the following nonlinear fusion formula. and video feature weights : in: For scene coefficients; For audio signal-to-noise ratio, Peak signal-to-noise ratio (PSNR) of the video; For link packet loss rate, Video block error rate; For network jitter, For transmission delay, =30ms =100ms is the threshold; =0.05、 =0.03 is the signal quality gain coefficient. =0.1、 =0.08 is the link impairment attenuation coefficient. =0.02、 =0.015 is the real-time penalty coefficient.

7. The audio and video signal fusion transmission method according to claim 1, characterized in that, The video semantic feature vector is extracted using the improved visual Transformer model, and its attention mask matrix M satisfies: in: x, y are the image patch indices, and k=3 is the size of the local attention window; For image blocks and The intersection and union ratio; A mask value of 1 indicates strong attention, 0.2 indicates weak attention, and 0 indicates no attention.

8. An audio and video signal fusion transmission device, characterized in that, include: The audio semantic feature extraction module performs Mel-frequency conversion and MFCC feature extraction on the input audio signal, captures temporal dependencies through an LSTM network, and generates a 512-dimensional audio semantic feature vector; and, The video semantic feature extraction module is used to sample keyframes of the input video signal, and uses a visual Transformer model to extract scene, action and expression features, generating a 512-dimensional video semantic feature vector. A cross-modal joint module is used to calculate the similarity weight between the audio semantic feature vector and the video semantic feature vector through a cross-modal attention mechanism, and to map the audio semantic feature vector and the video semantic feature vector to a shared semantic space based on the similarity weight to generate an audio-video joint feature matrix; the cross-modal attention mechanism includes a feature alignment layer, a similarity calculation layer and a dynamic weight allocation layer, wherein the dynamic weight allocation layer adjusts the fusion ratio of audio and video features in real time according to the scene type; The scene fusion module is used to output scene coefficients α through the scene classification model, and adjust the audio and video joint feature matrix based on α to obtain the fused feature flow: fused feature flow = α·audio semantic feature vector + (1-α)·video semantic feature vector, where the scene includes meeting scene, film and television scene and monitoring scene; The generation module is used to generate a physically unclonable session master key K_main through a quantum random number generator. The quantum random number generator is based on the principle of laser phase noise and outputs an entropy value ≥128 bits. The encapsulation module is used to encapsulate the session master key K_main using the Kyber-768 algorithm to generate the encrypted session key K_enc. The encryption module is used to perform frame-level dynamic encryption on the fused feature stream: subkey K_frame=HMAC-SHA256(K_enc, frame_index || timestamp), where frame_index is the frame sequence number and timestamp is a millisecond-level timestamp. Each frame of the fused feature stream is encrypted using the AES-256-GCM algorithm to obtain the encrypted fused feature stream. The transmission determination module is used to determine the transmission environment. If the transmission link is a wired environment, the encrypted fusion feature stream is transmitted using the GPMI Type-B interface protocol that supports bidirectional multi-stream. The protocol includes three parallel sub-streams: the fusion feature stream, the control signal stream, and the quantum key stream, with a total bandwidth ≥192Gbps. If the transmission link is a wireless environment, the encrypted fusion feature stream is transmitted using a dynamic congestion control algorithm based on the QUIC protocol. Multi-link aggregation is achieved through SD-ARC technology, with an end-to-end latency ≤20ms. When the packet loss rate is within 20%, data integrity is restored through forward error correction.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Robot video interaction method and system

    CN112651334A

  • Video transmission optimization method and device based on SDWAN

    CN116170371A

  • Optimization method and system based on audio and video fusion communication technology

    CN117135150A

  • Instance segmentation method and device, equipment and medium

    CN117372691A

  • Audio and video joint coding and decoding method and system based on generative artificial intelligence

    CN119583873A

Cited By

  • Cross-modal retrieval method supporting authorized access

    CN121351147A

  • A cross-modal retrieval method supporting authorized access

    CN121351147B

  • High-quality vehicle-mounted audio and video call method and system based on mobile internet

    CN121864771A