Audio and video cross-network conference system based on security data exchange
By introducing secure isolation switching devices and optimizing data transmission processes in audio and video cross-network conference systems, the problems of data transmission security risks and equipment performance degradation in cross-network audio and video conferences are solved, and audio and video transmission with high security, stability and compatibility are achieved.
Patent Information
- Application Number
- CN202510696181.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has problems of data transmission security risks and equipment performance degradation in cross-network audio and video conferencing, especially when data transmission is carried out in untrusted Internet networks, where encryption measures are lacking and equipment performance is affected by protocol unpacking and packetization operations.
The audio and video cross-network conferencing system based on secure data exchange is adopted, and strict security inspection and encryption processing is carried out by introducing a secure isolation switching device, and the data transmission process is optimized to reduce unnecessary protocol unpacking and packetization operations. A common data encapsulation and transmission protocol, such as the TCP protocol, only necessary protocol conversion and data encapsulation are carried out between the front and rear audio and video devices.
It improves the security and stability of data transmission, reduces equipment performance requirements, improves the fluency and stability of audio and video conferencing, and simplifies the system architecture, reduces operation and maintenance costs, and improves the compatibility and scalability of the system.
Smart Images

Figure CN120302003A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an audio and video cross-network conferencing system based on secure data exchange. Background Art
[0002] In today's digital age, many units with high security levels have set up internal work networks isolated from the Internet to ensure information security. However, there is still an urgent need for audio and video conferencing between these units and other units, and the other party may also be in an isolated internal network environment. At this time, the communication between the two parties needs to be connected through the untrusted Internet network. For example, for two isolated network systems A and C, an audio and video conference needs to be carried out through the Internet network B.
[0003] The existing technical solutions mainly implement this cross-network audio and video conferencing by deploying two groups of cross-network isolation devices. In this solution, the audio and video front and rear ends A1 and A2 need to perform unpacking, cross-networking, and repackaging of audio and video packets once, and the audio and video front and rear ends B1 and B2 also need to perform unpacking, cross-networking, and repackaging of audio and video packets once. When the front and rear end devices cross the network in the market, generally the same protocol is used for both incoming and outgoing. For example, when incoming, it is H323, and when outgoing, it is also H323. Its core working principle is to connect to the Internet through cross-network isolation devices, convert the internal network audio and video data into a format suitable for transmission on the Internet and then transmit it, and then convert the data back into a format suitable for reception by the internal network at the receiving end.
[0004] However, the existing technical solutions based on this core working principle have many problems. On the one hand, although the communication between the two parties is connected to the Internet through cross-network isolation devices, there is no corresponding encryption measure on the Internet in essence, which makes the data face great security risks during the transmission process. If encryption is to be carried out, special devices need to be used, which undoubtedly greatly increases the cost. On the other hand, when the two parties deploy cross-network devices for protocol cross-networking, in fact, two protocol unpacking and repackaging operations are performed, and two protocol proxies are made. Each time, protocol parsing and modification are required, which seriously affects the device performance and reduces the fluency and stability of the audio and video conference. Summary of the Invention
[0005] The embodiments of this application provide a technical solution for an audio and video cross-network conferencing system based on secure data exchange.
[0006] 1. An audio and video cross-network conferencing system based on secure data exchange, characterized by comprising: a first terminal (1), a first front and rear audio and video device (A1), a first secure isolation and exchange device (U1), a second front and rear audio and video device (A2), a third front and rear audio and video device (B2), a second secure isolation and exchange device (U2), a fourth front and rear audio and video device (B1), and a second terminal (2). The first terminal (1) and the first front and rear audio and video device (A1) are in the first network system, the second front and rear audio and video device (A2) and the third front and rear audio and video device (B2) are in the untrusted second network system, and the fourth front and rear audio and video device (B1) and the second terminal (2) are in the third network system, where:
[0007] The first terminal (1) is used to generate first audio and video conferencing data and transmit it to the first front and rear audio and video device (A1);
[0008] The first front and rear audio and video device (A1) is used to split and encapsulate the first audio and video conferencing data to obtain a second audio and video conferencing data packet and transmit it to the second front and rear audio and video device (A2) through the first secure isolation and exchange device (U1);
[0009] The second front and rear audio and video device (A2) is used to convert the second audio and video conferencing data packet into a general data packet, encapsulate the general data packet using the TCP protocol to obtain an encapsulated general data packet, and send it to the third front and rear audio and video device (B2);
[0010] The third front and rear audio and video device (B2) is used to receive the encapsulated general data packet and restore the general data packet to transmit it to the fourth front and rear audio and video device (B1) through the second secure isolation and exchange device (U2);
[0011] The fourth front and rear audio and video device (B1) is used to decompose the encapsulated general data packet to obtain a general data packet and restore the general data packet into a first audio data packet;
[0012] The second terminal (2) receives the restored first audio data packet and plays it to conduct a cross-network conference between the first terminal (1) and the second terminal (2).
[0013] The technical solution provided by the embodiments of the present application has at least the following beneficial effects:
[0014] The audio-video cross-network conferencing system proposed in this application introduces a secure data exchange mechanism during data transmission. Specifically, the system performs strict security checks and encryption processing on audio-video data before cross-network transmission through the first secure isolation and exchange device (U1) and the second secure isolation and exchange device (U2), ensuring the confidentiality, integrity, and availability of data during transmission. This mechanism effectively prevents data from being stolen or tampered with during transmission, greatly improving the security of data transmission.
[0015] In this application, by optimizing the data transmission process, unnecessary protocol unpacking and packing operations are reduced. Specifically, the system only performs necessary protocol conversion and data encapsulation between the audio-video front and rear devices, avoiding complex protocol parsing and modification on cross-network isolation devices. This design reduces the requirements for device performance and the device burden, thereby improving the fluency and stability of audio-video conferencing.
[0016] The audio-video cross-network conferencing system of this application adopts a common data encapsulation and transmission protocol (such as the TCP protocol), enabling the system to be compatible with various audio-video devices and network environments. This design not only improves the compatibility of the system but also facilitates the future expansion and upgrade of the system. For example, when new audio-video devices or network environments need to be connected, only the corresponding configuration and adjustment of the audio-video front and rear devices are required to achieve the rapid expansion and upgrade of the system. In addition, by optimizing the system architecture, this application reduces the number of deployed cross-network isolation devices. Specifically, the system only deploys secure isolation and exchange devices between the audio-video front and rear devices to achieve the secure transmission and exchange of audio-video data. This design simplifies the system architecture, reduces the operation and maintenance costs, and improves the overall efficiency of the system. Description of the Drawings
[0017] Figure 1 It is a flowchart of the audio-video cross-network conferencing system based on secure data exchange provided for the implementation of this application. Detailed Implementation Manner
[0018] As Figure 1 shown, an audio-video cross-network conferencing system based on secure data exchange includes:
[0019] The first terminal (1), the first front and rear audio-video device (A1), the first security isolation and exchange device (U1), the second front and rear audio-video device (A2), the third front and rear audio-video device (B2), the second security isolation and exchange device (U2), the fourth front and rear audio-video device (B1), and the second terminal (2). The first terminal (1) and the first front and rear audio-video device (A1) are in the first network system, the second front and rear audio-video device (A2) and the third front and rear audio-video device (B2) are in the untrusted second network system, and the fourth front and rear audio-video device (B1) and the second terminal (2) are in the third network system. Among them:
[0020] The first terminal (1) is used to generate the first audio-video conference data and transmit it to the first front and rear audio-video device (A1);
[0021] The first front and rear audio-video device (A1) is used to split and encapsulate the first audio-video conference data to obtain the second audio-video conference data packet and transmit it to the second front and rear audio-video device (A2) through the first security isolation and exchange device (U1);
[0022] The second front and rear audio-video device (A2) is used to convert the second audio-video conference data packet into a general data packet, encapsulate the general data packet using the TCP protocol to obtain the encapsulated general data packet, and send it to the third front and rear audio-video device (B2);
[0023] The third front and rear audio-video device (B2) is used to receive the encapsulated general data packet and restore the general data packet to transmit it to the fourth front and rear audio-video device (B1) through the second security isolation and exchange device (U2);
[0024] The fourth front and rear audio-video device (B1) is used to de-encapsulate the encapsulated general data packet to obtain the general data packet and restore the general data packet into the first audio data packet;
[0025] The second terminal (2) receives the restored first audio data packet and plays it to conduct a cross-network conference between the first terminal (1) and the second terminal (2).
[0026] Optionally, the general data packet includes MagicNumber, metadata length, message body length, metadata, and message body. Among them, MagicNumber is used to distinguish the type of data packet; the message body length is used to record the lengths of the metadata and the message body, the metadata is a record of the message attributes; and the message body is the actual generated message data.
[0027] Preferably, the general data packet format adopts a five-layer standardized structure to break through the limitations of the fixed format of traditional audio-video protocols and achieve cross-network system protocol-independent transmission. The technical details are as follows:
[0028] MagicNumber field (4 bytes)
[0029] Use a hexadecimal unique identifier (such as 0x54435041) to make a low-level distinction from traditional protocol packets such as H.323 / SIP through a binary bit pattern
[0030] Support an extended bit (the highest bit) to mark the packet type: 0 for audio and video data, 1 for control signaling
[0031] At the hardware level, the high-speed matching and filtering of MagicNumber can be achieved through FPGA, with a misjudgment rate lower than 0.001%
[0032] Length field design (4 bytes each)
[0033] Metadata Length: Record the number of bytes in the metadata block, supporting a dynamic length of 0 - 4GB
[0034] Payload Length: Use variable-length encoding (VarInt). For data less than 128 bytes, use single-byte encoding to improve the transmission efficiency of small packets
[0035] This application realizes the decoupling process of metadata and message body through a dual length field, reducing the parsing overhead by 20% compared with the traditional TLV format
[0036] Metadata block (dynamic and variable)
[0037] Use key-value structured storage, including the following core attributes:
[0038] Protocol ID: A 16-bit enumeration value (0x01 - H.323, 0x02 - SIP, 0xFF - custom protocol) Channel Index: An 8-bit unsigned integer, supporting the multiplexing of 256 concurrent audio and video channels
[0039] Seq No: A 32-bit cyclic counter, used in conjunction with a sliding window to achieve packet loss retransmission
[0040] Enc Flag: A 1-bit boolean value, combined with a MAC check bit (1 bit) to form a security control bit field
[0041] Extension mechanism: Support custom attribute extension through the type-length-value (TLV) format, such as adding a QoS priority to mark the message body block (binary data)
[0042] Support three data forms:
[0043] Original audio - video data: The RTP payload is directly encapsulated, preserving the timestamp and sequence number
[0044] Compressed data: When the compression flag in the metadata is 1, the LZ4 algorithm is used for compression, and the compression ratio can reach 5:1
[0045] Encrypted data: Encrypted in CBC mode based on the SM4 algorithm, and the IV vector is stored in a specific field of the metadata
[0046] II. Hardware - accelerated parsing mechanism
[0047] For the high - speed processing requirements of general data packets, a dedicated hardware parsing pipeline is designed:
[0048] Three - level parsing pipeline
[0049]
[0050] Optionally, the first audio - video front - and - rear device (A1) is also used to perform packet splitting, per - packet encryption, adding an encryption flag, and encapsulation on the first audio - video conference data in sequence to obtain the second audio - video conference data packet in ciphertext form to prevent the second audio - video conference data packet from being intercepted and cracked when transmitted in the non - trusted second network system.
[0051] Optionally, when the first audio - video front - and - rear device (A1) performs per - packet encryption on the data packets obtained by splitting the first audio - video conference data, multiple rounds of operations are performed based on FPGA to execute encryption in parallel using the encryption key. The round operations are based on the set key expansion parameters and round function parameters. The encryption key is generated based on the elliptic curve cryptosystem and can be updated according to the set update mechanism. When the key is rotated, a progressive switching strategy is adopted. First, the new - transmitted data packets are encrypted with the new key, and for the data packets that are already in the transmission process, the old key is still used for encryption. The encryption key can be configured as a shared key between the first audio - video front - and - rear device (A1) and the fourth audio - video front - and - rear device (B1).
[0052] Optionally, when the first audio - video front - and - rear device (A1) performs per - packet encryption on the data packets obtained by splitting the first audio - video conference data, a MAC value is generated based on the MAC algorithm that matches the encryption and the MAC value is appended to the tail of the encrypted data packet, so that the MAC value and the encrypted data packet are transmitted to the second audio - video front - and - rear device (A2) through the first security isolation and exchange device (U1). When the second audio - video front - and - rear device (A2) converts the second audio - video conference data packet into a general data packet, the MAC value is attached. The fourth audio - video front - and - rear device (B1) recalculates the MAC value after unpacking the encapsulated general data packet to obtain the general data packet and compares it with the attached MAC value to determine whether data tampering has occurred.
[0053] Optionally, when the first front and rear audio - video device (A1) splits the first audio - video conference data into data packets and encrypts each packet one by one, it uses the private key assigned to the first front and rear audio - video device (A1) to sign the encrypted data packet to generate signature information, and at the same time adds the timestamp information provided by the trusted timestamp service (TSA) to the encrypted data packet; the fourth front and rear audio - video device (B1) unpacks the encapsulated general data packet to obtain the general data packet, then uses the public key assigned to the first front and rear audio - video device (A1) to verify the signature information, and confirms the sending time of the data packet through the timestamp information.
[0054] In particular, in an application scenario, the specific implementation of the above - mentioned solution is as follows:
[0055] 1. Parallel encryption round function based on FPGA
[0056] Formula:
[0057] Where:
[0058] S i : The encryption state of the i - th round, representing the intermediate result of the segmented encryption of the current data packet to be processed
[0059] K i : The sub - key of the i - th round, derived from the master key generated by the elliptic curve cryptosystem through the key expansion algorithm
[0060] P: Fixed permutation matrix, with the dimension matching the encryption block length (e.g., a 32×32 matrix for AES - 256), used to achieve data diffusion
[0061] C i : The constant of the i - th round, generated by the hardware random number generator, and dynamically updated in each round to enhance the avalanche effect
[0062] R i : The random number of the i - th round, associated with the key expansion parameter, used to resist differential cryptanalysis
[0063] F: Non - linear round function, adopting a combination structure of S - box substitution and P - box permutation, implemented in FPGA through look - up tables and parallel logic
[0064] This application follows the design of the substitution - permutation network (SPN) structure, and the core lies in realizing the dual security mechanisms of confusion and diffusion through three - round operations. FPGA utilizes its hardware parallel characteristics to perform operations on S i and K iThe XOR operation, the non-linear transformation of the F function, and the permutation operation of the P matrix are decomposed into independent logic units for parallel execution. Taking 256-bit encryption as an example, each round operation can be split into 16 16-bit parallel processing units, and 1 data block can be processed per clock cycle through a pipeline architecture. The key expansion parameters and round function parameters are dynamically loaded by the hardware configuration module, supporting instruction set adaptation for multiple cryptographic algorithms such as AES and SM4. When the first audio-visual front and rear devices (A1) encrypt each packet of the data packet, the FPGA achieves an encryption throughput of more than 10 Gbps through this formula, meeting the real-time requirements of 4K video conferencing.
[0065] 2. Elliptic Curve Key Generation and Update Formula and Technical Principle
[0066] Formula:
[0067] Where:
[0068] K new : The newly generated encryption key, with a length matching the elliptic curve parameters (e.g., 48 bytes for the P-384 curve)
[0069] K old : The old encryption key, coexisting with the new key in the transmission link through a progressive switching strategy
[0070] d: The elliptic curve private key parameter, satisfying d ∈ [1, n - 1] (n is the curve order)
[0071] H: The SHA-3-512 hash function, outputting a 64-byte hash value
[0072] t: The UTC time provided by the trusted timestamp service (TSA), accurate to the millisecond level
[0073] ID: The device unique identifier, composed of a 48-bit MAC address and a 32-bit serial number
[0074] r: A 256-bit hardware true random number, generated through a thermal noise source
[0075] This application is based on the elliptic curve discrete logarithm problem (ECDLP), and the core operation K old ·d represents the point multiplication operation on the elliptic curve. In the prime field GF(p), let the elliptic curve equation be y 2 =x 3 +ax + b modp, K oldCorresponding to the public key point Q, the result of the dot product operation is d·Q, and its inverse operation is computationally infeasible. The hash function H is used to mix the timestamp, device ID, and random number to generate a key update factor to prevent key prediction attacks. When the first front and rear audio-visual device (A1) and the fourth front and rear audio-visual device (B1) perform key rotation, a progressive switching strategy is adopted: within the first 10 seconds, the usage probability of the new key increases from 10% to 100% according to the logistic function, and the transmitted data packets continue to be decrypted using the old key to ensure that the link is not interrupted. The key update mechanism supports two triggering modes: periodic update (default every 15 minutes) or based on the number of data packets (every 10,000 data packets), which is dynamically configured by the security policy module.
[0076] 3. Progressive Key Switching Probability Model
[0077] Formula:
[0078] Where:
[0079] P new (t): The probability of using the new key at time t, with a value range of [0,1]
[0080] γ: The switching rate parameter, dynamically adjusted by network latency (γ = 0.5 when the latency ≤ 50ms, and γ = 0.1 when the latency > 200ms)
[0081] t switch : The starting time point of key switching, determined by synchronizing the system clock with the TSA timestamp
[0082] Technical principle: This model uses the logistic function to achieve a smooth transition of key switching and avoid packet decryption failures caused by traditional hard switching. Mathematically, the slope of the function curve is the largest at t = t switch At this point, the usage probabilities of the old and new keys are each 50%. When implemented in hardware, the key management module of the A1 device maintains two key registers (the old key K old and the new key K new ), calculates P switch (t) by comparing the system time with t new , and compares it with a random number between 0 and 1 output by the hardware random number generator to determine the key used for the current data packet. When t - t switch > 10γ -1 , P new (t)> 0.99, and it is approximately considered that the full-scale switching is completed. This mechanism can control the packet loss rate caused by switching to less than 0.01% under a 100Mbps bandwidth, meeting the real-time requirements of video conferencing.
[0083] 4. MAC Value Generation and Verification
[0084] Formula: MAC = H(K MAC ‖P‖H(K MAC ‖P‖IV))
[0085] Where: MAC: Message Authentication Code, with a length consistent with the output of the hash function (e.g., 32 bytes for SHA-256)
[0086] KMAC: MAC generation key, associated with the encryption key through a key derivation function (KDF)
[0087] P: Encrypted data packet payload, containing the RTP header and audio / video data
[0088] IV: 128-bit initialization vector, generated by the counter mode (CTR) and incremented for each data packet
[0089] H: SHA-256 hash function, following the FIPS180-4 standard
[0090] This application implements a variant of the HMAC algorithm, enhancing collision resistance through double hashing. The first layer of hashing H(K MAC ‖P‖IV) is used to generate the content digest of the data packet. The second layer of hashing remixes this digest with the key and the data packet again to ensure a strong binding between the MAC value and the data packet content and the key. In the A1 device, the MAC calculation module and the encryption module work in parallel through a pipeline architecture: when the encrypted data packet enters the MAC calculation unit, it is first concatenated with K MAC and IV, generates an intermediate digest through a 256-bit hash circuit, and then is concatenated with the original input again for the second hashing, finally generating a 32-byte MAC value appended to the end of the data packet. After receiving the data packet, the B1 device recalculates the MAC value using the same K MAC and compares it with the received MAC value through an exclusive-OR comparator. If the number of different bits exceeds 3, it is determined that data has been tampered with, and the retransmission mechanism is triggered. This scheme can resist length extension attacks, and the MAC calculation latency is less than 10 microseconds in a 1Gbps network environment.
[0091] 6. Timestamp Signature Verification
[0092] Formula: Verify = (Sig = E pub (H(P‖TS))) ∧ (|t now -TS| < Δt)
[0093] Where:
[0094] Verify: Verification result, boolean value (true / false)
[0095] Sig: Received digital signature, with a length matching the elliptic curve signature algorithm (e.g., 64 bytes for ECDSA-P-256)
[0096] E pub : Elliptic curve public key verification function, with the hash value as the input and whether the signature is valid as the output
[0097] H: SHA-256 hash function, used to generate the content summary of the data packet
[0098] P: The complete content of the encrypted data packet (including the MAC value)
[0099] TS: The UTC time returned by the trusted timestamp service, in the format of ISO 8601
[0100] t now : The current time synchronized by the receiving device through NTP
[0101] Δt: Time tolerance, default set to 300 seconds (5 minutes)
[0102] This application ensures the authenticity and timeliness of the data packet through a dual verification mechanism. The signature verification part is based on the ECDSA algorithm. Device A1 uses the private key to sign H(P‖TS), and device B1 uses the public key to verify the correctness of the signature to ensure that the data packet has not been tampered with and comes from a legitimate device. The timestamp verification prevents replay attacks (such as an attacker intercepting an old data packet and resending it) by comparing the difference between the receiving time and the timestamp. In implementation, the timestamp service module interacts with the TSA server through the HTTPS protocol to obtain a timestamp certificate signed by the CA. This certificate contains the timestamp value, serial number, and certificate validity period. Device B1 maintains a time deviation table to record the clock offset with each TSA server for correcting the comparison error between t and TS. When the verification fails, the system marks the data packet as suspicious and triggers the key update process to prevent potential man-in-the-middle attacks. now The comparison error with TS. When the verification fails, the system marks the data packet as suspicious and triggers the key update process to prevent potential man-in-the-middle attacks.
[0103] 7. Trigger conditions for multi-factor key update
[0104] Formula: Trigger = (N > N thresh ) ∨ (Δt > T thresh ) ∨ (σ(K) > σ thresh )
[0105] Where:
[0106] Trigger: Key update trigger flag, boolean value (true / false)
[0107] N: The number of encrypted data packets since the last key update, real-time counted by the hardware counter
[0108] N thresh: Threshold of the number of data packets, default set to 10,000 (configurable range: 1,000 - 100,000)
[0109] Δt: Time interval since the last key update, accurately measured by the system timer
[0110] T thresh : Threshold of the time interval, default 15 minutes (configurable range: 5 - 60 minutes)
[0111] σ(K): Standard deviation of the key usage frequency, calculated through a sliding window (window size: 500 data packets)
[0112] σ thresh : Threshold of the standard deviation, default 0.3 (indicating that the usage frequency fluctuation exceeds 30% of the average value)
[0113] This application adopts a multi - dimensional trigger mechanism to balance security and system overhead. The threshold of the number of data packets N thresh Based on Shannon's cryptography theory, when the amount of encrypted data exceeds 2 n / 2 (n is the key length), the probability of brute - force cracking increases significantly. Therefore, for a 256 - bit key, the default threshold is set to 10,000 data packets (each data packet is about 1 KB on average, and the total data volume is about 10 MB, far lower than the security boundary of 2 128 . The threshold of the time interval T thresh Prevents statistical analysis attacks caused by long - term use of the same key and adjusts dynamically according to the network environment (for example, for satellite links with high latency, it can be set to 60 minutes). The standard deviation σ(K) is used to detect abnormal key usage patterns. When the key usage frequency suddenly increases or decreases (such as suffering from a DDoS attack or key leakage), an emergency update is triggered. In terms of hardware implementation, the key management module maintains three independent trigger detection units: The data packet counter uses a 64 - bit unsigned integer, the time interval is measured by a high - precision timer (resolution 1 ms), and the frequency standard deviation is calculated in real - time through a streaming statistics module in the FPGA. When any condition is met, the elliptic curve key update process is triggered, and a key update notification is sent to the peer device through the control channel.
[0114] Optionally, the third audio - video front - and - rear device (B2) is also used to de - encapsulate the encapsulated general data packets based on the TCP protocol to obtain general data packets and transmit them to the fourth audio - video front - and - rear device (B1) through the second security isolation and exchange device (U2). The fourth audio - video front - and - rear device (B1) is also used to de - encapsulate the encapsulated general data packets and then decrypt them to obtain general data packets in plaintext form, and restore the general data packets in plaintext form to the first audio data packets.
[0115] Optionally, when the fourth front and rear audio - video device (B1) decrypts the encapsulated general data packet after decompression, it performs multiple rounds of operations based on FPGA to execute decryption in parallel using the decryption key. The round operations are based on the set key expansion parameters and round function parameters. The decryption key is generated based on the elliptic curve cryptosystem and can be updated according to the set update mechanism. When rotating the key, a progressive switching strategy is adopted. First, the newly transmitted data packets are decrypted using the new key, and for the data packets that are already in the transmission process, the old key is continued to be used for decryption. The receiving key can be configured as a shared key between the first front and rear audio - video device (A1) and the fourth front and rear audio - video device (B1).
[0116] Optionally, when the fourth front and rear audio - video device (B1) decompresses the encapsulated general data packet to obtain the general data packet, it strips the general protocol layer information to extract the general metadata therefrom, and injects audio - video protocol layer information into the general metadata to restore the general data packet into the first audio data packet.
[0117] Optionally, when the second front and rear audio - video device (A2) converts the second audio - video conference data packet into a general data packet, it strips the audio - video protocol layer information therefrom to extract the core audio - video information and adds it as audio - video metadata, and encapsulates the audio - video metadata based on the format of the general data packet to form a general data packet.
[0118] Particularly, in a specific application scenario, the specific technical implementation of the above - mentioned solution is as follows:
[0119] 1. Parallel encryption / decryption round function based on FPGA
[0120]
[0121] Where: S i : Encryption / decryption state of the i - th round (128 - bit data block), K i : Sub - key of the i - th round (generated by expanding the master key), P: Fixed permutation matrix (16×16 bytes, realizing data diffusion), C i : Constant of the i - th round (enhancing the avalanche effect), R i : Random number of the i - th round (resisting differential attacks), θ i : Horizontal angle of the camera corresponding to the i - th data packet (- 180° to 180°), Γ(θ i ), Λ(θ i ), Ω(θ i ): Angle modulation functions, defined as:
[0122]
[0123] where α, β, γ are modulation coefficients (default 0.5), is matrix dot product, and M1, M2, M3 are angular perturbation matrices.
[0124] This application introduces angular modulation based on the traditional SPN network, and converts the camera angle information into the dynamic offset of the encryption parameter through trigonometric functions. The FPGA quickly calculates the angular modulation function through a look-up table, and dynamically adjusts the key and state parameters in each encryption round, so that adjacent data packets will generate different ciphertexts even if their contents are the same. For example, when the camera angle is facing the center (θ = 0°), Λ(θ i ) takes the maximum value, enhancing the confusion degree of the data in the central area and improving the security of key areas such as faces.
[0125] 2. Elliptic Curve Key Generation
[0126]
[0127] where: K new / K old : new and old encryption keys (256 bits), d: elliptic curve private key (d ∈ [1, n - 1], n is the curve order), H: SHA-3-256 hash function, t: timestamp (UTC millisecond level), ID: device identifier (64 bits), r: random number (256 bits), θ avg : average angle of the last 10 data packets, ΔK(θ avg ): angular derived key increment, defined as:
[0128]
[0129] where G is the elliptic curve base point, θ var is the angular variance, and S(θ var ) is a smoothing function:
[0130]
[0131] This application incorporates angular information into the key update process. ΔK(θ avg ) generates an angular-related key increment through elliptic curve dot product, so that the key is dynamically adjusted according to the camera angle. When the angle changes drastically (θ var is large), S(θ var ) approaches 1, enhancing the key change amplitude; otherwise, it remains stable. This design expands the key space to 2 256 ×360, effectively resisting time-based key prediction attacks.
[0132] 3. Progressive Key Switching Probability Model
[0133]
[0134] Where: P new (t): Probability of using a new key at time t, γ(θ rate ): Switching rate parameter modulated by the angle change rate:
[0135]
[0136] Where θ rate is the angle change rate (° / second), γ0 is the base rate (default 0.5), and δ is the modulation factor (default 0.3)
[0137] t switch : Starting time of key switching.
[0138] In this application, considering that when the camera rotates rapidly (θ rate is large), γ(θ rate ) increases, accelerating the key switching to ensure security in dynamic scenarios; when the angle is stable, the switching process slows down to reduce transmission interruptions caused by key updates. The FPGA calculates θ rate in real time through a Kalman filter and dynamically adjusts the slope of the logistic curve.
[0139] 4. MAC Value Generation and Verification
[0140] MAC = H(K MAC ‖P‖H(K MAC ‖P‖IV‖Θ))
[0141] K MAC : MAC key (128 bits) P: Encrypted data packet IV: Initialization vector (96 bits) Θ: Angle authentication code, defined as: Θ = H(θ curr ‖θ prev ‖θ pred ‖K θ ) Where θ curr is the current angle, θ prev is the previous frame angle, θ pred is the predicted angle (calculated through the autoregressive model AR(3)), and K θ is the angle authentication key
[0142] This application embeds the hash value of the angle sequence into the MAC calculation to ensure the integrity of the angle information. The receiving end verifies whether the angle has been tampered with by comparing the calculated Θ with the received Θ. If an attacker attempts to modify the video angle, even if the encrypted content remains unchanged, the MAC verification will fail, effectively preventing view spoofing attacks.
[0143] 5. Timestamp Signature Verification
[0144] Verify = (Sig = Epub (H(P‖TS‖θ TS )))^(|t now -TS|<Δt)^(|θ curr -θ TS |<Δθ)
[0145] Where: Sig: Digital signature (ECDSA - 256), E pub : Public key verification function, TS: Timestamp (UTC milliseconds), θ TS : Angle corresponding to the timestamp, θ curr : Currently received angle, Δt: Time tolerance (default 300ms), Δθ: Angle tolerance (default 5°).
[0146] This application adds angle consistency verification on the basis of traditional signature verification. The sender embeds the angle θ corresponding to the timestamp into the signature, and the receiver verifies whether the difference between the current angle and θ is within a reasonable range. If the angle mutation exceeds the threshold (such as Δθ), it is determined as a suspicious data packet, and an alarm will be triggered even if the timestamp and signature are valid. TS embedded in the signature, and the receiver verifies whether the difference between the current angle and θ TS is within a reasonable range. If the angle mutation exceeds the threshold (such as Δθ), it is determined as a suspicious data packet, and an alarm will be triggered even if the timestamp and signature are valid.
[0147] 6. Trigger condition formula for multi - factor key update (including angle entropy detection)
[0148] Trigger=(N>N thresh )∨(Δt>T thresh )∨(σ(K)>σ thresh )∨(H(Θ)<H thresh )
[0149] Where: H(Θ): Information entropy of the angle sequence Θ, calculated as:
[0150]
[0151] where p(θ i ) is the occurrence probability of the angle θ i
[0152] H thresh : Entropy threshold (default 0.8 bits / angle)
[0153] In this application, when the entropy value of the angle sequence is too low (such as when the camera is stationary for a long time), it indicates that the key usage pattern may become predictable. At this time, the key is updated even if other conditions (such as time, number of data packets) are not met. This design enables the system to adapt to different usage scenarios and achieve a dynamic balance between security and performance.
[0154] 7. Protocol conversion formula for angle perception (video frame reconstruction)
[0155]
[0156] Among them: F reconstructed : The reconstructed video frame, F i : The video frame of the i-th shard, w i (θ): Angle weight function:
[0157]
[0158] θ: The current camera angle, θ i : The optimal viewing angle of the i-th shard, α, β: Decay coefficients (default 0.01 and 0.1).
[0159] In this application, during the protocol conversion process, different-view video shards are dynamically weighted and merged according to the current camera angle. When the camera turns to a certain direction, the weight of the video shard in that direction increases, improving the visual quality. The B1 device intelligently selects the optimal shard combination through angle information to achieve seamless viewing angle switching.
[0160] 8. Angle-assisted TCP congestion control formula
[0161]
[0162] Among them: cwnd: Congestion window size, cwnd base : Base window size, θ rate : Angle change rate, θ rate_max : Maximum angle change rate (default 100° / second), η: Angle adjustment factor (default 0.3), RTT: Round-trip delay, RTT min : Minimum round-trip delay λ: Delay sensitivity (default 0.5).
[0163] When the camera rotates rapidly (θ rate is high), increase the congestion window to preferentially transmit data in the area of view change and reduce the sense of lag; when the angle is stable, return to the traditional congestion control strategy. This design enables the video conferencing system to more efficiently utilize network bandwidth in dynamic scenarios.
[0164] 10. Angle and audio synchronization
[0165] τ sync = τ audio + φ(θ) + ψ(θ rate )
[0166] Among them: τ sync : Synchronization timestamp, τ audio : Audio timestamp, φ(θ): Angle-related synchronization offset:
[0167]
[0168] ψ(θ rate ): Dynamic compensation related to the rate of change of angle:
[0169] A, B, C: Calibration parameters (default 5ms, 3ms, 50° / second), φ0: Phase offset (default π / 4)
[0170] In this application, the audio playback time is adjusted according to the camera angle to simulate the time difference of arrival caused by the change of the sound source direction in the real world. When the camera turns to the left, the audio of the left sound source is played in advance to enhance the sense of spatial immersion. The greater the rate of change of the angle, the greater the amount of dynamic compensation, ensuring the auditory coherence during rapid rotation.
[0171] Supported by the above embodiments, the embodiments of this application also provide an audio-visual cross-network conferencing system based on secure data exchange, which includes at least three security domains, and each security domain is configured with a terminal device and an edge processing device;
[0172] The first terminal in the first security domain generates multimedia data, which is subjected to protocol adaptation and security enhancement by the first edge processing device in the same domain;
[0173] The processed data is transmitted to the untrusted second security domain through the first isolation and exchange device, and is subjected to protocol conversion and encapsulation by the second edge processing device in this domain;
[0174] The encapsulated data is transmitted to the third edge processing device in the second security domain through the network layer protocol, and after being unpacked and format restored, it is transmitted to the third security domain through the second isolation and exchange device;
[0175] The fourth edge processing device in the third security domain performs secondary unpacking and protocol restoration on the data, and is received and presented by the second terminal in the same domain;
[0176] Each edge processing device supports dynamic key management, multi-round encryption operations, and cross-domain identity authentication. The isolation and exchange device uses physical isolation and data ferry mechanisms to ensure transmission security. The system realizes real-time multimedia communication across security domains through multi-level protocol conversion and content reconstruction.
[0177] It should be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.
[0178] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An audio and video cross-network conferencing system based on secure data exchange, characterized in that, Including: The first terminal (1), the first audio-video front and rear equipment (A1), the first security isolation and exchange device (U1), the second audio-video front and rear equipment (A2), the third audio-video front and rear equipment (B2), the second security isolation and exchange device (U2), the fourth audio-video front and rear equipment (B1), and the second terminal (2). The first terminal (1) and the first audio-video front and rear equipment (A1) are in the first network system. The second audio-video front and rear equipment (A2) and the third audio-video front and rear equipment (B2) are in the untrusted second network system. The fourth audio-video front and rear equipment (B1) and the second terminal (2) are in the third network system. Among them: The first terminal (1) is used to generate the first audio-video conference data and transmit it to the first audio-video front and rear equipment (A1); The first audio-video front and rear equipment (A1) is used to split the first audio-video conference data into packets, encapsulate them to obtain the second audio-video conference data packet, and transmit it to the second audio-video front and rear equipment (A2) through the first security isolation and exchange device (U1); The second audio-video front and rear equipment (A2) is used to convert the second audio-video conference data packet into a general data packet, encapsulate the general data packet using the TCP protocol to obtain the encapsulated general data packet, and send it to the third audio-video front and rear equipment (B2); The third audio-video front and rear equipment (B2) is used to receive the encapsulated general data packet and restore the general data packet to transmit it to the fourth audio-video front and rear equipment (B1) through the second security isolation and exchange device (U2); The fourth audio-video front and rear equipment (B1) is used to de-encapsulate the encapsulated general data packet to obtain the general data packet, and restore the general data packet to the first audio data packet; The second terminal (2) receives the restored first audio data packet and plays it to conduct a cross-network conference between the first terminal (1) and the second terminal (2).
2. The system according to claim 1, wherein The first audio-video front and rear equipment (A1) is also used to split the first audio-video conference data into packets in sequence, encrypt each packet one by one, add an encryption flag, and encapsulate it to obtain the second audio-video conference data packet in ciphertext form to prevent the second audio-video conference data packet from being intercepted and cracked when transmitted in the untrusted second network system.
3. The system according to claim 2, characterized in that, When the first audio-video front and rear equipment (A1) splits the first audio-video conference data into packets and encrypts each packet one by one, it performs multiple rounds of operations based on the FPGA to execute encryption in parallel using the encryption key. The round operations are based on the set key expansion parameters and round function parameters. The encryption key is generated based on the elliptic curve cryptosystem and can be updated according to the set update mechanism. When the key is rotated, a progressive switching strategy is adopted. First, the newly transmitted data packets are encrypted using the new key. For the data packets that are already in the transmission process, they continue to be encrypted using the old key. The encryption key can be configured as a shared key between the first audio-video front and rear equipment (A1) and the fourth audio-video front and rear equipment (B1).
4. The system according to claim 2, wherein When the first audio - video front - and - rear device (A1) divides the first audio - video conference data into packets and encrypts each packet one by one, it generates a MAC value based on the MAC algorithm that matches the encryption and attaches the MAC value to the tail of the encrypted packet, so that the MAC value and the encrypted packet are transmitted through the first security isolation and exchange device (U1) to the second audio - video front - and - rear device (A2), so that when the second audio - video front - and - rear device (A2) converts the second audio - video conference data packet into a general data packet, the MAC value is attached. The fourth audio - video front - and - rear device (B1) unpacks the encapsulated general data packet to obtain the general data packet and then recalculates the MAC value and compares it with the attached MAC value to determine whether data tampering has occurred.
5. The system according to claim 2, wherein When the first audio - video front - and - rear device (A1) divides the first audio - video conference data into packets and encrypts each packet one by one, it uses the private key assigned to the first audio - video front - and - rear device (A1) to sign the encrypted packet to generate signature information, and at the same time adds the timestamp information provided by the trusted timestamp service (TSA) to the encrypted packet; the fourth audio - video front - and - rear device (B1) unpacks the encapsulated general data packet to obtain the general data packet and then uses the public key assigned to the first audio - video front - and - rear device (A1) to verify the signature information and confirm the sending time of the data packet through the timestamp information.
6. The system according to claim 1, wherein The third audio - video front - and - rear device (B2) is also used to unpack the encapsulated general data packet based on the TCP protocol to obtain the general data packet and transmit it to the fourth audio - video front - and - rear device (B1) through the second security isolation and exchange device (U2). The fourth audio - video front - and - rear device (B1) is also used to unpack the encapsulated general data packet and then decrypt it to obtain the general data packet in plaintext form and restore the general data packet in plaintext form to the first audio data packet.
7. The system according to claim 4, characterized in that, When the fourth audio - video front - and - rear device (B1) unpacks and then decrypts the encapsulated general data packet, it performs multiple rounds of operations based on FPGA to execute decryption in parallel using the decryption key. The round operations are based on the set key expansion parameters and round function parameters. The decryption key is generated based on the elliptic curve cryptosystem and can be updated according to the set update mechanism. When the key is rotated, a progressive switching strategy is adopted. First, the newly transmitted data packet is decrypted using the new key, and for the data packets that are already in the transmission process, the old key is continued to be used for decryption. The receiving key can be configured as a shared key between the first audio - video front - and - rear device (A1) and the fourth audio - video front - and - rear device (B1).
8. The system according to claim 1, characterized in that, The general data packet includes MagicNumber, metadata length, message body length, metadata, and message body. Among them, MagicNumber is used to distinguish the type of data packet; the message body length is used to record the lengths of the metadata and the message body. The metadata is a record of the message attributes; the message body is the actual message data generated.
9. The system according to claim 1, wherein When the second audio-video front and rear device (A2) converts the second audio-video conference data packet into a general data packet, it strips the audio-video protocol layer information therein to extract the core audio-video information and adds it as audio-video metadata, and encapsulates the audio-video metadata based on the format of the general data packet to form a general data packet.
10. An audio-video cross-network conference system based on secure data exchange, characterized in that: It includes at least three security domains, and each security domain is configured with a terminal device and an edge processing device; The first terminal in the first security domain generates multimedia data, and the first edge processing device in the same domain performs protocol adaptation and security enhancement; The processed data is transmitted to the untrusted second security domain through the first isolation exchange device, and the second edge processing device in this domain performs protocol conversion and encapsulation; The encapsulated data is transmitted to the third edge processing device in the second security domain through the network layer protocol, and after de-encapsulation and format restoration, it is transmitted to the third security domain through the second isolation exchange device; The fourth edge processing device in the third security domain performs secondary de-encapsulation and protocol restoration on the data, and is received and presented by the second terminal in the same domain; Each edge processing device supports dynamic key management, multi-round encryption operations, and cross-domain identity authentication. The isolation exchange device uses physical isolation and data ferry mechanisms to ensure transmission security. The system realizes real-time multimedia communication across security domains through multi-level protocol conversion and content reconstruction.