Audio and video transmission method and device based on multi-modal AI enhancement and dynamic end-to-end encryption, equipment and storage medium

Through the audio and video transmission method of multimodal AI enhancement and dynamic end-to-end encryption, the problems of low efficiency and insufficient security of low-quality audio and video transmission are solved, and efficient and secure audio and video data transmission is achieved.

CN120751177APending Publication Date: 2025-10-03SHENZHEN JIUZHOU ELECTRIC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510824909.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional set-top boxes lack the ability to repair low-resolution, high-noise audio and video in real time. Fixed encryption algorithms are prone to delays or packet loss when the network fluctuates, and multimodal collaborative processing causes waste of computing power and security vulnerabilities.

Method used

Multimodal AI enhancement technology is used to process audio, video, and subtitle text data, combined with dynamic end-to-end encryption, to generate encryption decision instructions based on real-time network status, add tamper-proof watermarks, and transmit data through secure transmission channels.

Benefits of technology

It improves the transmission efficiency and security of low-quality audio and video, dynamically adapts to network conditions, reduces delays and resource waste, and enhances copyright protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751177A_ABST
    Figure CN120751177A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video transmission method and device based on multi-mode AI enhancement and dynamic end-to-end encryption, equipment and a storage medium, and relates to the technical field of digital audio and video processing and secure transmission, and the method comprises the steps: obtaining an original audio and video stream, and separating the original audio and video stream into audio data, video data and subtitle text data; performing multi-mode AI enhancement processing on the audio data, the video data and the subtitle text data to obtain enhanced audio and video data; generating an encryption decision instruction based on the real-time network state, and performing encryption processing on the enhanced audio and video data according to the encryption decision instruction to generate encrypted data with a tamper-proof watermark; and when the encrypted data passes integrity verification and copyright legality judgment, the encrypted data is transmitted through a secure transmission channel, and collaborative optimization of security and efficiency is realized through a core technology chain of multi-modal separation enhancement, dynamic encryption decision, hardware-level watermark binding and credible verification transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital audio and video processing and secure transmission technology, and in particular to an audio and video transmission method and apparatus, device, and storage medium based on multimodal AI enhancement and dynamic end-to-end encryption. Background Art

[0002] Traditional set-top boxes lack the ability to repair low-resolution, high-noise audio and video in real time, and cannot adequately process low-quality content, resulting in a poor user experience.

[0003] Existing fixed encryption algorithms are rigid. For example, encryption schemes (such as AES-256) are prone to delays or packet loss during network fluctuations and cannot dynamically adapt to transmission conditions. Regarding multimodal collaboration, the AI ​​processing and encryption processes for voice, images, and text are independent of each other, resulting in wasted computing power and security vulnerabilities.

[0004] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of the present invention is to provide an audio and video transmission method and device, equipment and storage medium based on multimodal AI enhancement and dynamic end-to-end encryption, aiming to solve the technical problems of low efficiency and insufficient security of low-quality audio and video transmission.

[0006] To achieve the above objectives, the present invention provides an audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption, the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption comprising the following steps:

[0007] Acquire an original audio and video stream, and separate the original audio and video stream into audio data, video data, and subtitle text data;

[0008] Performing multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data;

[0009] Generate an encryption decision instruction based on the real-time network status, encrypt the enhanced audio and video data according to the encryption decision instruction, and generate encrypted data with an anti-tampering watermark;

[0010] When the encrypted data passes the integrity verification and copyright legality determination, the encrypted data is transmitted through a secure transmission channel.

[0011] In one embodiment, the step of performing multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data includes:

[0012] Converting low-resolution video frames of the video data into high-resolution video frames through an artificial intelligence image processing model;

[0013] Eliminating background noise from the audio data using an audio processing algorithm to obtain noise-reduced audio data;

[0014] Identifying key content in the subtitle text data through natural language processing technology and generating metadata tags synchronized with the video timeline;

[0015] The high-resolution video frame, the noise-reduced audio data, and the metadata tag are recombined to obtain enhanced audio and video data.

[0016] In one embodiment, the step of generating an encryption decision instruction based on the real-time network status includes:

[0017] Real-time monitoring of network transmission parameters;

[0018] Inputting the network transmission parameters into a decision model for processing and analysis to obtain processing and analysis results;

[0019] generating, under high-quality network conditions, instructions using a high-security encryption algorithm based on the processing and analysis results;

[0020] An instruction for using a high-efficiency encryption algorithm is generated based on the processing and analysis results under low-quality network conditions.

[0021] In one embodiment, the step of encrypting the enhanced audio and video data according to the encryption decision instruction to generate encrypted data with an anti-tampering watermark includes:

[0022] Performing an encryption operation on the enhanced audio and video data to generate encrypted audio and video data;

[0023] Device-bound watermark information is added to the encrypted audio and video data to form an anti-tampering protection layer, thereby generating encrypted data with an anti-tampering watermark.

[0024] In one embodiment, before the step of transmitting the encrypted data through a secure transmission channel when the encrypted data passes the integrity verification and the copyright validity determination, the method further includes:

[0025] Extracting user identity information contained in the encrypted data with the tamper-proof watermark in a secure verification environment;

[0026] Matching the user identity information with the device registration information to obtain a user identity information matching result;

[0027] Perform integrity check on the received data and compare it with the original record to obtain the data integrity check result;

[0028] When the user identity information matching result is a successful match and the data integrity check result is a data integrity check passed, it is determined that the verification is passed.

[0029] In one embodiment, the method further comprises:

[0030] When any one of the user identity information matching result or the data integrity check result fails, the verification is determined to have failed, and the infringement handling operation is executed, the audio and video playback process is immediately terminated, the illegal use record is sent to the blockchain node, and a copyright violation warning message is displayed on the playback interface.

[0031] In one embodiment, when the encrypted data passes the integrity verification and the copyright validity determination, the step of transmitting the encrypted data through a secure transmission channel includes:

[0032] When the encrypted data passes the integrity verification and copyright validity determination, dividing the encrypted data into data units suitable for network transmission;

[0033] transmitting the data unit via an encrypted communication protocol and maintaining a forward security protection mechanism throughout the transmission process;

[0034] After the receiving end successfully receives all data units and confirms their integrity, the transmission process is completed.

[0035] In addition, to achieve the above objectives, the present invention also proposes an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, characterized in that the device includes:

[0036] A data acquisition module is used to acquire the original audio and video stream and separate the original audio and video stream into audio data, video data and subtitle text data;

[0037] An enhancement module, configured to perform multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data;

[0038] An encryption module, configured to generate an encryption decision instruction based on the real-time network status, and encrypt the enhanced audio and video data according to the encryption decision instruction to generate encrypted data with an anti-tampering watermark;

[0039] The transmission module is used to transmit the encrypted data through a secure transmission channel when the encrypted data passes the integrity verification and the copyright legality determination.

[0040] In addition, to achieve the above-mentioned purpose, the present invention also proposes an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, the device including: a memory, a processor, and an audio and video transmission program based on multimodal AI enhancement and dynamic end-to-end encryption stored on the memory and runnable on the processor, the audio and video transmission program based on multimodal AI enhancement and dynamic end-to-end encryption being configured to implement the steps of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption as described above.

[0041] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which is stored an audio and video transmission program based on multimodal AI enhancement and dynamic end-to-end encryption. When the audio and video transmission program based on multimodal AI enhancement and dynamic end-to-end encryption is executed by the processor, the steps of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption as described above are implemented.

[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption as described above.

[0043] One or more technical solutions proposed in this application have at least the following technical effects:

[0044] The field of digital audio and video processing and secure transmission technology includes: obtaining the original audio and video stream, separating the original audio and video stream into audio data, video data and subtitle text data; performing multimodal AI enhancement processing on the audio data, video data and subtitle text data to obtain enhanced audio and video data; generating encryption decision instructions based on real-time network status, encrypting the enhanced audio and video data according to the encryption decision instructions, and generating encrypted data with tamper-proof watermarks; when the encrypted data passes the integrity verification and copyright legitimacy judgment, the encrypted data is transmitted through a secure transmission channel, and the core technology chain of multimodal separation enhancement, dynamic encryption decision, hardware-level watermark binding and trusted verification transmission is used to achieve coordinated optimization of security and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 A flowchart illustrating the first embodiment of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application;

[0048] Figure 2 This is a system architecture diagram provided for the first embodiment of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption of this application;

[0049] Figure 3 A flow chart of the dynamic encryption switching algorithm provided in Example 1 of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption of this application;

[0050] Figure 4 This is a multimodal AI processing timing diagram provided in Example 1 of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption of this application;

[0051] Figure 5 This is a flowchart of the second embodiment of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application;

[0052] Figure 6 This is a schematic diagram of the module structure of an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption according to an embodiment of the present application;

[0053] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption in the embodiment of the present application.

[0054] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0055] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0056] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0057] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, etc. The following uses an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption as an example to illustrate this embodiment and the following embodiments.

[0058] Based on this, the embodiment of the present application provides an audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption of this application.

[0059] In this embodiment, the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption includes steps S10 to S40:

[0060] Step S10, obtaining the original audio and video stream, and separating the original audio and video stream into audio data, video data and subtitle text data;

[0061] In the specific implementation, the original audio and video signal sources that have not been processed are collected. These audio and video streams may come from different channels, such as live streams received through network transmission protocols (such as RTMP, HLS, etc.) or unprocessed media files stored locally; then the mixed audio and video streams are accurately split into independent audio parts, video parts and subtitle text parts, preparing for subsequent enhanced processing of each part.

[0062] It should be noted that if Figure 2 As shown, the system architecture includes:

[0063] Multimodal AI processing module:

[0064] The input unit receives the original audio and video stream (supports RTMP, HLS and other protocols) and separates the audio, video and subtitle text data.

[0065] The AI ​​enhancement unit achieves real-time super-resolution (SR) and denoising (Deblur) based on a lightweight CNN model (such as the improved ESRGAN); uses the RNN noise suppression algorithm (DENOISE) to separate human voices from background sounds and optimize sound quality; and extracts subtitle keywords through the NLP model, matching them with the voice content to generate dynamic metadata.

[0066] The compression coding unit uses AI-optimized encoders (such as H.266 / VVC) to dynamically adjust the compression rate based on the complexity of the enhanced data.

[0067] Dynamic encryption transmission module:

[0068] The encryption policy controller monitors network bandwidth, latency, and packet loss rate in real time (predicting the status in the next 5 seconds through the MLP neural network).

[0069] High-bandwidth stable network enables AES-256+ECC hybrid encryption (video stream AES encryption, key ECC distribution).

[0070] Low-bandwidth / high-latency networks switch to ChaCha20-Poly1305 lightweight encryption to reduce computational overhead.

[0071] Hardware acceleration unit: Integrates NPU (neural network processor) and encryption ASIC chip to process AI enhancement and encryption tasks in parallel, reducing CPU load.

[0072] Secure channel management module:

[0073] Establish a two-way authenticated TLS1.3 channel and support forward security (PFS); verify the content hash value through the blockchain node to prevent tampering by the middleman during transmission.

[0074] Step S20: performing multimodal AI enhancement processing on the audio data, video data, and subtitle text data to obtain enhanced audio and video data;

[0075] In practice, AI technology is used to optimize the separated audio, video, and subtitle data, improving the overall quality of the audio and video content and the user experience. By processing audio, video, and text data separately, the quality of each modality can be targeted and then recombined into an enhanced audio and video content.

[0076] In a feasible implementation, step S20 includes steps A11 to A14:

[0077] A11: Converts low-resolution video frames of video data into high-resolution video frames through an AI image processing model;

[0078] In the specific implementation, super-resolution models based on deep learning (such as the improved ESRGAN) are used. These models are trained on a large amount of image data and can identify and enhance image details. The model can learn the mapping relationship between low-resolution images and high-resolution images;

[0079] The separated low-resolution video frames are fed into the model frame by frame, and the model outputs corresponding high-resolution video frames. During this process, the model intelligently fills in details based on the learned features, making the video clearer and sharper, and improving the visual quality. For example, when processing 480p video, it can be upscaled to 720p or higher resolution while maintaining the video's smoothness and frame rate.

[0080] A12: Use an audio processing algorithm to eliminate background noise from the audio data to obtain noise-reduced audio data.

[0081] It's important to note that the audio processing algorithm uses an RNN (recurrent neural network) noise suppression algorithm (such as the DENOISE algorithm), which analyzes the time-frequency characteristics of audio signals to distinguish between human voices and background noise. For example, the algorithm learns the frequency range and characteristic patterns of human voices, as well as the characteristics of background noise (such as ambient noise and equipment noise).

[0082] The noise reduction process feeds the separated audio data into an algorithm, which processes the audio in real time, reducing or removing background noise while preserving the clarity of the human voice. The processed audio is purer, improving the signal-to-noise ratio and audibility.

[0083] A13: Uses natural language processing technology to identify key content in subtitle text data and generate metadata tags synchronized with the video timeline;

[0084] It should be noted that natural language processing technology uses NLP models to perform semantic analysis on subtitle text, extracting keywords, phrases, and important information. For example, the model can identify entity names, action descriptions, or time expressions in subtitles.

[0085] Generating metadata tags involves generating metadata tags corresponding to timestamps for key content identified based on the video's timeline. These tags can be used to accurately display subtitles during video playback and for subsequent content retrieval and analysis. For example, key dialogue in subtitles can be identified and accurately timestamped for them, allowing the relevant subtitles to be displayed synchronously with video playback.

[0086] A14: Recombines high-resolution video frames, denoised audio data, and metadata tags to obtain enhanced audio and video data.

[0087] In its implementation, the processed high-resolution video frames, noise-reduced audio data, and metadata tags are reassembled according to audio and video format specifications and synchronization requirements. For example, video container formats (such as MP4 and MKV) are used to encapsulate the video, audio, and metadata together, ensuring their synchronization during playback. This integration significantly improves the quality of the audio and video data, providing users with a superior viewing experience.

[0088] Step S30: Generate an encryption decision instruction based on the real-time network status, encrypt the enhanced audio and video data according to the encryption decision instruction, and generate encrypted data with an anti-tampering watermark;

[0089] In the specific implementation, the encryption strategy is dynamically adjusted according to the network conditions to ensure the security and integrity of audio and video data during transmission, and an anti-tampering watermark is added to enhance copyright protection.

[0090] like Figure 3 As shown in the figure, the dynamic encryption switching process includes:

[0091] Collect network status parameters (bandwidth B, latency D, packet loss rate L); input a pre-trained MLP model (input layer 3 nodes, hidden layer 8 nodes, output layer 2 nodes), and output encryption mode scores (Score_AES, Score_ChaCha); if Score_AES > threshold α and B ≥ 10Mbps, enable AES-256; otherwise, enable ChaCha20.

[0092] Multimodal enhancement collaboration rules include: video enhancement priority > audio enhancement > text processing to ensure key image quality; encryption computing resource allocation and AI enhancement tasks share the NPU computing power pool, and conflicts are avoided through time-slice rotation scheduling.

[0093] This strategy deeply couples multimodal AI enhancement (video, audio, and text) with dynamic encryption strategies to achieve triple optimization of "quality, security, and efficiency." It also proposes an encryption decision model based on an MLP network, breaking through the limitations of traditional fixed threshold switching. It can be integrated into existing set-top box chipsets and integrated with mainstream DRM solutions (such as Widevine and FairPlay) via an SDK. At 2Mbps bandwidth, it improves PSNR by 4dB and reduces encryption latency by 22%.

[0094] In a feasible implementation, step S30 includes steps A21 to A24:

[0095] A21: Real-time monitoring of network transmission parameters;

[0096] It should be noted that network transmission parameters refer to the continuous collection of parameters such as network bandwidth, latency, and packet loss rate. These parameters reflect the current network transmission capacity and service quality. For example, information such as bandwidth utilization, packet transmission latency, and packet loss per second can be obtained through network monitoring tools or APIs.

[0097] Monitoring: During network transmission, these parameters are collected in real time to promptly understand changes in network status. For example, network bandwidth, latency, and packet loss rate data are collected every 100 milliseconds to provide an accurate basis for subsequent encryption decisions.

[0098] A22: Input the network transmission parameters into the decision model for processing and analysis to obtain processing and analysis results;

[0099] It should be noted that the decision model uses a multi-layer perceptron (MLP) neural network model based on machine learning. This model has been trained with a large amount of network status data and encryption requirement data, and can predict the appropriate encryption mode based on the input network parameters.

[0100] In the specific implementation, the network transmission parameters (such as bandwidth, delay, and packet loss rate) monitored in real time are normalized and input into the decision model, and the model outputs the encryption mode score.

[0101] A23: Generates instructions using a high-security encryption algorithm based on the processing and analysis results under high-quality network conditions;

[0102] In a specific implementation, when the network bandwidth is sufficient (e.g., B ≥ 10 Mbps), the delay is low, and the packet loss rate is small, the network is considered to be in a high-quality state.

[0103] Based on the score output by the decision model and network conditions, an instruction is generated to select a high-security encryption algorithm (such as AES-256).

[0104] A24: Generates instructions using a high-efficiency encryption algorithm based on processing and analysis results under low-quality network conditions.

[0105] In a specific implementation, when the network bandwidth is limited (e.g., B<3Mbps), the delay is high (e.g., D>200ms), and the packet loss rate is high, the network is considered to be in a low-quality state.

[0106] Generate instructions to select a high-efficiency encryption algorithm based on the decision model's score and network conditions.

[0107] In a feasible implementation, step S30 further includes steps A31 to A32:

[0108] A31: Perform encryption processing on the enhanced audio and video data to generate encrypted audio and video data;

[0109] In a specific implementation, the enhanced audio and video data is encrypted and converted according to the encryption algorithm selected by the encryption decision instruction. For example, the audio and video data is encrypted using an encryption algorithm, and the original data is converted into ciphertext to ensure the confidentiality of the data during transmission.

[0110] A32: Add device-bound watermark information to the encrypted audio and video data to form an anti-tampering protection layer and generate encrypted data with an anti-tampering watermark.

[0111] It should be noted that watermark information containing the unique identifier of the device (such as the device hardware ID) is generated and embedded into the encrypted data using a digital watermark embedding algorithm. For example, using watermarking technology based on pixel-level micro-shifting, information such as the device ID and timestamp is embedded into the video image to form an anti-tampering protection layer.

[0112] The watermark information is tightly coupled with the encrypted data, and any tampering with the data will destroy the watermark, allowing detection of illegal data modifications. For example, if the watermark information in the encrypted data is inconsistent with the device binding information, it is determined that the data may have been tampered with, triggering the corresponding security mechanism.

[0113] Step S40: When the encrypted data passes the integrity verification and copyright validity determination, the encrypted data is transmitted through a secure transmission channel.

[0114] In a specific implementation, before transmission, a hash value is calculated for the encrypted data using a hash algorithm (such as SHA-256) and compared with the hash value provided by the sender to ensure that the data has not been tampered with during processing and encryption. For example, the SHA-256 hash value of the encrypted data is calculated and verified to be consistent with the hash value generated by the sender.

[0115] Verify the copyright information of encrypted data to ensure that the data comes from a legitimate source and has the appropriate authorization. For example, check whether the data complies with the requirements of the DRM (Digital Rights Management) scheme and verify the relevant copyright certificates and authorization information.

[0116] Once the encrypted data passes integrity verification and copyright validation, it is transmitted through a secure transmission channel. For example, TLS 1.3 is used to encrypt data transmission, preventing man-in-the-middle attacks and data leaks, ensuring that audio and video data is transmitted securely and reliably to its destination.

[0117] One implementation of this strategy is about live streaming in a low-bandwidth environment, such as live broadcasting of sports events over a 4G / 5G mobile network. When the network bandwidth fluctuates dramatically (from 5Mbps to 2Mbps) and content transmission security needs to be guaranteed, the system will perform the following operations: First, multimodal AI enhancement processing is performed. For video, the input 480p video stream (H.264 encoding, bit rate 1.5Mbps) is upgraded to 720p resolution in real time using a lightweight ESRGAN model. At the same time, the AI ​​encoder (based on H.266) reduces the output bit rate to 1.0Mbps, with a compression rate of approximately 33%, while maintaining a PSNR (peak signal-to-noise ratio) of no less than 38dB; for audio, background noise and human voice are separated, and the signal-to-noise ratio is increased from 15dB to 25dB through the RNN denoising model; for text processing, key timestamps in the live subtitles are extracted and dynamic metadata is generated and embedded in the video stream. Dynamic encryption switching is then performed, collecting bandwidth, latency, and packet loss rate data every 100ms and feeding it into a pre-trained MLP model (the input layer is the normalized B / D / L value, and the output layer is the AES / ChaCha20 switching probability). When the bandwidth is less than 3Mbps and the latency exceeds 200ms, the ChaCha20-Poly1305 algorithm is enabled to reduce encryption latency (measured from 15ms to 6ms). In terms of hardware acceleration, the NPU allocates 70% of computing power to the ESRGAN model and 30% to ChaCha20 encryption. The ASIC chip accelerates key generation to ensure smooth video. Finally, secure transmission verification is performed. The encrypted data stream is transmitted through a TLS1.3 channel and a SHA-256-based blockchain hash value is attached to ensure content integrity. When multimodal processing starts (0-10ms), the video, text, and audio threads are processed in parallel. The video thread receives the video stream and starts the ESRGAN super-resolution model in 0ms, continuing processing until 50ms to complete. At 25ms, keyframes are detected, triggering dynamic HDR to SDR mapping. The text thread starts subtitle keyword extraction in 5ms and completes metadata generation in 15ms. The audio thread starts RNN noise reduction processing in 10ms and ends in 40ms. During this period, a bandwidth drop to 3Mbps is detected in 20ms, triggering a switch in the encryption algorithm. Data synchronization and encryption are performed in 30-50ms. The text thread completes subtitle synchronization in 30ms and releases resources to the compression and encoding module. After the video and audio threads are fully completed in 50ms, the dynamic encryption process begins. In terms of hardware acceleration collaboration, the NPU allocates computing power according to priority within 0-50ms, with video enhancement accounting for 70%, audio noise reduction accounting for 20%, and the remaining 10% used for initial encryption key generation. The dynamic encryption response is 20-50ms. At 20ms, it switches to ChaCha20 encryption based on the MLP prediction result (bandwidth 2Mbps, latency 250ms). The ASIC chip completes the encrypted data packet encapsulation through hardware acceleration at 50ms. The total latency is controlled within 55ms, which is a 40% improvement compared to traditional solutions.

[0118] Another implementation method of this strategy is to protect the copyright of paid on-demand content. When on-demand 4K HDR encrypted movies (protected by DRM-X 4.0), it is necessary to adapt to non-HDR display devices and prevent illegal recording. The system will perform multimodal AI enhancement processing. In terms of HDR→SDR dynamic mapping, the BT.2020 color gamut is converted to BT.709 through a ResNet-based tone mapping network, and the peak brightness is reduced from 1000nit to 300nit; in terms of audio optimization, multi-channel audio (such as Dolby Atmos) is identified and downmixed to stereo output, while using AI sound field expansion technology to maintain the sense of space; in terms of subtitle adaptation, the subtitle position is dynamically adjusted to avoid occlusion of HDR highlight areas, and the font size is adaptively scaled with resolution. In terms of hardware-level copyright protection, the video stream is encrypted using AES-256, and the key is distributed through the ECC algorithm and bound to the device hardware ID (such as the TEE embedded certificate). The ASIC chip generates an invisible watermark containing the user account and timestamp (such as pixel-level micro-offset) and embeds it into each frame. In terms of anti-piracy mechanisms, the hardware ID and watermark consistency are verified through TEE during playback. If illegal recording (such as HDCP cracking) is detected, playback is immediately terminated and reported to the blockchain. Performance verification data shows that the full encryption process (including watermarking) takes ≤ 20ms, while traditional solutions take ≥ 50ms. After HDR to SDR conversion, the SSIM (structural similarity) is ≥ 0.92, with no perceptible difference.

[0119] It should be noted that in the above implementations:

[0120] In terms of dynamic adaptability, AI enhancement and encryption strategies are dynamically adjusted based on real-time environmental parameters (bandwidth, device performance), avoiding resource waste caused by fixed strategies.

[0121] In the hardware collaboration mode, the NPU+ASIC chips have a clear division of labor: the NPU focuses on AI computing, the ASIC handles security tasks such as encryption / watermarking, and hardware queue scheduling is used to avoid resource contention.

[0122] In terms of compliance and compatibility, it supports API interfaces for mainstream DRM standards (Widevine, FairPlay) and complies with GDPR / CCPA data privacy requirements.

[0123] This strategy significantly reduces the stuttering rate in low-bandwidth image quality enhancement; greatly shortens the encryption switching response time through MLP prediction technology; and greatly improves hardware resource utilization by using the NPU+CPU method.

[0124] like Figure 4As shown, starting at 0ms, the video thread initiates ESRGAN super-resolution processing, which continues until 50ms, completing the high-definition reconstruction task. At 5ms, the text thread begins subtitle extraction, completing keyword extraction and outputting structured metadata in 15ms. At 10ms, the audio thread initiates RNN denoising, completing audio purification in 40ms. At 20ms, the key node synchronization controller triggers encryption switching decisions to respond to network fluctuations in real time. At 25ms, the video thread triggers HDR tone mapping to dynamically optimize color. At 30ms, the synchronization controller confirms the readiness of multimodal data, completing the three-way synchronization of audio, video, and text. Finally, at 50ms, three milestones are achieved: video super-resolution reconstruction is complete, audio denoising is complete, and multimodal data synchronization is complete. At this point, the compression encoding module immediately starts and compresses the synchronized high-quality data for output. The entire process achieves multi-threaded parallel processing and coordinated optimization through precise timing control, completing all enhancement operations and generating the output stream within 50ms.

[0125] This embodiment provides an audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption, and the field of digital audio and video processing and secure transmission technology, including: obtaining original audio and video streams, separating the original audio and video streams into audio data, video data and subtitle text data; performing multimodal AI enhancement processing on the audio data, video data and subtitle text data to obtain enhanced audio and video data; generating encryption decision instructions based on real-time network status, encrypting the enhanced audio and video data according to the encryption decision instructions, and generating encrypted data with an anti-tampering watermark; when the encrypted data passes the integrity verification and copyright legitimacy judgment, the encrypted data is transmitted through a secure transmission channel, and the core technology chain of multimodal separation enhancement, dynamic encryption decision, hardware-level watermark binding and trusted verification transmission is used to achieve coordinated optimization of security and efficiency.

[0126] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 5 Before step S40, steps S401 to S404 are also included:

[0127] Step S401, extracting user identity information contained in encrypted data with tamper-proof watermark in a security verification environment;

[0128] In practice, user identity information contained in encrypted data with a tamper-evident watermark is extracted within a secure verification environment. This means extracting user identity information from encrypted and tamper-evident watermarked data within a secure verification environment (one with strict security measures to prevent data tampering or leakage). User identity information can include user account numbers, device IDs, and other identifying information.

[0129] Step S402: Match the user identity information with the device registration information to obtain a user identity information matching result;

[0130] In the specific implementation, the user identity information is matched and compared with the device registration information to obtain the user identity information matching result. This is to compare the extracted user identity information with the information left when the device was registered to confirm whether they are consistent.

[0131] Step S403: Perform an integrity check on the received data and compare it with the original record to obtain a data integrity check result;

[0132] In the specific implementation, the integrity check operation is performed on the received data and compared with the original record to obtain the data integrity check result. The integrity check of the received data is to check whether the data has been tampered with or damaged during the transmission process.

[0133] Step S404: When the user identity information matching result is a successful match and the data integrity check result is a passed data integrity check, it is determined that the verification is passed.

[0134] In a specific implementation, when the user identity information matching result is a successful match and the data integrity check result is a data integrity check passed, it is determined that the verification is passed.

[0135] In a feasible implementation, step S404 further includes step A41:

[0136] A41: When either the user identity information matching result or the data integrity check result fails, the verification is determined to have failed, infringement handling is performed, the audio and video playback process is immediately terminated, an illegal use record is sent to the blockchain node, and a copyright violation warning message is displayed on the playback interface.

[0137] In specific implementations, if either the user identity information match result or the data integrity check result fails, verification is determined to have failed, and infringement handling is executed, immediately terminating the audio and video playback process, sending the illegal use record to the blockchain node, and displaying a copyright violation warning message on the playback interface. If the user identity information match fails or the data integrity check fails, verification is considered failed. At this time, a series of infringement handling operations are executed, such as immediately stopping the audio and video playback process, sending the illegal use record to the blockchain node, and displaying a copyright violation warning message to the user on the playback interface.

[0138] Furthermore, when the encrypted data passes integrity verification and copyright validity determination, the encrypted data is segmented into data units suitable for network transmission;

[0139] Transmit data units through encrypted communication protocols and maintain forward security protection mechanisms throughout the entire transmission process;

[0140] After the receiving end successfully receives all data units and confirms their integrity, the transmission process is completed.

[0141] In specific implementations, once the encrypted data passes integrity verification and copyright validation, it is segmented into data units suitable for network transmission. Once the encrypted data passes integrity verification and its copyright is determined to be valid, the encrypted data is broken into smaller pieces, which are more suitable for network transmission. Because network transmission sometimes has data size limitations, segmenting the data into smaller pieces makes transmission easier and more efficient.

[0142] Transmit the data units using an encrypted communication protocol, maintaining forward security throughout the entire transmission process. This means using an encrypted communication protocol to transmit the previously segmented data units. Furthermore, forward security must be maintained throughout the entire transmission process.

[0143] The transmission process is completed when the receiving end successfully receives all data units and verifies their integrity. This means that when the receiving end successfully receives all the divided data units and verifies that the data is complete, the entire transmission process is completed.

[0144] This embodiment provides an audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption. Through user identity authentication and data integrity verification, it ensures that only legitimate users can use audio and video content and that the content has not been tampered with during transmission. By performing infringement processing operations when verification fails, it can effectively deter and prevent the illegal use of audio and video content and protect the rights and interests of copyright holders. In addition, the use of encrypted communication protocols and forward security protection mechanisms further enhances the security of audio and video data during transmission, so that audio and video content can be effectively protected in various network environments.

[0145] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.

[0146] This application also provides an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, please refer to Figure 6 , the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption includes:

[0147] The data acquisition module 10 is used to obtain the original audio and video stream and separate the original audio and video stream into audio data, video data and subtitle text data;

[0148] Enhancement module 20, configured to perform multimodal AI enhancement processing on audio data, video data, and subtitle text data to obtain enhanced audio and video data;

[0149] The encryption module 30 is used to generate an encryption decision instruction based on the real-time network status, encrypt the enhanced audio and video data according to the encryption decision instruction, and generate encrypted data with an anti-tampering watermark;

[0150] The transmission module 40 is used to transmit the encrypted data through a secure transmission channel when the encrypted data passes the integrity verification and copyright validity determination.

[0151] The audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application adopts the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption in the above-mentioned embodiment, which can solve the technical problems of low efficiency and insufficient security of low-quality audio and video transmission. Compared with the prior art, the beneficial effects of the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application are the same as the beneficial effects of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided by the above-mentioned embodiment, and the other technical features of the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0152] In one embodiment, the enhancement module 20 is further configured to convert low-resolution video frames of the video data into high-resolution video frames through an artificial intelligence image processing model;

[0153] Eliminate background noise from audio data through an audio processing algorithm to obtain noise-reduced audio data;

[0154] Use natural language processing technology to identify key content in subtitle text data and generate metadata tags synchronized with the video timeline;

[0155] The high-resolution video frames, denoised audio data, and metadata tags are recombined to obtain enhanced audio and video data.

[0156] In one embodiment, the enhancement module 20 is further configured to monitor network transmission parameters in real time;

[0157] Input the network transmission parameters into the decision model for processing and analysis to obtain processing and analysis results;

[0158] Generate instructions using high-security encryption algorithms based on processing and analysis results under high-quality network conditions;

[0159] Generate instructions using a high-efficiency encryption algorithm based on the processing and analysis results under low-quality network conditions.

[0160] In one embodiment, the enhancement module 20 is further configured to perform an encryption operation on the enhanced audio and video data to generate encrypted audio and video data;

[0161] Add device-bound watermark information to the encrypted audio and video data to form an anti-tampering protection layer and generate encrypted data with an anti-tampering watermark.

[0162] In one embodiment, the enhancement module 20 is further configured to extract user identity information contained in the encrypted data with the tamper-proof watermark in a secure verification environment;

[0163] Match and compare the user identity information with the device registration information to obtain a user identity information matching result;

[0164] Perform integrity check on the received data and compare it with the original record to obtain the data integrity check result;

[0165] When the user identity information matching result is a successful match and the data integrity check result is a data integrity check passed, it is determined that the verification is passed.

[0166] In one embodiment, the enhancement module 20 is further configured to determine that the verification has failed when either the user identity information matching result or the data integrity check result fails, execute infringement processing operations, immediately terminate the audio and video playback process, send an illegal use record to the blockchain node, and display a copyright violation warning message on the playback interface.

[0167] In one embodiment, the enhancement module 20 is further configured to segment the encrypted data into data units suitable for network transmission when the encrypted data passes integrity verification and copyright validity determination;

[0168] Transmit data units through encrypted communication protocols and maintain forward security protection mechanisms throughout the entire transmission process;

[0169] After the receiving end successfully receives all data units and confirms their integrity, the transmission process is completed.

[0170] The present application provides an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption. The audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption in the above-mentioned embodiment one.

[0171] Reference below Figure 7 , which shows a schematic structural diagram of an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption suitable for implementing the embodiments of the present application. The audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0172] like Figure 7As shown, the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 to RAM (Random Access Memory) 1004. Various programs and data required for the operation of the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption are also stored in RAM 1004. The processing device 1001, ROM 1002 and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0173] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0174] The audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application adopts the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption in the above-mentioned embodiment, which can solve the technical problems of low efficiency and insufficient security of low-quality audio and video transmission. Compared with the prior art, the beneficial effects of the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption provided by this application are the same as the beneficial effects of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided by the above-mentioned embodiment, and the other technical features of the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption are the same as the features disclosed in the method of the previous embodiment, and will not be repeated here.

[0175] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0176] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0177] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption in the above-mentioned embodiment.

[0178] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash memory), optical fiber, CD-ROM (CD-Read Only Memory), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0179] The above-mentioned computer-readable storage medium can be included in the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption; or it can exist alone without being assembled into the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption.

[0180] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, the audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption: the field of digital audio and video processing and secure transmission technology, including: obtaining the original audio and video stream, separating the original audio and video stream into audio data, video data and subtitle text data; performing multimodal AI enhancement processing on the audio data, video data and subtitle text data to obtain enhanced audio and video data; generating encryption decision instructions based on real-time network status, encrypting the enhanced audio and video data according to the encryption decision instructions, and generating encrypted data with an anti-tampering watermark; when the encrypted data passes the integrity verification and copyright legitimacy judgment, the encrypted data is transmitted through a secure transmission channel.

[0181] The computer program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0182] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0183] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0184] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption, which can solve the technical problems of low efficiency and insufficient security of low-quality audio and video transmission. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided in the above-mentioned embodiment, and will not be repeated here.

[0185] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption.

[0186] The computer program product provided in this application can address the technical issues of low efficiency and insufficient security in low-quality audio and video transmission. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption provided in the above-mentioned embodiment, and will not be elaborated here.

[0187] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption, characterized in that: The method comprises: Acquire an original audio and video stream, and separate the original audio and video stream into audio data, video data, and subtitle text data; Performing multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data; Generate an encryption decision instruction based on the real-time network status, encrypt the enhanced audio and video data according to the encryption decision instruction, and generate encrypted data with an anti-tampering watermark; When the encrypted data passes the integrity verification and copyright legality determination, the encrypted data is transmitted through a secure transmission channel.

2. The method according to claim 1, wherein The step of performing multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data includes: Converting low-resolution video frames of the video data into high-resolution video frames through an artificial intelligence image processing model; Eliminating background noise from the audio data using an audio processing algorithm to obtain noise-reduced audio data; Identifying key content in the subtitle text data through natural language processing technology and generating metadata tags synchronized with the video timeline; The high-resolution video frame, the noise-reduced audio data, and the metadata tag are recombined to obtain enhanced audio and video data.

3. The method according to claim 1, wherein The step of generating an encryption decision instruction based on the real-time network status includes: Real-time monitoring of network transmission parameters; Inputting the network transmission parameters into a decision model for processing and analysis to obtain processing and analysis results; generating, under high-quality network conditions, instructions using a high-security encryption algorithm based on the processing and analysis results; An instruction for using a high-efficiency encryption algorithm is generated based on the processing and analysis results under low-quality network conditions.

4. The method according to claim 1, wherein The step of encrypting the enhanced audio and video data according to the encryption decision instruction to generate encrypted data with an anti-tampering watermark includes: Performing an encryption operation on the enhanced audio and video data to generate encrypted audio and video data; Device-bound watermark information is added to the encrypted audio and video data to form an anti-tampering protection layer, thereby generating encrypted data with an anti-tampering watermark.

5. The method according to claim 1, wherein Before the step of transmitting the encrypted data through a secure transmission channel when the encrypted data passes integrity verification and copyright legitimacy determination, the method further includes: Extracting user identity information contained in the encrypted data with the tamper-proof watermark in a secure verification environment; Matching the user identity information with the device registration information to obtain a user identity information matching result; Perform integrity check on the received data and compare it with the original record to obtain the data integrity check result; When the user identity information matching result is a successful match and the data integrity check result is a data integrity check passed, it is determined that the verification is passed.

6. The method according to claim 5, wherein The method further comprises: When any one of the user identity information matching result or the data integrity check result fails, the verification is determined to have failed, and the infringement handling operation is executed, the audio and video playback process is immediately terminated, the illegal use record is sent to the blockchain node, and a copyright violation warning message is displayed on the playback interface.

7. The method according to claim 1, wherein The step of transmitting the encrypted data through a secure transmission channel when the encrypted data passes the integrity verification and the copyright legitimacy determination includes: When the encrypted data passes the integrity verification and copyright validity determination, dividing the encrypted data into data units suitable for network transmission; transmitting the data unit via an encrypted communication protocol and maintaining a forward security protection mechanism throughout the transmission process; After the receiving end successfully receives all data units and confirms their integrity, the transmission process is completed.

8. An audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, characterized in that: The device comprises: A data acquisition module is used to acquire the original audio and video stream and separate the original audio and video stream into audio data, video data and subtitle text data; An enhancement module, configured to perform multimodal AI enhancement processing on the audio data, the video data, and the subtitle text data to obtain enhanced audio and video data; An encryption module, configured to generate an encryption decision instruction based on the real-time network status, and encrypt the enhanced audio and video data according to the encryption decision instruction to generate encrypted data with an anti-tampering watermark; The transmission module is used to transmit the encrypted data through a secure transmission channel when the encrypted data passes the integrity verification and the copyright legality determination.

9. An audio and video transmission device based on multimodal AI enhancement and dynamic end-to-end encryption, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the audio and video transmission method based on multimodal AI enhancement and dynamic end-to-end encryption are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Internal and external network audio and video secure transmission method and system based on cloud platform

    CN121441655A

  • A cloud platform-based internal and external network audio and video security transmission method and system

    CN121441655B