Content enhancement method applied to audio and video and related equipment

By employing high-efficiency video encoding, audio pixel decoding, multimodal structured tagging, and Z-shaped reconstruction repair technology, the problem of the inability of existing audio and video enhancement technologies to dynamically adjust has been solved, thereby improving audio and video quality and providing efficient system support.

CN121486593APending Publication Date: 2026-02-06ZHONGLI INTELLIGENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511760882.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing audio and video enhancement technologies cannot dynamically adjust according to different audio and video characteristics and actual needs, resulting in limited enhancement effects and affecting user experience and the accuracy of information delivery.

Method used

It employs high-efficiency video encoding, audio pixel decoding, multimodal structured tagging, dynamic rule base to generate decision instructions, data augmentation processing, and Z-shaped reconstruction and repair technology to achieve dynamic enhancement and mapping of audio and video streams.

Benefits of technology

It effectively enhances audio and video quality, provides high-quality audio and video support, meets the real-time and security requirements of professional scenarios, and reduces system latency and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486593A_ABST
    Figure CN121486593A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of artificial intelligence, and relates to a content enhancement method applied to audio and video and related equipment, and the method comprises the steps: receiving a high-efficiency video code transmitted by a transmitting end of a KVM seat cooperation system; performing audio pixel decoding processing on the high-efficiency video code to obtain original audio pixel data; performing multi-modal structured marking processing on the original audio pixel data to obtain multi-modal structured identification data; generating a decision instruction according to a preset dynamic rule base and the multi-modal structured identification data; performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data; performing Z-shaped recombination repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream; and mapping the enhanced audio and video stream into a standard audio and video source, and returning the standard audio and video source to the KVM system. According to the invention, the audio and video quality is effectively enhanced, and high-quality audio and video support is provided for the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and related equipment for enhancing audio and video content. Background Technology

[0002] In KVM-based collaborative systems, audio and video transmission and processing are core functions. However, in practical applications, due to factors such as network bandwidth limitations and encoding compression losses, the audio and video data transmitted from the sending end to the receiving end often suffers from quality degradation, such as audio distortion and video blurring. This not only affects the user's audiovisual experience but may also lead to inaccurate information transmission in scenarios with high audio and video quality requirements, such as remote conferencing and monitoring command, thus impacting work efficiency and decision-making accuracy.

[0003] Most existing audio and video enhancement technologies process a single modality, lacking the ability to comprehensively process audio and video. Furthermore, their processing methods are relatively fixed and cannot be dynamically adjusted according to different audio and video characteristics and actual needs, resulting in limited enhancement effects. Summary of the Invention

[0004] The purpose of this application is to propose a content enhancement method and related equipment for audio and video, so as to solve the problem that existing audio and video enhancement technologies cannot be dynamically adjusted according to different audio and video characteristics and actual needs.

[0005] To address the aforementioned technical problems, this application provides a content enhancement method for audio and video, employing the following technical solution:

[0006] Receive high-efficiency video encoding sent by the transmitter of the KVM agent collaboration system;

[0007] The high-efficiency video encoding is subjected to audio pixel decoding to obtain the original audio pixel data;

[0008] The original audio pixel data is subjected to multimodal structured tagging processing to obtain multimodal structured tagging data;

[0009] Decision instructions are generated based on a preset dynamic rule base and the multimodal structured identifier data;

[0010] The original audio pixel data is augmented according to the decision instruction to obtain augmented audio pixel data.

[0011] The enhanced audio pixel data is subjected to Z-shaped reconstruction and repair processing to obtain an enhanced audio and video stream;

[0012] The enhanced audio and video stream is mapped to a standard audio and video source, and the standard audio and video source is sent back to the KVM system.

[0013] To address the aforementioned technical problems, this application also provides a content enhancement device for audio and video, employing the following technical solution:

[0014] A high-efficiency video encoding acquisition module is used to receive high-efficiency video encoding sent by the sending end of the KVM agent collaboration system;

[0015] An audio pixel decoding module is used to perform audio pixel decoding processing on the high-efficiency video encoding to obtain the original audio pixel data;

[0016] A multimodal structured tagging module is used to perform multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured tagging data;

[0017] The decision instruction generation module is used to generate decision instructions based on a preset dynamic rule base and the multimodal structured identifier data.

[0018] The data augmentation module is used to perform data augmentation processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data;

[0019] The Z-shaped reconstruction and repair module is used to perform Z-shaped reconstruction and repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream;

[0020] The audio and video stream mapping module is used to map the enhanced audio and video stream to a standard audio and video source, and to send the standard audio and video source back to the KVM system.

[0021] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution:

[0022] It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the content enhancement method applied to audio and video as described above.

[0023] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below:

[0024] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the content enhancement method applied to audio and video as described above.

[0025] This application provides a content enhancement method for audio and video, comprising: receiving high-efficiency video encoding sent by the transmitter of the KVM agent collaboration system; performing audio pixel decoding processing on the high-efficiency video encoding to obtain original audio pixel data; performing multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured identifier data; generating decision instructions based on a preset dynamic rule base and the multimodal structured identifier data; performing data enhancement processing on the original audio pixel data according to the decision instructions to obtain enhanced audio pixel data; performing Z-shaped reconstruction and repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream; mapping the enhanced audio and video stream to a standard audio and video source, and transmitting the standard audio and video source back to the KVM system. Compared with the prior art, this application effectively enhances the audio and video quality and maps the enhanced audio and video stream to a standard audio and video source for transmission back to the KVM system, providing high-quality audio and video support for the system. Attached Figure Description

[0026] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0028] Figure 2 This is a flowchart illustrating the implementation of the content enhancement method for audio and video provided in this application embodiment;

[0029] Figure 3 This is a schematic diagram of the structure of a conference privacy protection and content enhancement system based on real-time multimodal data processing provided in this application embodiment;

[0030] Figure 4 This is a schematic diagram of the structure of the distributed high-definition video transmission device provided in the embodiments of this application;

[0031] Figure 5 This is a schematic diagram of the structure of the distributed high-definition video receiving device provided in the embodiments of this application;

[0032] Figure 6 This is a schematic diagram of the structure of the content enhancement device for audio and video provided in the embodiments of this application;

[0033] Figure 7 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0035] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0036] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0037] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0038] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0039] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0040] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0041] It should be noted that the content enhancement method for audio and video provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the content enhancement device for audio and video is generally set in the server / terminal device.

[0042] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0043] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of the audio / video content enhancement method according to this application. The above-described audio / video content enhancement method, applied to a KVM workstation collaboration system, includes steps S201, S202, S203, S204, S205, S206, and S207.

[0044] In this application embodiment, a KVM (Keyboard Video Mouse) seat collaboration system refers to a system that enables access to and control of a computer via a direct connection to a keyboard, video, or mouse (KVM) port. Multiple operators can access and control multiple remote server hosts through a single keyboard, mouse, and monitor. Video conferencing on such systems is typically used in high-risk, highly specialized scenarios such as emergency command, production scheduling, and broadcast control.

[0045] In this embodiment of the application, the operation flow of the KVM agent collaboration system is as follows:

[0046] Sending link: KVM distributed input nodes compress and encode the raw video stream using an H.265 encoder, forming the basic link for IP network transmission;

[0047] Receiver decoding: The receiver's "subframe decoding unit" decodes the encoded data to restore the original pixel information of the subframe, ensuring that subsequent reconstruction and repair processes are based on complete and valid image data, and guaranteeing the accuracy of image quality restoration.

[0048] End-to-end synchronization: The PTPv2 clock protocol is used to align the nanosecond-level timestamps between the sending and receiving ends, ensuring precise matching of audio and video streams in the time dimension.

[0049] In step S201, the high-efficiency video encoding sent by the transmitter of the KVM agent collaboration system is received.

[0050] In this application embodiment, High Efficiency Video Coding (HEVC), also known as H.265, is a video compression standard jointly developed by the International Telecommunication Union (ITU-T) Video Coding Experts Group and the ISO / IEC Moving Picture Experts Group, and was officially released in 2013. As an upgrade to H.264 / AVC, its core objective is to increase the data compression rate by approximately 50% while maintaining the same image quality, supporting ultra-high-definition video with a maximum resolution of 7680×4320. This standard has been incorporated into the ITU-T H.265 and ISO / IEC 23008-2 international standards. HEVC adopts a block-based hybrid coding framework, including intra-frame prediction, transform coding, and entropy coding techniques. Its intra-frame prediction modes have been expanded from 9 in H.264 to 35 (including 33 angle modes, DC mode, and Planar mode), and a quadtree coding unit structure (supporting 64×64 to 8×8 segmentation) and RQT adaptive transform have been introduced. The encoding process includes reference sample adaptive smoothing, prediction mode selection, residual transformation, and coefficient scanning, employing a context-based encoding strategy and sample adaptive offset optimization technology. This standard supports 720P high-definition video transmission at bitrates of 1-2 Mbps with only approximately a 2% increase in hardware consumption, and is widely used in surveillance, video conferencing, and ultra-high-definition streaming media transmission.

[0051] In step S202, the high-efficiency video encoding is subjected to audio pixel decoding processing to obtain the original audio pixel data.

[0052] In this embodiment, a distributed command center leadership meeting is used as the specific implementation scenario. It deeply integrates multimodal real-time analysis and intelligent rendering modules to form a complete closed loop of "analysis-decision-execution". The receiving end completes real-time decoding of the H.265 bitstream through the subframe decoding unit, synchronously parses the embedded timestamp metadata, restores the original pixel data of various color encoding formats, and achieves low-latency transmission through a high-speed bus interface to ensure seamless integration with the next module. Specific implementation methods include:

[0053] (1) High-quality multimodal source acquisition:

[0054] Capture lossless or low-loss audio and video streams to provide broadcast-quality raw data that surpasses that of conventional video conferencing systems for all subsequent AI processing, ensuring the accuracy and realism of the final result. Specifically:

[0055] • Video Stream Capture: Captures the screen composition of a designated workstation from the video output card of the KVM system via physical interfaces (HDMI, DP) or IP streams (RTSP protocol, etc.). Supports multiple high-definition and ultra-high-definition resolutions, covering a wide range from Full HD to Ultra HD, and features high frame rate processing capabilities; supports multiple color sampling formats and is compatible with other industry-standard color encoding methods; end-to-end capture latency is kept low to meet the needs of real-time processing scenarios.

[0056] • Audio Stream Capture: Captures the audio channels of a specified seat from the digital audio mixing matrix of the KVM system. It has multimodal audio format adaptation capabilities, supports industry standard audio encoding schemes that meet high fidelity requirements, and achieves accurate capture and lossless transmission of multi-source mixed audio signals, meeting the stringent requirements of professional scenarios for sound quality reproduction.

[0057] (2) Multimodal data synchronization and association:

[0058] Building a precise spatiotemporal alignment foundation enables cross-modal analysis, specifically:

[0059] • By assigning a unified timestamp to video frames and audio segments at the same point in time and aligning the buffers, we can ensure that visual behaviors and audio events correspond precisely in time, laying the foundation for collaborative analysis and processing of multimodal data.

[0060] In step S203, the original audio pixel data is subjected to multimodal structured labeling processing to obtain multimodal structured label data.

[0061] In this embodiment of the application, after the receiving end completes subframe decoding and metadata parsing, the newly added multimodal analysis unit immediately starts high-precision real-time marking.

[0062] In step S204, decision instructions are generated based on a preset dynamic rule base and multimodal structured identifier data.

[0063] In this embodiment of the application, in the scenario of a distributed command center leadership meeting, after the multimodal analysis unit completes the structured labeling, the dynamic rule engine triggers a precise strategy based on the labeling results and the real-time context.

[0064] In step S205, the original audio pixel data is subjected to data augmentation processing according to the decision instruction to obtain enhanced audio pixel data.

[0065] In this embodiment of the application, in the scenario of a distributed command center leadership meeting, the dynamic rule engine and the real-time AI rendering and generative content enhancement module form a strict closed loop of "decision-execution", and the real-time AI rendering and generative content enhancement module (step 180) performs pixel-level operations.

[0066] In step S206, the enhanced audio pixel data is subjected to Z-shaped reconstruction and repair processing to obtain the enhanced audio and video stream.

[0067] In this embodiment of the application, in the scenario of a leadership meeting in a distributed command center, after real-time AI rendering is executed, it needs to be repaired by Z-shaped reconstruction to ensure the integrity of the image. Specifically, the purified subframes output by the AI ​​rendering module are subjected to Z-shaped reconstruction (with delay compensation, time ≤8ms) and weighted interpolation repair to solve the problem of disordered / missing subframes caused by single-path transmission, and to ensure that the reconstructed I-frames have no tearing or blurring defects, and the color restoration accuracy ΔE<2.

[0068] In step S207, the enhanced audio and video stream is mapped to a standard audio and video source, and the standard audio and video source is sent back to the KVM system.

[0069] In this embodiment of the application, in the scenario of a distributed command center leadership meeting, the seamless system integration and output module serves as the final link, transmitting the processed enhanced signal back to the KVM system in the form of a standard device, thereby achieving seamless integration with the existing workflow.

[0070] This application provides a method for enhancing audio and video content, comprising: receiving high-efficiency video encoding sent by a transmitter of a KVM agent collaboration system; performing audio pixel decoding on the high-efficiency video encoding to obtain original audio pixel data; performing multimodal structured tagging on the original audio pixel data to obtain multimodal structured identifier data; generating decision instructions based on a preset dynamic rule base and the multimodal structured identifier data; performing data enhancement processing on the original audio pixel data according to the decision instructions to obtain enhanced audio pixel data; performing Z-shaped reconstruction and repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream; mapping the enhanced audio and video stream to a standard audio and video source, and transmitting the standard audio and video source back to the KVM system. Compared with the prior art, this application effectively enhances audio and video quality and maps the enhanced audio and video stream to a standard audio and video source for transmission back to the KVM system, providing high-quality audio and video support for the system.

[0071] In some optional implementations of the embodiments of this application, the multimodal structured identification data includes identity identification information. The step of performing multimodal structured tagging processing on the original audio pixel data to obtain the multimodal structured identification data specifically includes the following steps:

[0072] The raw audio pixel data is processed by avatar pixel recognition to obtain the raw avatar pixel data.

[0073] Read the system database, perform similarity matching on the original avatar pixel data based on the preset avatar data in the system database, and obtain the similarity matching result;

[0074] The similarity matching results are filtered according to the preset avatar matching conditions to obtain the avatar filtering results;

[0075] Obtain the identity information corresponding to the avatar screening results, and mark the original audio pixel data according to the identity information to obtain multimodal structured identification data;

[0076] The decision-making instructions include enabling or disabling virtual avatars, and the steps of generating decision-making instructions based on a preset dynamic rule base and multimodal structured identifier data include the following steps:

[0077] Determine whether identity information falls under the scope of privacy protection;

[0078] If the identity information is a privacy-protected object, a decision instruction to enable the virtual avatar will be generated;

[0079] If the identity information is not subject to privacy protection, a decision instruction to disable the virtual avatar will be generated.

[0080] The steps involved in performing data augmentation on the original audio pixel data according to the decision instructions to obtain augmented audio pixel data include the following steps:

[0081] When the decision instruction is to enable virtual avatars, dynamic 3D virtual avatar pixel data corresponding to the original avatar pixel data is generated according to the MediaPipe architecture, and the original avatar pixel data of the original audio pixel data is replaced with dynamic 3D virtual avatar pixel data to obtain enhanced audio pixel data.

[0082] When the decision instruction is to disable virtual avatars, no data augmentation processing will be performed on the original audio pixel data.

[0083] In this application embodiment, the high-precision real-time tagging of this application specifically includes an identity tag. Specifically, this application generates a corresponding identity tag based on a preset 1:1 headshot comparison strategy when a facial similarity > 95% is detected. For example, when the facial similarity of Chief Engineer Li (user.id=003) is detected to be > 95%, the tag “Face Protection_user_003_Timestamp[t9-t10]” is generated.

[0084] In this embodiment, when the marker “Chief Engineer Li (user.id=003)_Face Protection_Timestamp [t9-t10]” is detected and user.role = “Advanced Administrator”, the rule engine automatically executes “IF user.role==Advanced Administrator THEN enable(Face_Filtering) AND set(Replacement_Mode=Avatar)”, calling MediaPipe to generate a dynamic 3D virtual image, ensuring that facial details are not leaked and the lip-sync error is <30ms; in some optional implementations of this embodiment, if the user is a guest, the “Blur” mode is switched to blur the image, thereby achieving differentiated control of identity sensitivity.

[0085] In this application embodiment, the enhancement processing of identity marker data is face virtualization and enhancement. Specifically, in response to the "Enable Avatar Mode" command, 478 3D facial key points are detected through MediaPipe Face Mesh to drive a high-precision 3D virtual image in real time (frame rate ≥30fps, natural expression).

[0086] In some optional implementations of the embodiments of this application, the above-mentioned multimodal structured identification data includes behavioral identification information. The step of performing multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured identification data specifically includes the following steps:

[0087] Based on the HRNet-W48+Bi-LSTM architecture, the raw pixel data of the original audio pixel data is processed for behavior recognition to obtain behavior identification information;

[0088] The original audio pixel data is labeled based on the behavioral identification information to obtain multimodal structured identification data;

[0089] The decision-making instructions include the steps of triggering virtual subtitle prompts and not triggering virtual subtitle prompts, and generating decision-making instructions based on a preset dynamic rule base and multimodal structured identifier data, specifically including the following steps:

[0090] Obtain system mode information corresponding to the original audio pixel data, and determine whether the behavior identifier information belongs to a non-scene action based on the system mode information;

[0091] If the behavior identifier information belongs to a non-scene action, then a decision instruction to trigger the virtual subtitle prompt is generated;

[0092] If the behavior identifier information does not belong to a non-scene action, then a decision instruction that does not trigger virtual subtitle prompts is generated;

[0093] The steps involved in performing data augmentation on the original audio pixel data according to the decision instructions to obtain augmented audio pixel data include the following steps:

[0094] When the decision instruction is to trigger a virtual caption prompt, a behavioral reminder caption is generated from the original audio pixel data to obtain enhanced audio pixel data;

[0095] If the decision instruction is to not trigger virtual subtitle prompts, then no data augmentation processing will be performed on the original audio pixel data.

[0096] In this application embodiment, the high-precision real-time tagging of this application specifically includes behavior identification tags. Specifically, this application uses the HRNet-W48+Bi-LSTM architecture to identify human behavior and generate structured tags, such as "user_001_frequent distraction_number of times ≥3_timestamp[t1-t2]" and "user_002_lying on the table_duration ≥5s_timestamp[t3-t4]", etc.

[0097] In this embodiment, after identifying a person's behavior, the application will evaluate the behavior marker based on the context information and correct the user behavior through a behavior identification strategy. For example, when the context information is system.mode="Emergency Command" and time="14:00", the rule engine evaluates the "user_001_frequent distraction_number of times ≥3_timestamp[t1-t2]" marker and triggers the "focus reminder" strategy, which overlays the "Please stay focused" prompt with virtual subtitles. If system.mode is switched to "Night shift", the Behavior_Sensitivity_Level is automatically adjusted to "Low", relaxing the judgment threshold for behaviors such as "yawning" to avoid excessive intervention during informal periods.

[0098] In this embodiment of the application, the behavior identification mark data enhancement processing is behavior repair and enhancement. Specifically, in response to the "focus reminder" instruction, a "Please stay focused" prompt is superimposed on the edge of the screen through a virtual subtitle engine.

[0099] In some optional implementations of the embodiments of this application, the step of performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data specifically includes the following steps:

[0100] Retrieve historical cached data corresponding to the original audio pixel data;

[0101] Based on the StarGAN-v2 architecture and frame history cache data, semantically guided translation is performed on the human body region in the current frame of the original audio pixel data to generate standard pose pixel data.

[0102] Replace the human body region in the current frame of the original audio pixel data with standard pose pixel data.

[0103] In this embodiment, the augmentation processing of behavior identifier data can also be performed by calling StarGAN-v2 with a 30-frame history cache as a style reference to perform semantically guided translation of the human body region in the current frame, generate a standard pose and replace it (single frame time < 10ms).

[0104] In some optional implementations of the embodiments of this application, the above-mentioned multimodal structured identifier data includes sensitive word identifier information. The step of performing multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured identifier data specifically includes the following steps:

[0105] The original audio data is processed by the speech recognition model to obtain the original audio text data.

[0106] The original audio text data is processed by sensitive word matching based on the sensitive word database to obtain sensitive word identification information;

[0107] The original audio pixel data is labeled based on the sensitive word identification information to obtain multimodal structured identification data;

[0108] The decision-making instructions include audio mute instructions. The steps for generating decision-making instructions based on a preset dynamic rule base and multimodal structured identifier data specifically include the following steps:

[0109] Retrieve audio mute commands corresponding to sensitive word identifiers from the dynamic rule base;

[0110] The steps involved in performing data augmentation on the original audio pixel data according to the decision instructions to obtain augmented audio pixel data include the following steps:

[0111] The original audio pixel data is muted according to the audio mute command to obtain enhanced audio pixel data.

[0112] In this embodiment, the high-precision real-time tagging specifically includes sensitive word identification tagging. Specifically, through the mainstream speech recognition service model, indecent content and sensitive words are identified, and structured tags such as "indecent words_nonsense_timestamp[t11-t12]" and "sensitive words_security risks_timestamp[t13-t14]" are generated. The content review service model is integrated to achieve sensitive word matching without directly triggering silencing or highlighting, ensuring collaborative processing with the dynamic rule engine.

[0113] In this embodiment, when seat.tag = "Finance Department", the rule engine dynamically loads the "Financial_Terms" sensitive word library to accurately monitor words such as "stock price" and "financial report". If the "vulgar words_nonsense_timestamp[t11-t12]" tag is detected, the "Audio_Mute" strategy is triggered to truncate the audio segment and insert a "beep" prompt sound to ensure the professionalism of the meeting audio.

[0114] In this application embodiment, the sensitive word identification mark data enhancement processing is voice purification and enhancement. Specifically, in response to the “Audio_Mute” command, the sensitive word time period is precisely muted or replaced at the sample level based on VAD and timestamp (mute positioning error <30ms, audio waveform transition is smooth).

[0115] In some optional implementations of the embodiments of this application, the high-precision real-time tagging of this application also includes category tagging. Specifically, this application uses the YOLOv7-tiny model to identify various types of objects (such as unauthorized personnel, pets, and spilled water cups) and generates region coordinates and category tags (such as "unauthorized personnel_region[x1,y1,x2,y2]_timestamp[t5-t6]" "spilled water cup_region[x3,y3,x4,y4]_timestamp[t7-t8]", etc.).

[0116] In some optional implementations of the embodiments of this application, the data augmentation processing of this application also includes scene repair and enhancement. Specifically, in response to the "scene filling" instruction, the LaMa model is called with the seat static background template as a priori to fill in the irrelevant object area.

[0117] In some optional implementations of the embodiments of this application, the above-mentioned standard audio and video source includes a standard video source and a standard audio source. The step of mapping the enhanced audio and video stream to the standard audio and video source and sending the standard audio and video source back to the KVM system specifically includes the following steps:

[0118] Create a virtual camera device based on the V4L2 driver layer;

[0119] The enhanced video stream is mapped to a standard video source based on the virtual camera device;

[0120] Create a virtual microphone device based on the ALSA driver layer;

[0121] The enhanced audio stream of the enhanced audio and video streams is mapped to a standard audio source based on the virtual microphone device.

[0122] In this embodiment of the application, the method for converting the processed enhanced signal into a standard device form can be as follows:

[0123] • Video Output: A virtual camera device, KVM_Virtual_Cam, is created through the V4L2 driver layer. The enhanced video stream (such as repaired behavior and poses, filled-in scene areas, and virtualized face images) after zigzag reconstruction and repair is mapped to a standard video source. The KVM conferencing module automatically recognizes this device as a "camera" signal source, supporting real-time transmission of high-definition video to meet the needs of various scenarios.

[0124] • Audio Output: A virtual microphone device, KVM_Virtual_Mic, is created through the ALSA driver layer, mapping the enhanced audio stream (such as muted sensitive word segments, smoothly transitioned waveforms, etc.) to a standard audio source. The KVM conferencing module automatically recognizes this device as a "microphone" signal source, supporting high-fidelity audio transmission to meet the needs of various scenarios.

[0125] The KVM system treats KVM_Virtual_Cam and KVM_Virtual_Mic as ordinary communication sources, requiring no modification to existing meeting module code or workflows. For example, when the system switches to "Emergency Command" mode, the "Attention Reminder" command (such as virtual caption overlay) triggered by the dynamic rule engine (step 170) will be directly output through the virtual camera, without the upper-layer application needing to be aware of the underlying processing logic. Through GPU hardware acceleration and CPU multi-threading optimization, the system resource usage (CPU / GPU) is ensured to be less than 20% under full load. For example, StarGAN-v2 behavior repair and LaMa scene filling are processed in parallel through CUDA cores, avoiding excessive consumption of computing resources.

[0126] In practical applications, Figure 3 A schematic diagram of a conference privacy protection and content enhancement system based on real-time multimodal data processing is shown, such as... Figure 3As shown, the transmitting end 1210 compresses the HDMI source video stream into an H.265 bitstream through the distributed high-definition video transmitting device 1212, generates priority subframes through keyframe splitting, and transmits them through the multi-layer network 1220 according to the FEC redundancy strategy; the receiving end 1230's distributed high-definition video receiving device 1231 first performs subframe decoding and SEI metadata parsing to restore the high dynamic range pixel format adapted to subsequent processing requirements, and obtains context information such as timestamps; the core AI analysis module is based on a dynamic rule engine, combining user roles (such as senior administrators enabling Avatar mode, etc.), system modes (such as emergency command triggering focus reminders, etc.) and agent attributes (such as financial... (Ministry of Commerce sensitive word monitoring, etc.) drives the generative AI model to perform pixel-level enhancements—using StarGAN-v2 to achieve single-frame behavior repair, LaMa to complete scene filling, and MediaPipe to achieve lip-sync and face virtualization; after the image integrity is processed by the reconstruction and repair unit, KVM_Virtual_Cam / Mic is created through the virtual device unit integrated by the KVM system, supporting direct connection to the large screen 1232a or the console 1232b via HDMI output, achieving transparent integration with end-to-end latency <30ms and resource consumption <20%, so that end users can enjoy the enhanced professional-grade meeting experience without modifying the code, while meeting the requirements of privacy protection, content compliance, and real-time performance.

[0127] In practical applications, Figure 4 A schematic diagram of a distributed high-definition video transmission device has been presented, such as... Figure 4 As shown, the distributed high-definition video transmitter integrates three core modules: the compression and extraction unit achieves efficient processing of the original source through lossless / low-loss capture and H.265 compression encoding, and synchronously embeds SEI metadata to support end-to-end synchronization; the keyframe splitting unit relies on the FPGA chip to achieve precise grid splitting (error < 0.05 pixels), adds pixel overlap areas and generates single-path transmission markers (splitting time < 1ms), and works with the single-path scheduling module to sort according to the priority of "Class A (keyframe) → Class B (non-keyframe)", dynamically configuring 10%-30% FEC redundancy to adapt to network conditions and ensure transmission reliability; the subframe definition unit encapsulates subframes with CRC-32 checksums and unique ID markers, and is compatible with the existing KVM single-path protocol through a standard IP network interface, realizing traceability and integrity verification of the transmission process.

[0128] In practical applications, Figure 5The distributed high-definition video receiving device adopts a modular architecture to achieve end-to-end processing, integrating four core modules: The decompression and metadata parsing unit restores high-precision pixel data in multiple formats through efficient decoding and low-latency transmission technology; the intelligent rendering and enhancement unit deploys generative AI models based on the Axera650 chip computing platform, performing behavior repair, scene filling, and face virtualization based on dynamic rule engine strategies, and combining metadata to achieve context-aware enhancement, ensuring smooth processing and visual consistency; the reconstruction, repair, and output unit ensures image integrity through efficient reconstruction and repair algorithms, outputting ultra-high-definition video frames that support high dynamic range standards and maintain high color accuracy; and the system integration and virtual device unit creates virtual devices based on a standard framework, achieving seamless backhaul and compatibility of enhanced audio and video streams, ensuring overall low-latency transmission and efficient resource utilization, and seamlessly integrating with existing workflows.

[0129] In summary, this application constructs a comprehensive capability enhancement system for KVM-based conferences by deeply integrating high-fidelity source capture, multimodal real-time analysis, and generative AI enhancement technologies through a full-link collaborative architecture. At the source capture end, it supports cross-resolution, high-frame-rate, and multi-color encoding format video stream adaptation and lossless transmission of multimodal audio formats, ensuring end-to-end low-latency capture. By integrating a large voice service model, it achieves accurate and efficient processing of speech recognition and sensitive content review. Relying on a dynamic rule engine to drive pixel-level enhancement of generative AI models such as behavior repair, scene completion, and face virtualization, combined with nanosecond-level timestamp alignment technology, it ensures audio and video synchronization. Finally, through virtual device units, it achieves professional-grade display compatibility and high dynamic range presentation. While significantly reducing system latency, it enhances the security of conference content with high recall and precision, possessing multiple values: improving conference security, reducing deployment costs by 60%, and promoting the deep integration of AI and KVM. This meets the stringent requirements of professional scenarios for real-time performance, accuracy, and security.

[0130] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0131] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0133] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0134] Further reference Figure 6 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an audio and video content enhancement device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0135] like Figure 6 As shown, the content enhancement device 200 applied to audio and video in this embodiment includes:

[0136] The high-efficiency video encoding acquisition module 210 is used to receive high-efficiency video encoding sent by the sending end of the KVM agent collaboration system;

[0137] The audio pixel decoding module 220 is used to perform audio pixel decoding processing on high-efficiency video encoding to obtain the original audio pixel data.

[0138] The multimodal structured tagging module 230 is used to perform multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured tagging data;

[0139] The decision instruction generation module 240 is used to generate decision instructions based on a preset dynamic rule base and multimodal structured identifier data;

[0140] Data augmentation module 250 is used to perform data augmentation processing on the original audio pixel data according to the decision instructions to obtain enhanced audio pixel data;

[0141] Z-shaped reconstruction and repair module 260 is used to perform Z-shaped reconstruction and repair processing on enhanced audio pixel data to obtain enhanced audio and video streams;

[0142] The audio and video stream mapping module 270 is used to map the enhanced audio and video stream to a standard audio and video source and send the standard audio and video source back to the KVM system.

[0143] In this embodiment, a content enhancement device 200 for audio and video is provided, comprising: a high-efficiency video encoding acquisition module 210, used to receive high-efficiency video encoding sent by a transmitter of a KVM workstation collaboration system; an audio pixel decoding module 220, used to perform audio pixel decoding processing on the high-efficiency video encoding to obtain original audio pixel data; a multimodal structured tagging module 230, used to perform multimodal structured tagging processing on the original audio pixel data to obtain multimodal structured identifier data; a decision instruction generation module 240, used to generate decision instructions based on a preset dynamic rule base and multimodal structured identifier data; a data enhancement module 250, used to perform data enhancement processing on the original audio pixel data according to the decision instructions to obtain enhanced audio pixel data; a Z-shaped reconstruction and repair module 260, used to perform Z-shaped reconstruction and repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream; and an audio and video stream mapping module 270, used to map the enhanced audio and video stream to a standard audio and video source and send the standard audio and video source back to the KVM system. Compared with existing technologies, this application effectively enhances audio and video quality and maps the enhanced audio and video streams as standard audio and video sources back to the KVM system, providing high-quality audio and video support for the system.

[0144] In some optional implementations of the embodiments of this application, the aforementioned multimodal structured identification data includes identity identification information, the aforementioned decision instructions include enabling virtual avatars and disabling virtual avatars, the multimodal structured tagging module includes: an avatar pixel recognition submodule, a similarity matching submodule, a filtering submodule, and an identity identification tagging submodule, the decision instruction generation module includes: an identity identification judgment submodule, a first decision instruction generation submodule, and a second decision instruction generation submodule, and the data augmentation module includes: a first data augmentation submodule and a second data augmentation submodule, wherein:

[0145] The avatar pixel recognition submodule is used to perform avatar pixel recognition processing on the raw pixel data of the raw audio pixel data to obtain the raw avatar pixel data.

[0146] The similarity matching submodule is used to read the system database, perform similarity matching processing on the original avatar pixel data based on the preset avatar data in the system database, and obtain the similarity matching result;

[0147] The filtering submodule is used to filter the similarity matching results according to preset avatar matching conditions to obtain the avatar filtering results;

[0148] The identity tagging submodule is used to obtain identity information corresponding to the avatar screening results, and to tag the original audio pixel data according to the identity information to obtain multimodal structured tag data;

[0149] The identity identification determination submodule is used to determine whether identity identification information belongs to the privacy protection object;

[0150] The first decision instruction generation submodule is used to generate a decision instruction to enable the virtual avatar if the identity information belongs to the privacy protection object.

[0151] The second decision instruction generation submodule is used to generate a decision instruction to disable the virtual avatar if the identity information does not belong to the privacy protection object.

[0152] The first data augmentation submodule is used to generate dynamic 3D virtual image pixel data corresponding to the original avatar pixel data according to the MediaPipe architecture when the decision instruction is to enable virtual image, and replace the original avatar pixel data of the original audio pixel data with the dynamic 3D virtual image pixel data to obtain augmented audio pixel data.

[0153] The second data augmentation submodule is used to prevent data augmentation of the original audio pixel data when the decision instruction is to disable the virtual avatar.

[0154] In some optional implementations of the embodiments of this application, the aforementioned multimodal structured identification data includes behavioral identification information, the aforementioned decision instructions include triggering virtual subtitle prompts and not triggering virtual subtitle prompts, the multimodal structured tagging module includes: a behavioral recognition submodule and a behavioral identification tagging submodule, the decision instruction generation module includes: a behavioral identification judgment submodule, a third decision instruction generation submodule and a fourth decision instruction generation submodule, and the data augmentation module includes: a third data augmentation submodule and a fourth data augmentation submodule, wherein:

[0155] The behavior recognition submodule is used to perform behavior recognition processing on the raw pixel data of the raw audio pixel data according to the HRNet-W48+Bi-LSTM architecture to obtain behavior identification information.

[0156] The behavior identification tagging submodule is used to tag the raw audio pixel data according to the behavior identification information to obtain multimodal structured identification data;

[0157] The behavior identification judgment submodule is used to obtain system mode information corresponding to the original audio pixel data, and to determine whether the behavior identification information belongs to non-scene actions based on the system mode information.

[0158] The third decision instruction generation submodule is used to generate a decision instruction that triggers virtual subtitle prompts if the behavior identification information belongs to a non-scene action.

[0159] The fourth decision instruction generation submodule is used to generate a decision instruction that does not trigger virtual subtitle prompts if the behavior identification information does not belong to a non-scene action.

[0160] The third data enhancement submodule is used to generate behavioral reminder subtitles from the original audio pixel data when the decision instruction is to trigger virtual subtitle prompts, thus obtaining enhanced audio pixel data;

[0161] The fourth data enhancement submodule is used to prevent data enhancement of the original audio pixel data when the decision instruction is not to trigger virtual subtitle prompts.

[0162] In some optional implementations of the embodiments of this application, the above-mentioned data enhancement module includes:

[0163] The history cache retrieval submodule is used to retrieve the history cache data corresponding to the original audio pixel data;

[0164] The standard pose pixel acquisition submodule is used to perform semantic-guided translation of the human body region in the current frame of the original audio pixel data based on the StarGAN-v2 architecture and frame history cache data, and generate standard pose pixel data.

[0165] The Standard Pose Pixel Replacement submodule is used to replace the human body region in the current frame of the original audio pixel data with standard pose pixel data.

[0166] In some optional implementations of the embodiments of this application, the multimodal structured identifier data includes sensitive word identifier information, the decision instruction includes an audio mute instruction, the multimodal structured tagging module includes an audio recognition submodule, a sensitive word matching submodule, and a sensitive word tagging submodule, the decision instruction generation module includes an audio mute instruction generation submodule, and the data enhancement module includes an audio mute submodule, wherein:

[0167] The audio recognition submodule is used to perform audio recognition processing on the raw audio pixel data based on the speech recognition model to obtain the raw audio text data.

[0168] The sensitive word matching submodule is used to perform sensitive word matching on the original audio text data according to the sensitive word library to obtain sensitive word identification information;

[0169] The sensitive word tagging submodule is used to tag the original audio pixel data according to the sensitive word tagging information to obtain multimodal structured tagging data;

[0170] The audio mute instruction generation submodule is used to obtain audio mute instructions corresponding to sensitive word identification information from the dynamic rule base;

[0171] The audio mute submodule is used to mute the original audio pixel data according to the audio mute command, so as to obtain enhanced audio pixel data.

[0172] In some optional implementations of the embodiments of this application, the aforementioned standard audio and video source includes a standard video source and a standard audio source, and the aforementioned audio and video stream mapping module includes:

[0173] The virtual camera device creation submodule is used to create virtual camera devices based on the V4L2 driver layer;

[0174] The video stream mapping submodule is used to map the enhanced video stream of the enhanced audio and video stream to a standard video source based on the virtual camera device;

[0175] The Virtual Microphone Device Creation Submodule is used to create a virtual microphone device based on the ALSA driver layer;

[0176] The audio stream mapping submodule is used to map the enhanced audio stream of the enhanced audio / video stream to a standard audio source based on the virtual microphone device.

[0177] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 7 , Figure 7 This is a basic structural block diagram of a computer device according to an embodiment of this application.

[0178] Computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected via a system bus. It should be noted that only computer device 300 with components 310-330 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0179] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0180] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as the hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device 300. In the embodiments of this application, the memory 310 is typically used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions applied to audio and video content enhancement methods. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.

[0181] In some embodiments, processor 320 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 320 is typically used to control the overall operation of computer device 300. In embodiments of this application, processor 320 is used to execute computer-readable instructions stored in memory 310 or to process data, such as executing computer-readable instructions applied to audio and video content enhancement methods.

[0182] The network interface 330 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 300 and other electronic devices.

[0183] The computer equipment provided in this application effectively enhances audio and video quality and maps the enhanced audio and video streams as standard audio and video sources back to the KVM system, providing high-quality audio and video support for the system.

[0184] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described content enhancement method applied to audio and video.

[0185] The computer-readable storage medium provided in this application effectively enhances audio and video quality and maps the enhanced audio and video streams as standard audio and video sources back to the KVM system, providing high-quality audio and video support for the system.

[0186] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0187] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A content enhancement method applied to audio-video, characterized in that, The method is applied to a receiving end of a KVM agent collaboration system, and comprises the following steps: Receiving high-efficiency video coding sent by a sending end of the KVM agent collaboration system; Performing audio pixel decoding processing on the high-efficiency video coding to obtain original audio pixel data; Performing multi-modal structured labeling processing on the original audio pixel data to obtain multi-modal structured identification data; Generating a decision instruction according to a preset dynamic rule base and the multi-modal structured identification data; Performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data; Performing Z-shaped reorganization repair processing on the enhanced audio pixel data to obtain an enhanced audio and video stream; Mapping the enhanced audio and video stream into a standard audio and video source, and returning the standard audio and video source to the KVM system.

2. The content enhancement method for audio-video according to claim 1, wherein, The multi-modal structured identification data comprises identity information, and the step of performing multi-modal structured labeling processing on the original audio pixel data to obtain multi-modal structured identification data comprises the following steps: Performing portrait pixel recognition processing on original pixel data of the original audio pixel data to obtain original portrait pixel data; Reading a system database, performing similarity matching processing on the original portrait pixel data according to preset portrait data of the system database to obtain a similarity matching result; Performing screening processing on the similarity matching result according to a preset portrait matching condition to obtain a portrait screening result; Obtaining identity information corresponding to the portrait screening result, and performing labeling processing on the original audio pixel data according to the identity information to obtain the multi-modal structured identification data; The decision instruction comprises enabling a virtual image and not enabling the virtual image, and the step of generating a decision instruction according to a preset dynamic rule base and the multi-modal structured identification data comprises the following steps: Determining whether the identity information belongs to a privacy protection object; If the identity information belongs to a privacy protection object, a decision instruction for enabling a virtual image is generated; If the identity information does not belong to a privacy protection object, a decision instruction for not enabling a virtual image is generated; The step of performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data comprises the following steps: When the decision instruction is to enable a virtual image, dynamic 3D virtual image pixel data corresponding to the original portrait pixel data is generated according to a MediaPipe architecture, and the original portrait pixel data of the original audio pixel data is replaced with the dynamic 3D virtual image pixel data to obtain the enhanced audio pixel data; When the decision instruction is not to enable a virtual image, the original audio pixel data is not subjected to data enhancement processing.

3. The content enhancement method for audio-visuals as claimed in claim 1 wherein, The multi-modal structured identification data comprises behavior identification information, and the step of performing multi-modal structured labeling processing on the original audio pixel data to obtain multi-modal structured identification data comprises the following steps: According to the HRNet-W48+Bi-LSTM architecture, the original pixel data of the original audio pixel data is subjected to behavior recognition processing to obtain the behavior identification information; According to the behavior identification information, the original audio pixel data is subjected to labeling processing to obtain the multi-modal structured identification data; The decision instruction includes triggering a virtual subtitle prompt and not triggering a virtual subtitle prompt, and the step of generating a decision instruction according to a preset dynamic rule library and the multi-modal structured identification data specifically includes the following steps: Obtain system mode information corresponding to the original audio pixel data, and determine whether the behavior identification information belongs to a non-scene action according to the system mode information; If the behavior identification information belongs to a non-scene action, a decision instruction triggering a virtual subtitle prompt is generated; If the behavior identification information does not belong to a non-scene action, a decision instruction not triggering a virtual subtitle prompt is generated; The step of performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data specifically includes the following steps: When the decision instruction is to trigger a virtual subtitle prompt, an action reminder subtitle is generated in the original audio pixel data to obtain the enhanced audio pixel data; When the decision instruction is not to trigger a virtual subtitle prompt, no data enhancement processing is performed on the original audio pixel data.

4. The content enhancement method for audio-video according to claim 3, wherein, The step of performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data specifically includes the following steps: Obtain historical cache data corresponding to the original audio pixel data; According to the StarGAN-v2 architecture and the frame historical cache data, the human body region in the current frame of the original audio pixel data is subjected to semantic guided translation to generate standard posture pixel data; The human body region in the current frame of the original audio pixel data is replaced with the standard posture pixel data.

5. The content enhancement method for audio-visuals as claimed in claim 1 wherein, The multi-modal structured identification data includes sensitive word identification information, and the step of performing multi-modal structured labeling processing on the original audio pixel data to obtain multi-modal structured identification data specifically includes the following steps: According to a speech recognition model, the original audio data of the original audio pixel data is subjected to audio recognition processing to obtain original audio text data; According to a sensitive word library, the original audio text data is subjected to sensitive word matching processing to obtain sensitive word identification information; According to the sensitive word identification information, the original audio pixel data is subjected to labeling processing to obtain the multi-modal structured identification data; The decision instruction includes an audio mute instruction, and the step of generating a decision instruction according to a preset dynamic rule library and the multi-modal structured identification data specifically includes the following steps: In the dynamic rule library, an audio mute instruction corresponding to the sensitive word identification information is obtained; The step of performing data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data specifically includes the following steps: According to the audio mute instruction, the original audio pixel data is subjected to audio mute processing to obtain the enhanced audio pixel data.

6. The content enhancement method for audio-visuals as claimed in claim 1 wherein, The standard audio-video source includes a standard video source and a standard audio source, and the step of mapping the enhanced audio-video stream into the standard audio-video source and returning the standard audio-video source to the KVM system specifically includes the following steps: A virtual camera device is created according to the V4L2 driver layer; An enhanced video stream of the enhanced audio-video stream is mapped into the standard video source according to the virtual camera device; A virtual microphone device is created according to the ALSA driver layer; An enhanced audio stream of the enhanced audio-video stream is mapped into the standard audio source according to the virtual microphone device.

7. A content enhancement device applied to audio and video, characterized in that, It comprises: An HEVC acquisition module is configured to receive HEVC sent by a sending terminal of the KVM agent collaboration system; An audio pixel decoding module is configured to perform audio pixel decoding processing on the HEVC to obtain original audio pixel data; A multi-modal structured labeling module is configured to perform multi-modal structured labeling processing on the original audio pixel data to obtain multi-modal structured identification data; A decision instruction generation module is configured to generate a decision instruction according to a preset dynamic rule base and the multi-modal structured identification data; A data enhancement module is configured to perform data enhancement processing on the original audio pixel data according to the decision instruction to obtain enhanced audio pixel data; A Z-shaped reorganization repair module is configured to perform Z-shaped reorganization repair processing on the enhanced audio pixel data to obtain an enhanced audio-video stream; An audio-video stream mapping module is configured to map the enhanced audio-video stream into a standard audio-video source and return the standard audio-video source to the KVM system.

8. The content enhancement device for audio-video as claimed in claim 7, wherein, The multi-modal structured identification data includes identity information, the multi-modal structured labeling module includes a portrait pixel recognition submodule, a similarity matching submodule, a screening submodule, and an identity identification labeling submodule, the decision instruction generation module includes an identity identification judgment submodule, a first decision instruction generation submodule, and a second decision instruction generation submodule, and the data enhancement module includes a first data enhancement submodule and a second data enhancement submodule, wherein: The portrait pixel recognition submodule is configured to perform portrait pixel recognition processing on original pixel data of the original audio pixel data to obtain original portrait pixel data; The similarity matching submodule is configured to read a system database, perform similarity matching processing on the original portrait pixel data according to preset portrait data of the system database to obtain a similarity matching result; The screening submodule is configured to perform screening processing on the similarity matching result according to a preset portrait matching condition to obtain a portrait screening result; The identity identification labeling submodule is configured to obtain identity information corresponding to the portrait screening result and perform labeling processing on the original audio pixel data according to the identity information to obtain the multi-modal structured identification data; The identity identification judgment submodule is configured to judge whether the identity information belongs to a privacy protection object. The first decision instruction generation submodule is configured to generate a decision instruction to enable a virtual image if the identity information belongs to a privacy protection object. The second decision instruction generation submodule is configured to generate a decision instruction not to enable a virtual image if the identity information does not belong to a privacy protection object. The first data enhancement submodule is configured to, when the decision instruction is to enable a virtual image, generate dynamic 3D virtual image pixel data corresponding to the original avatar pixel data according to a MediaPipe architecture, replace original avatar pixel data of the original audio pixel data with the dynamic 3D virtual image pixel data, and obtain the enhanced audio pixel data. The second data enhancement submodule is configured to, when the decision instruction is not to enable a virtual image, not perform data enhancement processing on the original audio pixel data. 9.A computer device, comprising a memory and a processor, and characterized in that, The memory stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the content enhancement method for audio and video according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the steps of the content enhancement method for audio and video according to any one of claims 1 to 6.