Cloud networking monitoring video lightweight processing method and system based on U-ACE

Through the U-ACE unified AI core engine, the multi-task module is integrated and the deep collaboration technology is adopted, the information bottlenecks and resource waste problems of traditional separated AI models in cloud network monitoring video processing, and efficient and intelligent video processing and user experience optimization are achieved.

CN120388322AActive Publication Date: 2025-07-29JINHUA LINGJIANG INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510781094.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-29
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The traditional separated AI model has problems of information bottlenecks, waste of resources and poor scenario adaptability in cloud networked surveillance video processing, making it difficult to achieve end-to-end optimization and efficient processing.

Method used

U-ACE's unified AI core engine is adopted, integrating the perception and priority decision module (PPM), the deep analysis and prediction enhancement module (APEM), the intelligent compression and lightweight module (ICM) and the user preference and intelligent presentation module (UPSM). Through shared encoder, attention modulation mechanism and joint training, deep collaboration and information sharing between each module is achieved.

Benefits of technology

It improves the efficiency and resource utilization of surveillance video processing, enhances the analysis quality of key content, optimizes the user experience, and responds to scene changes through an adaptive optimization mechanism to achieve lightweight and intelligent system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388322A_ABST
    Figure CN120388322A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud networking monitoring video lightweight processing method and a cloud networking monitoring video lightweight processing system based on U-ACE. The core of the method is to construct and operate a multi-task and multi-mode AI core engine U-ACE. The processing flow comprises the following steps of: sensing an input video in real time and judging a processing priority through a sensing and priority decision module PPM of the U-ACE; for high-priority content, a deep analysis and prediction enhancement module APEM of the U-ACE is used for analysis, prediction and enhancement; self-adaptive content compression is carried out on a non-high-priority or non-key area through an intelligent compression module (ICM) of the U-ACE; in combination with user preferences, enhanced display and layered rendering are carried out by using the user preferences of the U-ACE and an intelligent presentation module UPSM; and running parameters and strategies for dynamically adjusting the engine are fed back in real time through an engine self-optimization module ESOM of the U-ACE. The processing efficiency is improved, the resource consumption is reduced, and the key content presentation and the user experience are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a lightweight processing method for cloud-connected monitoring videos based on U-ACE. Background Art

[0002] In the field of modern video monitoring and processing, cloud-connected monitoring systems generate a vast amount of video data, and the real-time, efficient, and intelligent processing of this data is a major challenge currently faced. Traditional video processing methods and early AI-based methods usually adopt multiple separate AI models or modules for specific subtasks (such as scene perception, object detection, video encoding, user interaction analysis, etc.). This separated architecture has many inherent limitations: First, each separated model is usually designed and trained independently, and the information interaction between them is often one-way and coarse-grained, lacking deep feature sharing and fine-grained real-time collaboration mechanisms. This leads to information bottlenecks and redundant calculations in the overall processing flow, making it difficult to achieve end-to-end smooth optimization. Second, due to the separation of models, it is difficult to optimize computing resources and storage resources uniformly and dynamically from a global perspective of the system. Each model may pursue the optimal performance of its own task and consume excessive resources, resulting in low overall resource utilization, especially in complex and changing monitoring scenarios, where it is unable to adapt flexibly. Third, the overall response ability of the separated models to scene changes is weak. For example, the judgment of the scene perception module may not be transmitted to the analysis module and the compression module in a timely and effective manner to adjust their internal strategies, resulting in a decline in system performance or loss of key information in the event of an emergency or drastic environmental change. Finally, improving the entire system requires coordinating the upgrades of multiple independent models, resulting in high development and maintenance costs.

[0003] Therefore, there is an urgent need for a core AI technology that can overcome the limitations of traditional separated models, achieve breakthroughs in the intelligence, lightweight, adaptability, and overall performance of monitoring video processing, and effectively address the technical challenges faced in building such a unified system. Summary of the Invention

[0004] The object of the present invention is to provide a method and system for intelligent lightweight processing of cloud-connected monitoring videos based on U-ACE. By constructing a highly integrated, multi-task, multi-modal AI core engine (Unified AI Core Engine, U-ACE), it realizes the full-link integrated intelligent processing of monitoring videos from perception, analysis, prediction, enhancement, compression to presentation and optimization. The present invention not only proposes the architecture of U-ACE, but also proposes specific solutions to the technical challenges faced in constructing such a unified engine (such as multi-task conflicts, difficult gradient optimization, complex information flow control, large computational resource requirements, etc.), and reveals the resulting synergistic gains, thereby overcoming the limitations of traditional separate models, improving processing efficiency, reducing resource consumption, enhancing the quality of key content, and optimizing the user experience.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for lightweight processing of cloud-connected monitoring videos based on U-ACE, which is executed by a multi-task, multi-modal AI core engine (U-ACE). The U-ACE adopts a deep Encoder-Decoder architecture and integrates a Perception and Priority Decision Module (PPM), a Deep Analysis and Prediction Enhancement Module (APEM), an Intelligent Compression and Lightweight Module (ICM), a User Preference and Intelligent Presentation Module (UPSM), and an Engine Self-Optimization Module (ESOM). The method at least includes the following steps: Real-time perception and priority decision step S1: The PPM of the U-ACE uses the early features extracted by the shared encoder of the U-ACE to real-time perceive the scene type, complexity and change situation of the input monitoring video, and combines with the dynamic resource management unit inside the U-ACE to dynamically determine the processing priority of the video content and preliminarily plan the computing resources; High-priority content deep processing step S2: For the video content determined to be of high priority in step S1, the APEM of the U-ACE uses the deep spatio-temporal features extracted by the shared encoder to perform refined semantic analysis, key information adaptive enhancement, and forward prediction of the scene dynamic change pattern; among them, the scene perception result output by the PPM in step S1 affects the utilization of the shared encoder features by the APEM through the attention modulation mechanism; the prediction result of the APEM is internally fed back to the PPM to assist the dynamic adjustment of its priority decision and guide the enhancement parameters of the APEM itself; Intelligent Compression of Non-Critical Content - Step S3: For the content determined to be of non-high priority in Step S1 or the non-critical regions in the video frames identified by APEM in Step S2, the ICM of the U-ACE uses the features extracted by the shared encoder or the output of the PPM and performs deep compression using a content-adaptive neural compression algorithm. Its compression strategy is jointly modulated by the priority determined in Step S1 and the analysis results of APEM in Step S2; User-Preference-Driven Intelligent Rendering - Step S4: The UPSM of the U-ACE learns the user's historical interaction data to build a user preference model, and combines the high-priority content processed in Step S2 and the data compressed in Step S3 to perform customized enhanced display and refined hierarchical rendering of the content of user interest; Engine Closed-Loop Self-Optimization - Step S5: The ESOM of the U-ACE, based on deep reinforcement learning, continuously monitors the overall operating state of the U-ACE, the performance metrics of each module such as PPM, APEM, ICM, UPSM, and user feedback, and dynamically adjusts the decision threshold of PPM in Step S1, the analysis depth and enhancement level of APEM in Step S2, the compression intensity of ICM in Step S3, the rendering strategy of UPSM in Step S4, and the operating parameters of the shared encoder and the dynamic resource management unit inside the U-ACE; Among them, the U-ACE realizes the efficient cooperation and information sharing of the functions of each module in Steps S1 to S5 through the shared encoder features, the attention modulation mechanism, and the end-to-end joint training strategy including the co-gradient regularization term, so as to overcome the limitations brought by the traditional separate models.

[0006] Preferably, the dynamic change pattern prediction part of APEM in Step S2 uses a multi-head prediction decoder (Multi-Head Prediction Decoder). This decoder, based on the features encoded by CST-Transformer, predicts multiple types of future information in parallel and ensures the cooperation and rationality between different prediction heads through the internal prediction consistency regularization loss (Prediction Consistency Regularization Loss) to cope with the target conflict problem in multi-task learning.

[0007] Preferably, the neural compression algorithm of ICM in Step S3 uses a learned context-based autoregressive entropy model (Learned Context-based Autoregressive Entropy Model), and combines the semantic segmentation map or saliency map output by APEM in Step S2 to allocate fewer bits to the latent variable representation of the non-critical regions, realizing semantic-guided efficient compression, so that the prediction module of APEM in Step S2 indirectly benefits from the image compressibility features learned by ICM, generating a collaborative gain.

[0008] Preferably, the user preference model of the UPSM in step S4 is an Explainable Graph Attention Network (EGAT).

[0009] Preferably, the action space of the deep reinforcement learning agent of the ESOM in step S5 is designed as a hierarchical structure (Hierarchical Action Space). The top-level action selects the macro policy, and the bottom-level action fine-tunes the specific parameters of each module of the U-ACE in steps S1 to S4 under the selected macro policy to improve the learning efficiency and the robustness of the policy and cope with the optimization challenges brought by the huge parameter space of the unified engine.

[0010] Preferably, a simplified mathematical form of the collaborative gradient regularization term is L_sgr = λ_sgr * Σ_ {i≠j} (1 - cos(grad(L_i), grad(L_j))), where L_i and L_j are the gradients of the loss functions of different task modules that execute steps S1 to S4 inside the U-ACE on the shared parameters, cos(·, ·) represents the cosine similarity, and λ_sgr is the regularization coefficient. This regularization term promotes the collaboration of parameter updates by punishing the significant inconsistency in the gradient directions of different tasks to alleviate the gradient conflict and catastrophic forgetting problems in multi-task learning.

[0011] Preferably, the dynamic resource management unit inside the U-ACE can achieve fine-grained resource adaptability inside the engine based on the priority determination of the PPM in step S1 and the optimization instructions of the ESOM in step S5, specifically including: adjusting the effective computational depth by selecting different variants of the shared encoder model with preset different computational depths or using a dynamic network structure that supports early exit, dynamically activating or deactivating different decoder branches in the APEM for specific tasks, and selecting different complexity levels of the neural compression models preset in the ICM.

[0012] Preferably, the attention modulation mechanism in step S2 includes: the scene category embedding vector e_sc and the complexity score c output by the PPM in step S1 are multiplied or concatenated channel by channel with the feature map F_ape_in obtained by the APEM from the shared encoder and then processed through a small convolutional network to generate the modulated feature map F_ape_mod = AttentionModulator(F_ape_in, e_sc, c), where the AttentionModulator network learns to dynamically adjust the channel response or the spatial attention region of F_ape_in according to the scene information, enabling the APEM to more effectively utilize the features most relevant to the current scene.

[0013] The present invention also discloses a lightweight processing system for cloud-connected monitoring videos based on U-ACE, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned lightweight processing method for cloud-connected monitoring videos based on U-ACE is implemented.

[0014] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned lightweight processing method for cloud-connected monitoring videos based on U-ACE is implemented.

[0015] The technical solution of the present invention may include the following beneficial effects: The present invention discloses a lightweight processing method and system for cloud-connected monitoring videos based on U-ACE. Through a unified U-ACE architecture, a shared encoder, an internal attention modulation mechanism, and joint training, unprecedented deep cooperation and information sharing among functional modules are achieved; feature reuse avoids redundant calculations; ICM achieves precise depth compression under the guidance of PPM and APEM; the global optimization of ESOM ensures the allocation of resources according to demand, thereby realizing the resource utilization efficiency and system lightweight; U-ACE responds to scene changes as a whole, and the strategies of its internal modules can be adjusted quickly and consistently, and effectively solves technical problems such as multi-task conflicts and information flow control through an internal cooperation mechanism; APEM tightly couples analysis, prediction, and enhancement to form a positive cycle, improving the analysis quality of key content; ESOM performs end-to-end adaptive optimization of the entire U-ACE based on hierarchical DRL to ensure the long-term efficient operation of the system; high integration may bring performance improvements such as compressed feature-assisted prediction, further enhancing the overall intelligence level of the system. Description of the Drawings

[0016] Figure 1 It is a schematic diagram of the overall functional architecture and internal main data flow of the unified AI core engine U-ACE in the method of the present invention.

[0017] Figure 2 It is an exemplary multi-level structural block diagram of the cascaded spatio-temporal Transformer (CST-Transformer) encoder adopted in U-ACE of the present invention.

[0018] Figure 3 It is an exemplary network structural block diagram of the multi-head prediction decoder and conditional GAN enhancement part of the APEM module in U-ACE of the present invention.

[0019] Figure 4 It is a schematic diagram of the hierarchical reinforcement learning agent adopted by the ESOM module in U-ACE of the present invention interacting with U-ACE for self-optimization.

[0020] Figure 5 It is a line chart comparing the performance of the method of the present invention (U-ACE) with that of a benchmark algorithm (separation model) in terms of the mean average precision (mAP) of key object detection on different test subsets.

[0021] Figure 6 It is a comparison chart showing the "synergistic effect" of the method of the present invention (U-ACE) with a benchmark algorithm and ablation variants of U-ACE in terms of the comprehensive performance score. Detailed implementation manners

[0022] To further understand the content of the present invention, the present invention will be described in detail with reference to the accompanying drawings and embodiments. The following further details the present application with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that only parts related to the invention are shown in the drawings for the sake of convenience of description.

[0023] As Figure 1 shown, a lightweight processing method for cloud-connected network monitoring videos based on U-ACE in this embodiment is executed by a unified multi-task and multi-modal AI core engine (Unified AI Core Engine, U-ACE). The U-ACE adopts a deep Encoder-Decoder architecture and integrates a perception and priority decision-making module (PPM), a deep analysis and prediction enhancement module (APEM), an intelligent compression and lightweight module (ICM), a user preference and intelligent presentation module (UPSM), and an engine self-optimization module (ESOM). Specifically, it may include the following steps: Real-time perception and priority decision-making step S1: Executed by the PPM module of U-ACE. The PPM uses the shared encoder of U-ACE (for example, a cascaded spatio-temporal Transformer, CST-Transformer, see the appendix Figure 2The early features extracted (such as the output features F_s1 and F_s2 from Stage1 or Stage2 of CST-Transformer, which retain more details and are relatively fast to compute). PPM real-time perceives the scene type of the input monitoring video (such as outputting scene category labels through a lightweight MLP classification head), the scene complexity (such as outputting a complexity score of 0-1 through an MLP regression head), and the significant changes in the scene (such as by comparing feature differences within consecutive time windows or training a dedicated change detection head). Based on these perception results, PPM interacts with the dynamic resource management unit (RMU) inside U-ACE, dynamically determines the processing priority of the current video content (for example, divided into high, medium, and low levels), and preliminarily plans the computing resources to be consumed by subsequent modules such as APEM and ICM (for example, suggesting that RMU allocate more GPU cores for high-priority content, or activating a more complex analysis branch in APEM).

[0024] High-priority content in-depth processing step S2: For the video content determined to be high-priority by PPM in step S1, it is executed by the APEM module of U-ACE. APEM uses the deep spatio-temporal features extracted by the U-ACE shared encoder (especially its deep output or fused features, such as F_s3 or F_shared of CST-Transformer) for refined processing: Refined semantic analysis: Build a target detection head (such as a Transformer-based DETR decoder) and an instance segmentation head (such as a Mask2Former decoder) on top of the F_shared feature, and output high-precision target bounding boxes, category labels, instance masks, and possible action recognition results.

[0025] Adaptive enhancement of key information: Adopt a conditional generative adversarial network (Conditional GAN). Its generator takes the F_shared feature and the scene conditions determined by PPM in step S1 (such as low light, rain, fog, etc.) as inputs, and generates enhanced video frames (such as denoising, super-resolution, de-motion blur). The discriminator then evaluates the authenticity and quality of the generated results.

[0026] Forward-looking prediction of scene dynamic change patterns: A multi-head prediction decoder, based on the F_shared feature, parallelly predicts the key target movement trajectories in the short term in the future, the occurrence probability of medium-term scene events (such as crowd gathering trends), etc.

[0027] Collaborative mechanism 1 - Attention modulation: The scene perception results output by the PPM in step S1 (such as the scene category embedding vector e_sc and the complexity score c) affect the utilization of the shared encoder features F_ape_in by the APEM through a specific attention modulation mechanism (AttentionModulator). For example, AttentionModulator can be a small neural network, whose inputs are F_ape_in, e_sc, and c, and the output is the modulated feature map F_ape_mod = AttentionModulator(F_ape_in, e_sc, c). This modulation process can learn to achieve channel dimension recalibration of F_ape_in (amplifying the feature channels related to the current scene) or spatial dimension region focusing (making the APEM pay more attention to specific spatial regions related to the current scene type or complexity).

[0028] Collaborative mechanism 2 - Prediction feedback: The prediction results of the APEM (especially the medium-term event probability and short-term trajectory) are internally fed back to the PPM module in step S1. The decision-making network of the PPM can take these prediction features as additional inputs to achieve more forward-looking priority adjustment. At the same time, the prediction results of the APEM also dynamically guide its own enhancement parameters (for example, allocating more enhancement resources in advance to the targets on the predicted trajectory).

[0029] Intelligent compression step S3 for non-critical content: For the content determined by the PPM in step S1 as not of high priority, or the non-critical regions (such as static backgrounds, non-attended objects) in the video frames identified by the APEM in step S2, it is executed by the ICM module of the U-ACE. The ICM uses the features extracted by the U-ACE shared encoder (which can be shallow or middle-layer features F_s1, F_s2 with less computational complexity) or the direct output of the PPM (such as the saliency map), and adopts a content-adaptive neural compression algorithm (for example, a context-based autoregressive entropy model).

[0030] Collaborative mechanism 3 - Compression strategy modulation: The compression strategies of the ICM (such as the target bitrate, quantization parameter) are jointly and finely modulated by the priority (P) determined by the PPM in step S1 and the analysis results of the APEM in step S2 (such as the semantic segmentation map, indicating which regions are backgrounds and which are objects of specific categories), to achieve differential deep compression of "one frame, one strategy; one region, one strategy".

[0031] User Preference-Driven Intelligent Rendering Step S4: Executed by the UPSM module of U-ACE. The UPSM learns the user's historical interaction data to build a user preference model, and combines the high-priority content processed in Step S2 and the data compressed in Step S3 to perform customized enhanced display and refined hierarchical rendering of the content the user is interested in. The data integration and rendering process is as follows: In Step S2, APEM analyzes and enhances high-priority video content (such as raw pixel data or pre-decoded data) (e.g., through Conditional GAN). The enhanced high-priority content (e.g., pixel data of several key frames or short video segments) is temporarily stored in a high-quality buffer or, if necessary, undergoes a lightweight, near-lossless intra-frame encoding (e.g., using JPEG-LS or low-parameter intra-frame HEVC / AV1 encoding) for easy management and fast decoding; In Step S3, ICM performs deep neural compression on non-high-priority content or non-critical regions to generate a compressed bitstream. During rendering, this part of the bitstream is decoded in real time to restore the pixel data, but its quality is generally lower than that of high-priority content; The UPSM dynamically obtains the pixel data of the corresponding regions from the high-quality buffer (for the high-priority regions the user is interested in) and the output of the ICM decoder (for other regions) according to the current user's attention region (predicted by the user preference model or determined by real-time interaction) and the content priority of each region. Subsequently, the UPSM uses a graphics rendering interface (such as OpenGL, Vulkan, DirectX) to synthesize these image regions from different sources and of different qualities into the final display frame in real time. For example, the target the user is interested in and its adjacent regions use the high-quality data enhanced by APEM, while the distant background uses the data after ICM compression and decoding. The UPSM will handle the smooth transition between different quality regions (e.g., through Alpha blending or edge feathering) to enhance the visual effect. Through this mechanism, the UPSM can significantly reduce the overall data volume while ensuring the high-quality presentation of key content, achieving efficient intelligent rendering.

[0032] First, the UPSM analyzes the user's long-term historical interaction data (viewing records, zooming, panning, marking, etc.) and possible explicit feedback through its internal user preference learning sub-module (e.g., using the Explainable Graph Attention Network EGAT with enhanced interpretability) to build a dynamically updated and interpretable user preference model. Then, the UPSM combines this preference model, the high-quality and high-priority content processed by APEM in Step S2, and the lightweight data compressed by ICM in Step S3 to perform intelligent rendering: Customized Enhanced Display: Clearly overlay the analysis results of APEM in Step S2 (such as detection boxes, trajectories, event tags) on the video, and highlight specific targets or event trajectories according to user preferences.

[0033] Fine-grained hierarchical rendering: Dynamically adjust the rendering quality according to the focus area that the user is currently interested in (which can be determined by mouse hovering, eye tracking, or prediction using the user preference model) and the priority of the content itself. For the focus area that the user highly concerns, preferentially use the high-quality data enhanced by APEM in step S2; for other areas, use the data compressed by ICM in step S3 or the low-resolution version.

[0034] Engine closed-loop self-optimization step S5: Executed by the ESOM module of U-ACE. ESOM adopts deep reinforcement learning (DRL) technology (e.g., hierarchical DRL) and acts as a meta-controller to adaptively optimize the entire U-ACE. The DRL agent of ESOM first goes through a large number of offline training phases, exploring the system performance under various parameter configurations using historical monitoring data and simulation environments, and learning a robust initial policy. After the actual deployment of the system, ESOM mainly performs online fine-tuning. This online adjustment is periodic (e.g., evaluating and adjusting decisions every few minutes or hours) or event-driven (e.g., triggered when a continuous performance decline or significant change in user feedback is detected), rather than adjusting all parameters in real-time frame by frame, to balance the timeliness of optimization and the stability of the system. The policy network of ESOM (especially the top-level macro-policy selection part in hierarchical DRL) is designed as a lightweight network structure to ensure that it can quickly (e.g., in milliseconds) output decisions after receiving the current system state. To avoid system performance fluctuations caused by overly frequent or drastic parameter adjustments, the output actions of ESOM will go through a smoothing processing module (e.g., using exponential moving average to update target parameters, or setting maximum step size and change frequency limits for parameter adjustment). In addition, when executing ESOM instructions, the dynamic resource management unit also considers the current load and task queue of the system to avoid instantaneous overload. The dynamic resource management unit (RMU) inside the U-ACE then performs fine-grained resource adaptive regulation according to the instructions output by ESOM. For example, if the instruction requires reducing the computational load of the shared encoder, RMU will switch to a pre-trained CST-Transformer model variant with a shallower number of layers, or for the CST-Transformer architecture that supports the early exit mechanism, activate an earlier exit point to obtain the required features, thereby reducing the computational amount. Similarly, RMU can dynamically activate or deactivate the decoder branches for specific semantic analysis tasks in APEM according to the instructions (e.g., only retaining basic object detection and closing complex behavior analysis branches in low-priority scenarios), and select a predefined neural compression model that matches the current resource budget and content priority for ICM (e.g., switching from a high-bitrate high-quality model to a low-bitrate high-compression model). RMU is also responsible for submitting resource request hints (such as GPU utilization targets, task priorities) to the underlying operating system or computing platform, and optimizing task scheduling and batch size inside U-ACE to indirectly affect and efficiently utilize the allocated computing resources.ESOM continuously monitors the overall operating status of U-ACE (such as end-to-end latency, throughput, GPU / CPU occupancy), the perception accuracy and decision rationality of the PPM module in step S1, the analysis accuracy (mAP) of the APEM module in step S2, prediction error, enhancement quality (MOS score simulation), the compression ratio and reconstruction quality (PSNR / SSIM) of the ICM module in step S3, the user satisfaction simulation score of the presentation effect of the UPSM module in step S4, and real-time feedback from the user side (such as the number of video freezes, explicit scores).

[0035] Based on these high-dimensional state information, the DRL agent of ESOM dynamically adjusts the following parameters and strategies: The scene complexity determination threshold and priority division rule of PPM in step S1.

[0036] The analysis task depth and breadth of APEM in step S2 (for example, which specific analysis subtasks to activate, select different complexity versions of the model), the level of key information enhancement.

[0037] The compression intensity and target bitrate range of ICM in step S3.

[0038] The presentation strategy details of UPSM in step S4 (for example, the number of levels and quality thresholds of hierarchical rendering).

[0039] The computational depth (for example, dynamically skipping certain layers or modules) or the number of activated attention heads of the internal shared encoder (such as CST-Transformer) of U-ACE.

[0040] The resource allocation strategy of the internal dynamic resource management unit (RMU) of U-ACE. The goal of ESOM is to learn an optimal strategy to achieve a dynamic and globally optimal balance among processing efficiency, analysis quality, lightweightness, resource consumption, and user experience.

[0041] Among them, the dynamic sparse attention mechanism of CST-Transformer, for example, can predict a Top-K subset of Keys for each Query, or predict a binary mask to indicate which Keys are relevant, for each Query through a lightweight MLP or a method based on local feature similarity before calculating the attention scores between each Query and all Keys. In this way, the attention calculation is only performed on this sparse subset, reducing the computational amount, which is crucial for processing high-resolution or long-sequence videos and is an effective means to cope with the computational challenges of unified large models.

[0042] For the multi-head prediction decoder and conditional GAN enhancement part of the APEM module, please refer to the appendix Figure 3Among them, the prediction consistency regularization loss in the multi-head prediction decoder of APEM can be designed, for example, to encourage the alignment in time between the end point of the short-term trajectory prediction and the occurrence probability of the mid-term event (such as "an object enters a certain area"). If the short-term trajectory predicts that the object will enter area A at time t, then the probability of "an object enters near area A at time t" in the mid-term event prediction should also increase accordingly. This loss can be a soft constraint based on KL divergence, cross-entropy, or certain logical rules.

[0043] Among them, the learnable context autoregressive entropy model adopted by ICM, its context can include: (1) spatially encoded adjacent latent variables; (2) latent variables corresponding to the previous frame encoded in time; (3) auxiliary information (side information) related to the current encoding block extracted from the shared features of CST-Transformer, such as texture complexity, motion vectors (if available), and semantic class labels output by APEM. These rich context information enables the entropy model to more accurately predict the probability distribution of the current latent variable, thereby achieving higher compression efficiency. The prediction module of APEM may indirectly learn some statistical characteristics of the scene from these rich contexts (especially semantic classes and motion information). Even if these characteristics are initially for compression service, they may have a positive impact on the prediction task and form a synergistic gain.

[0044] The ESOM module uses a hierarchical reinforcement learning agent to interact with U-ACE for self-optimization. For details, please refer to the appendix Figure 4Among them, for the hierarchical DRL of ESOM, the high-level policy network outputs a discrete macro pattern label M (such as "efficiency first", "quality first", "balance"). The low-level policy network is a parameterized function π_low(a_low|s_sub,M), which outputs specific low-level actions a_low (i.e., the parameter adjustment values of each module of U-ACE) under the condition of the given current fine-grained state s_sub and the pattern M selected by the high level. This hierarchical structure decomposes complex optimization problems, with the high level responsible for strategic decision-making and the low level responsible for tactical execution, which helps to improve the stability and efficiency of learning. RMU then follows the instructions in a_low output by ESOM. For example, if the instruction requires reducing the computational depth of CST-Transformer, RMU will modify the computational graph so that a part of the intermediate layer computations are skipped. Among them, in the specific implementation of the collaborative gradient regularization term, the calculation of the gradient grad_shared(L_k) refers to the gradient vector of the task loss L_k with respect to all network parameters (mainly the parameters of CST-Transformer) that participate in the calculation of this task and belong to the shared part. Calculating the cosine similarity of gradients between all pairs of tasks and summing them with weights can obtain a scalar that measures the overall degree of gradient conflict. Adding it to the total loss and minimizing it can encourage collaboration.

[0045] The present invention also discloses a lightweight processing system for cloud-connected network monitoring videos based on U-ACE, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned lightweight processing method for cloud-connected network monitoring videos based on U-ACE.

[0046] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, it implements the above-mentioned lightweight processing method for cloud-connected network monitoring videos based on U-ACE.

[0047] The simulation environment and test data are as follows: Simulation hardware platform: CPU: Intel Xeon Gold 6248R @ 3.00GHz (24 Cores) or a multi-core processor of the same level GPU: NVIDIA Tesla V100 (32GB HBM2) or NVIDIA A100 (40 / 80GB HBM2e) or a high-performance computing card such as NVIDIA GeForce RTX 4090 (24GB GDDR6X) Memory: 256GB DDR4 / DDR5 RAM or higher Storage: 10TB NVMe SSD or a high-speed storage array Simulation software environment: OS: Ubuntu 20.04 LTS / 22.04 LTS Deep learning framework: PyTorch 1.12+ (recommended 2.x) or TensorFlow 2.10+ CUDA Version: 11.6+ Video codec library: FFmpeg 5.0+, optional dedicated neural video codec library Programming language: Python 3.9+ Other libraries: OpenCV, NumPy, Pandas, Matplotlib / Seaborn (for result visualization), RLlib / Stable Baselines3 (for DRL implementation) Test datasets: Public datasets: UA-DETRAC: Contains various complex traffic scenarios for vehicle detection, tracking, and scene complexity analysis.

[0048] MOTChallenge (MOT17, MOT20, MOT23): Contains crowd scenes under different lighting, density, and occlusion conditions for multi-object tracking and scene dynamics analysis.

[0049] VisDrone2019 / UAVDT: Videos taken from the perspective of drones, characterized by small targets, large perspective changes, and complex backgrounds, for testing the model's perception and robustness to small targets.

[0050] ActivityNet / Kinetics-700: Large-scale action recognition datasets, which can be used for pre-training of spatio-temporal feature extractors in PPM and APEM or evaluation of some tasks.

[0051] DAVIS / YouTube-VOS: Video object segmentation datasets, which can be used to evaluate the initial saliency of PPM or the performance of the segmentation branch of APEM.

[0052] Self-built / industry-specific datasets: Video data collected and annotated for specific monitoring application scenarios (such as smart city intersections, industrial park perimeters, inside shopping malls, specific events such as detecting left-behind objects, abnormal crowd behaviors, etc.). The data should cover different weather, lighting, time periods, and include various normal activities and simulated abnormal events. The total amount of the dataset is recommended to be more than several hundred hours, and fine-grained annotation should be carried out, including scene types, complexity levels, bounding boxes and IDs of key objects, trajectories, event types and spatio-temporal ranges, user attention areas (obtained by simulating user behaviors), etc.

[0053] Input Video Specification: The performance comparison experiment is mainly based on the scenario where the input video resolution is 1080p (1920×1080) @ 25 / 30fps. For core computing modules such as CST-Transformer, its computing load is significantly related to the input resolution. The U-ACE proposed in the present invention aims to seek a balance between performance and efficiency under different resolution inputs through dynamic sparse attention, dynamically adjusting the effective computing depth of the encoder, and the priority decision of PPM.

[0054] Performance Comparison: To verify the effectiveness of the method of the present invention, a baseline algorithm is selected for comparison. Baseline Algorithm: A typical enhanced AI monitoring system that uses fixed high-quality H.264 / AVC or H.265 / HEVC encoding (e.g., constant bitrate or constant quality mode) for all input video streams, and then uniformly uses advanced object detection models such as YOLOv5 / YOLOv8 for real-time analysis. This baseline algorithm does not include scene-adaptive dynamic resource allocation, differential processing flow, deep neural compression, user preference learning, and system-level closed-loop feedback optimization mechanism.

[0055] The performance metrics and comparison results are shown in the following figure and presented in the form of a line chart in the accompanying drawings: Indicator / Condition Baseline Algorithm Method of the Present Invention mAP@UA-DETRAC (Sunny Day) 0.72 0.76 mAP@UA-DETRAC (Night) 0.65 0.70 mAP@MOT17 (Indoor) 0.68 0.72 mAP@MOT17 (Outdoor Complex) 0.60 0.65 Average Video Loading Time (ms) 1200 ms 800 ms (Average) User Satisfaction Score (1 - 5) 3.2 4.0 (Average) Result Analysis (elaborated in detail in conjunction with Figure 5 and Figure 6 ) Analysis of Accuracy and Intelligence Level (refer to Figure 5 ) Figure 5The line chart shows the comparison between the method of the present invention (for high-priority content) and the baseline algorithm in terms of the mean average precision (mAP) of key object detection on multiple challenging test subsets (such as the sunny and night subsets of UA-DETRAC, and the indoor and outdoor complex subsets of MOT17). Although the overall system of the present invention pursues light weight and high efficiency, by preferentially allocating computing resources to high-value scenarios through PPM and performing in-depth and refined analysis (including feature enhancement and context awareness) by APEM, the accuracy of core AI tasks (such as object detection) in these key scenarios has been improved or at least not lower than that of the baseline algorithm. Specifically, for the high-priority regions identified by PPM and deeply processed by APEM in the scenario (such as regions containing key moving objects or specific events), the accuracy of key object detection (mAP) is effectively guaranteed or even improved. For example, on the UA-DETRAC night subset, the mAP increased from 0.65 to 0.70, and on the MOT17 outdoor complex subset, it increased from 0.60 to 0.65. This reflects the intelligent strategy of the method of the present invention and the ability to improve the analysis quality through the combination of complex model design and scenarios. This improvement is mainly due to the following synergistic effects: U-ACE focuses computing resources on the refined semantic analysis and adaptive enhancement of key information by APEM for high-priority regions (such as Conditional GAN, which performs denoising, super-resolution, or contrast enhancement on high-priority content in the APEM processing flow), effectively improving the visibility and distinguishability of objects in these regions. PPM ensures that regions containing key information are preferentially sent to APEM for high-quality processing and enhancement, and then this part of high-quality data is retained or encoded in high quality during the final presentation. And ICM is responsible for processing non-high-priority content determined by PPM or non-key regions identified by APEM, achieving a large amount of compression. Here, the mAP improvement is the evaluation result for the high-priority content processed by the APEM module, rather than the evaluation of the non-key regions extremely compressed by ICM. For non-key regions, the information fidelity will be reduced due to deep compression, but its goal is to maximize the compression efficiency.

[0056] User experience: As shown in the data in the performance comparison table, the method of the present invention significantly reduces the average video loading time (for example, from 1200 ms to 800 ms) due to the reduction of the overall bit rate, hierarchical rendering, and user-preference-driven presentation. At the same time, the comprehensive user satisfaction score obtained through user research or collecting implicit user feedback (such as the stuttering rate, task completion efficiency) has also been improved (for example, from 3.2 points to 4.0 points, with a full score of 5 points). This shows that the present invention has also achieved good results in optimizing the user interaction experience. The user satisfaction is comprehensively evaluated in the following ways: Objective performance metrics: Recorded and compared the objective performance data of key interactions, including the average video loading time (the time from when the user initiates viewing to the appearance of the first frame, which is 800ms on average for the method of the present invention and 1200ms for the baseline algorithm), the alarm delay when high-priority events occur, and the system response smoothness when simulating users' target tracking and playback operations (for example, by counting the number of stutters or frame loss rates per unit time, the method of the present invention is significantly lower than the baseline algorithm).

[0057] Subjective evaluation: Conducted a simulated user study by recruiting 100 participants with experience in using surveillance videos. Showed the video clips processed by the method of the present invention and the video clips processed by the baseline algorithm (the clip content, scene complexity, and event types were matched and balanced) to the participants, and asked them to complete a series of typical surveillance tasks (such as finding specific targets, identifying abnormal behaviors, evaluating video clarity, etc.). After completing the tasks, the participants were required to fill out a questionnaire based on a Likert five-point scale to evaluate the two methods from multiple dimensions such as video clarity, loading speed, interaction smoothness, and efficiency of obtaining key information. Based on the comprehensive scores, the average user satisfaction score of the method of the present invention was 4.0 (out of 5), while that of the baseline algorithm was 3.2 points.

[0058] Consideration of hierarchical rendering delay: The hierarchical rendering mechanism of UPSM has considered potential rendering delays in its design. Through the rapid prediction of the user's area of interest (for example, based on historical preferences or lightweight real-time interaction analysis) and the asynchronous loading and efficient caching mechanism for different quality data streams, it is ensured that when the user's focus point switches, the loading and rendering delay of high-quality content is controlled within [for example: within 150ms], which has little impact on the user experience. In the above subjective evaluation, the smoothness of rendering switching is also one of the evaluation dimensions. The present invention optimizes the user experience of interacting with surveillance videos through the user-preference-driven presentation and hierarchical rendering of UPSM, combined with the improvement of the overall processing efficiency.

[0059] Furthermore, in order to verify the contribution of each core module and collaborative mechanism inside the U-ACE of the present invention to the overall performance, ablation experiments were conducted, and the results are as Figure 6 shown. Figure 6 Shows the comparison of the method of the present invention (complete U-ACE), the baseline algorithm (separate model), and several U-ACE ablation variants (for example, U-ACE lacking APEM feedback PPM, U-ACE lacking a shared encoder, and U-ACE lacking ESOM optimization) in terms of comprehensive performance scoring metrics.

[0060] From Figure 6It can be clearly seen that the complete U-ACE method obtains the highest comprehensive performance score, significantly outperforming the baseline algorithm and all ablation variants, which intuitively reflects the "synergistic effect" generated by the unified engine architecture proposed in the present invention and its internal cooperation mechanism. The specific synergistic effects and the resulting performance improvements at least include: 1. By comparing the performance of "complete U-ACE" and "U-ACE (without internal feedback)", it can be seen that when the prediction results of APEM are fed back to PPM to assist its dynamic adjustment of priorities, PPM can make more forward-looking and accurate decisions, thus more effectively allocating computing resources to the content that truly requires in-depth processing. At the same time, the scene perception results output by PPM in step S1 affect APEM's utilization of the shared encoder features through the attention modulation mechanism, making APEM's analysis more targeted. This two-way cooperation between PPM and APEM jointly improves the effectiveness of high-priority content processing and the capture rate of key information, making a significant contribution to the comprehensive performance score.

[0061] 2. By comparing the performance of "complete U-ACE" and "U-ACE (without shared encoder)", it shows that the entire U-ACE architecture is based on the shared encoder features and an end-to-end joint training strategy including a collaborative gradient regularization term, which fundamentally promotes information sharing among modules, reduces redundant calculations, and alleviates conflicts in multi-task learning, providing a solid foundation for the synergistic effect between specific modules.

[0062] 3. By comparing the performance of "complete U-ACE" and "U-ACE (without ESOM optimization)", it shows that the ESOM module can enable U-ACE as a whole to adapt to different input video characteristics, operating environments, and user needs by continuously monitoring based on deep reinforcement learning and dynamically adjusting the operating parameters of each module of U-ACE (the decision threshold of PPM, the analysis depth of APEM, the compression intensity of ICM, the presentation strategy of UPSM) and core components (shared encoder, dynamic resource management unit), achieving a dynamic and better balance among processing efficiency, resource consumption, analysis quality, and user experience. This global closed-loop self-optimization ability is not possessed by separate models or other integrated models lacking such mechanisms and is crucial for improving the long-term comprehensive performance of the system in complex and changing scenarios.

[0063] In summary, through the collaborative work of a highly integrated and functionally complex AI model and in-depth optimization for the monitoring application scenario, the present invention achieves significant improvements in multiple dimensions such as the efficiency, lightweight, intelligent analysis accuracy, and user experience of monitoring video processing, and has important practical application value.

[0064] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A lightweight processing method for cloud-connected monitoring videos based on U-ACE, characterized in that, This method is executed by the multi-task and multi-modal AI core engine U-ACE, which adopts a deep Encoder-Decoder architecture and integrates a Perception and Priority Decision Module (PPM), a Deep Analysis and Prediction Enhancement Module (APEM), an Intelligent Compression and Lightweight Module (ICM), a User Preference and Intelligent Presentation Module (UPSM), and an Engine Self-Optimization Module (ESOM). This method at least includes the following steps: The PPM of the U-ACE uses the early features extracted by the shared encoder of the U-ACE to real-time perceive the scene type, complexity, and change situation of the input monitoring video, and combines with the dynamic resource management unit inside the U-ACE to dynamically determine the processing priority of the video content and preliminarily plan the computing resources; For the video content determined to be of high priority, the APEM of the U-ACE uses the deep spatio-temporal features extracted by the shared encoder to perform refined semantic analysis, adaptive enhancement of key information, and forward prediction of the dynamic change pattern of the scene; among them, the scene perception result output by the PPM affects the utilization of the shared encoder features by the APEM through the attention modulation mechanism; the prediction result of the APEM is internally fed back to the PPM to assist in the dynamic adjustment of its priority decision and guide the enhancement parameters of the APEM; For the content determined to be of non-high priority or the non-key regions in the video frames identified by the APEM, the ICM of the U-ACE uses the features extracted by the shared encoder or the output of the PPM to perform deep compression using a content-adaptive neural compression algorithm, and its compression strategy is jointly modulated by the determined priority and the analysis result of the APEM; The UPSM of the U-ACE learns the user's historical interaction data to build a user preference model, and combines the processed high-priority content and the compressed data to perform customized enhanced display and refined hierarchical rendering of the content concerned by the user; The ESOM of the U-ACE, based on deep reinforcement learning, continuously monitors the overall operating state of the U-ACE, the performance metrics of each module of the PPM, APEM, ICM, and UPSM, as well as the user feedback, and dynamically adjusts the decision threshold of the PPM, the analysis depth and enhancement level of the APEM, the compression intensity of the ICM, the presentation strategy of the UPSM, and the operating parameters of the shared encoder and the dynamic resource management unit inside the U-ACE; among them, the U-ACE realizes the efficient cooperation and information sharing of the functions of each module through the shared encoder features, the attention modulation mechanism, and an end-to-end joint training strategy including a collaborative gradient regularization term; 2. The method according to claim 1, wherein The dynamic change pattern prediction part of the APEM adopts a multi-head prediction decoder, which is based on the features encoded by CST-Transformer, parallelly predicts various types of future information, and ensures the cooperation and rationality between different prediction heads through an internal prediction consistency regularization loss.

3. The method according to claim 1, wherein The neural compression algorithm of the ICM adopts a learnable context-based autoregressive entropy model and combines the semantic segmentation map or saliency map output by the APEM to allocate fewer bits to the latent variable representation of non-critical regions, achieving semantically guided efficient compression, so that the prediction module of the APEM indirectly benefits from the image compressibility features learned by the ICM, generating a synergistic gain.

4. The method according to claim 1, characterized in that, The user preference model of the UPSM is an interpretable enhanced graph attention network.

5. The method according to claim 1, characterized in that, The action space of the deep reinforcement learning agent of the ESOM is designed as a hierarchical structure. The top-level action selects the macro policy, and the bottom-level action fine-tunes the specific parameters of each module of the U-ACE under the selected macro policy.

6. The method according to claim 1, wherein The collaborative gradient regularization term is: L_sgr = λ_sgr * Σ_ {i≠j} (1 - cos(grad(L_i), grad(L_j))), Where L_i and L_j are the gradients of the loss functions of different task modules inside the U-ACE on the shared parameters respectively, cos(·,·) represents the cosine similarity, and λ_sgr is the regularization coefficient. This regularization term promotes the synergy of parameter updates by punishing the significant inconsistency in the gradient directions of different tasks.

7. The method according to claim 6, characterized in that, The dynamic resource management unit inside the U-ACE can achieve fine-grained resource adaptability inside the engine based on the priority determination of the PPM and the optimization instructions of the ESOM, specifically including: adjusting the effective computational depth by selecting different variants of the shared encoder model with preset different computational depths or using a dynamic network structure that supports early exit, dynamically activating or deactivating different decoder branches for specific tasks in the APEM, and selecting the neural compression models with different complexity levels preset in the ICM.

8. The method according to claim 1, characterized in that, The attention modulation mechanism includes: the scene category embedding vector e_sc and complexity score c output by the PPM are multiplied or concatenated with the feature map F_ape_in obtained by the APEM from the shared encoder channel by channel and then processed through a small convolutional network to generate the modulated feature map F_ape_mod = AttentionModulator(F_ape_in, e_sc, c), where the AttentionModulator network learns to dynamically adjust the channel response or spatial attention region of F_ape_in according to the scene information, enabling the APEM to more effectively utilize the features most relevant to the current scene.

9. A lightweight processing system for cloud-connected monitoring videos based on U-ACE, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the U-ACE-based lightweight processing method for cloud-connected monitoring videos according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the U-ACE-based lightweight processing method for cloud-connected monitoring videos according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • DeepStream-based monitoring video analysis method and system

    CN116824480A

  • Man-machine cooperation intelligent control system based on AIGC

    CN119940425A

  • Intelligent agent architecture based on multi-modal large model

    CN120046645A

  • Ai-enhanced simulation and modeling experimentation and control

    US20240348663A1

  • Multimodal heterogeneous feature fusion-based compact video event description method

    WO2023050295A1

Cited By

  • Underwater image intelligent real-time enhancement method and device and electronic equipment

    CN122024031A