Deep synthesis service instantaneous algorithm

By improving the deep synthesis service algorithm, and optimizing the deep synthesis process using lightweight networks and dynamic resource allocation, the problem of slow processing speed of existing algorithms in scenarios with high real-time requirements is solved, and a low-latency, high-efficiency deep synthesis service is achieved.

CN120953450APending Publication Date: 2025-11-14ZHIXIN LEADER (HANGZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511095554.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing deep synthesis service algorithms are slow in real-time applications, making it difficult to meet user experience requirements.

Method used

We employ an improved neural network model, a lightweight network structure, a dynamic computing resource allocation mechanism, and a multi-dimensional quality assessment mechanism, combined with adaptive noise removal, format conversion, and feature extraction, to optimize the deep synthesis algorithm.

Benefits of technology

It achieves low-latency processing of deep synthesis service instantaneous algorithms in real-time interactive scenarios, improves the utilization of computing resources, enhances the visual effects and scalability of synthesized content, and reduces service deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953450A_ABST
    Figure CN120953450A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep synthesis, and discloses a deep synthesis service instantaneous algorithm, comprising the following steps: preprocessing input data, including data cleaning, format conversion and feature extraction; inputting the preprocessed data into an improved neural network model for feature learning and synthesis calculation, wherein the improved neural network model is optimized by introducing an attention mechanism and a lightweight network structure; a dynamic computing resource allocation mechanism is adopted, and corresponding computing resources are allocated to the deep synthesis tasks according to real-time computing loads and task priorities; and outputting a synthesis result. Through data preprocessing optimization, lightweight model design and dynamic resource allocation, the average delay of the algorithm for processing a frame of 256 * 256 pixel image is reduced to be within 80ms, the requirement of a real-time interaction scene is met, the delay is reduced by more than 60% compared with an existing algorithm, and the real-time performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep synthesis technology, specifically to a deep synthesis service instantaneous algorithm. Background Technology

[0002] With the rapid development of artificial intelligence technology, deep synthesis technology has been widely used in fields such as images, videos, and audio. By utilizing deep learning models, deep synthesis technology can generate highly realistic synthetic content, bringing revolutionary changes to industries such as film and television production, virtual reality, and advertising creativity.

[0003] However, existing deep synthesis service algorithms typically suffer from slow processing speed and poor real-time performance. This is mainly because deep synthesis tasks require substantial computational resources, and the data preprocessing and model calculation processes are complex and time-consuming. In some application scenarios with high real-time requirements, such as virtual avatar generation in real-time video conferencing and real-time character synthesis in online interactive games, existing algorithms struggle to meet the demands, severely impacting user experience. Therefore, we propose an instantaneous deep synthesis service algorithm. Summary of the Invention

[0004] The purpose of this invention is to provide a deep synthesis service instantaneous algorithm to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a deep synthesis service instantaneous algorithm, comprising the following steps:

[0006] S1. Preprocess the input data, including data cleaning, format conversion and feature extraction;

[0007] S2. The preprocessed data is input into the improved neural network model for feature learning and synthesis computation. The improved neural network model is optimized by introducing an attention mechanism and a lightweight network structure.

[0008] S3. A dynamic computing resource allocation mechanism is adopted to allocate corresponding computing resources to deep synthesis tasks based on real-time computing load and task priority.

[0009] S4. Output the synthesized result.

[0010] Optionally, the data cleaning includes removing noisy data, filling in missing values, and correcting outliers. The removal of noisy data employs an adaptive threshold filtering algorithm, which dynamically adjusts the filtering threshold according to the noise distribution of the data to improve the accuracy of noise removal.

[0011] Optionally, the lightweight network structure is achieved by reducing the number of network layers, reducing the number of convolutional kernels, and using grouped convolutions. The number of groups in the grouped convolutions can be dynamically adjusted according to the feature dimensions of the input data, thereby reducing computational redundancy while ensuring the integrity of feature extraction.

[0012] Optionally, the dynamic computing resource allocation mechanism includes real-time monitoring of the CPU utilization, memory usage, and network bandwidth of computing nodes, and adjusting the allocation ratio of computing resources based on the monitoring results;

[0013] When the load on a single computing node exceeds a preset threshold, some tasks are automatically offloaded to backup computing nodes to achieve load balancing.

[0014] Optionally, the improved neural network model employs a mixed-precision training method during training to accelerate the training speed;

[0015] The model uses 16-bit floating-point numbers for calculations during forward propagation and switches to 32-bit floating-point numbers during backpropagation to update parameters. At the same time, a gradient clipping threshold is set to prevent gradient explosion.

[0016] Optionally, the feature extraction employs principal component analysis or independent component analysis. In the principal component analysis process, feature contribution weights are introduced, and principal components with contribution values ​​below a preset threshold are removed to compress the data dimensionality.

[0017] Optionally, before outputting the synthesis result, a step of quality assessment of the synthesis result is included. The quality assessment uses peak signal-to-noise ratio and structural similarity index. When the assessment result is lower than the preset standard, an algorithm parameter fine-tuning mechanism is automatically triggered to re-perform the synthesis calculation.

[0018] Optionally, the attention mechanism in the improved neural network model adopts a spatiotemporal joint attention module, which can simultaneously assign weights to the spatial features and time series features of the input data, thereby improving the model's accuracy in synthesizing dynamic sequence data.

[0019] Compared with existing technologies, this invention provides a deep synthesis service instantaneous algorithm, which has the following beneficial effects:

[0020] 1. This deep synthesis service instantaneous algorithm, through data preprocessing optimization, lightweight model design, and dynamic resource allocation, reduces the average latency of processing a 256×256 pixel image frame to less than 80ms, meeting the requirements of real-time interactive scenarios. Compared with existing algorithms, it reduces latency by more than 60%, thus improving real-time performance. Employing a multi-dimensional quality assessment and adaptive adjustment mechanism, the average PSNR of the synthesized content reaches over 34dB, SSIM reaches over 0.93, and FID is less than 10, resulting in realistic visual effects that are difficult to distinguish from real content with the naked eye.

[0021] 2. This deep synthesis service instantaneous algorithm, with its dynamic computing resource allocation mechanism, improves computing resource utilization by more than 40%, avoids resource waste, and reduces service deployment costs. It performs particularly well in multi-task concurrent scenarios. The algorithm can dynamically adjust parameters according to the needs of different application scenarios (such as real-time priority or quality priority), and supports synthesis tasks from low resolution (128×128) to high resolution (1024×1024), with good scalability and flexibility.

[0022] 3. This deep synthesis service instant algorithm is implemented based on the PyTorch framework and supports various deployment environments such as CPU, GPU, and edge computing devices. It can be deployed via Docker containerization, making it easy to quickly integrate into existing service systems and lowering the application threshold. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the process structure of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] This invention provides a technical solution: a deep synthesis service instantaneous algorithm, comprising the following steps:

[0026] Adaptive noise removal employs an adaptive threshold filtering algorithm for different types of noise (such as Gaussian noise, salt-and-pepper noise, and impulse noise). For Gaussian noise, the standard deviation σ of the local image region is calculated. When the difference between the pixel value and the region mean exceeds 3σ, it is identified as noise and smoothed. For salt-and-pepper noise, median filtering is used. The filtering window size is dynamically adjusted according to the noise density: a 3×3 window is used when the noise density is below 10%, and a 5×5 window is used when it is above 10%. The filtering result is calculated, where n is the window radius, x is the input image, and y is the filtered image.

[0027] High-efficiency format conversion employs a hardware-accelerated method, leveraging the parallel computing capabilities of the GPU to convert image data from RGB format to the tensor format required by the model. Parallel computation for pixel value normalization is implemented through CUDA kernel functions, changing the traditional pixel-by-pixel processing method to block-level processing, resulting in an 8-10x speedup.

[0028] Lightweight feature extraction combines principal component analysis (PCA) with an incremental learning strategy to construct a basic feature space during the initial training phase and update the feature space through incremental PCA during real-time processing, avoiding the need to recalculate the covariance matrix and eigenvectors for each processing step. Neural network model computation, the core of deep synthesis, is achieved through structural innovation and training optimization, significantly improving computational speed while maintaining synthesis quality.

[0029] The lightweight Transformer architecture introduces sparse attention and depthwise separable convolutions on top of the traditional Transformer model. The sparse attention mechanism only calculates attention weights at key locations in the input sequence. For example, in facial image synthesis, it only focuses on key regions such as eyes, mouth, and nose, reducing attention computation by more than 70%. The depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution.

[0030] Mixed precision training and inference: During the model training phase, mixed precision of FP16 and FP32 is used. FP16 is used to store activation values ​​and gradients, while FP32 is used to update weights. At the same time, loss scaling technology is introduced to avoid gradient underflow. During the inference phase, FP16 precision is used throughout, and TensorRT is used for model optimization. Through layer fusion, quantization-aware training and other methods, the model inference speed is improved by 2-3 times.

[0031] Dynamic model pruning adjusts the model structure based on the complexity of the input data. For simple data (such as static background images), 30%-50% of the network layers are pruned; for complex data (such as dynamic facial expressions), the complete network structure is retained. Pruning is based on the importance score of each layer, which is determined by calculating the gradient contribution of the layer output to the loss function. Only network layers with scores higher than a threshold are retained.

[0032] To ensure the real-time performance of the algorithm in multi-task concurrent scenarios, this invention designs a load-aware dynamic computing resource allocation mechanism: real-time load monitoring, which collects indicators such as CPU utilization, GPU memory usage, memory bandwidth, and network I / O of computing nodes to build a load assessment model; and task priority scheduling, which sets priorities according to the real-time requirements of tasks and user levels, with priorities divided into 5 levels (levels 1-5), and tasks with higher priorities are given priority to obtain computing resources.

[0033] Elastic resource scaling automatically expands compute nodes using cloud-native technologies when current computing resources are found to be insufficient to meet the real-time requirements of all tasks. It utilizes Kubernetes' HorizontalPodAutoscaler to dynamically adjust the number of Pod replicas, with scaling response time controlled within 5 seconds.

[0034] For output and quality assessment, this invention establishes a closed-loop quality control mechanism to ensure the quality of the synthesized content. Multi-dimensional quality assessment is implemented, introducing perceptual hash similarity (PHash) and Fraser Inception distance (FID) in addition to traditional peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) metrics. PHash evaluates the visual consistency of the synthesized content by calculating the hash value of the image and comparing the Hamming distance; FID evaluates the realism of the synthesized content by comparing the distribution differences between real and synthesized data in the InceptionV3 model feature space.

[0035] Adaptive quality adjustment automatically triggers a parameter adjustment mechanism when the quality assessment result falls below a preset threshold. This includes measures such as increasing the sampling temperature during model inference, increasing the number of generation iterations, and adjusting the attention weight distribution until the synthesis quality meets the requirements. For scenarios with extremely high real-time requirements, a dynamic balance strategy between quality and speed can be implemented, sacrificing some quality while ensuring latency, or appropriately increasing latency in scenarios where quality is prioritized.

[0036] As one application of this embodiment:

[0037] In video conferencing scenarios, after adopting the algorithm of this invention, the motion latency of the virtual avatar is reduced from 300ms to 70ms, and users can hardly feel the delay, greatly improving the naturalness of the interaction; in live streaming scenarios, the transformation of the anchor's cartoon avatar can be completed in real time, and the audience experience is significantly improved; in intelligent customer service scenarios, the response speed of the virtual human is shortened from 1.5 seconds to 0.6 seconds, and the service efficiency is improved by 150%.

[0038] The present invention has been described in detail above. However, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, any modifications or improvements that do not depart from the spirit of the present invention are within the scope of protection of the present invention.

Claims

1. A deep synthesis service instantaneous algorithm, characterized in that, Includes the following steps: S1. Preprocess the input data, including data cleaning, format conversion and feature extraction; S2. The preprocessed data is input into the improved neural network model for feature learning and synthesis computation. The improved neural network model is optimized by introducing an attention mechanism and a lightweight network structure. S3. A dynamic computing resource allocation mechanism is adopted to allocate corresponding computing resources to deep synthesis tasks based on real-time computing load and task priority. S4. Output the synthesized result.

2. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The data cleaning includes removing noisy data, filling in missing values, and correcting outliers. The removal of noisy data uses an adaptive threshold filtering algorithm, which dynamically adjusts the filtering threshold according to the noise distribution of the data to improve the accuracy of noise removal.

3. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The lightweight network structure is achieved by reducing the number of network layers, reducing the number of convolutional kernels, and using grouped convolutions. The number of groups in the grouped convolutions can be dynamically adjusted according to the feature dimensions of the input data, thereby reducing computational redundancy while ensuring the integrity of feature extraction.

4. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The dynamic computing resource allocation mechanism includes real-time monitoring of CPU utilization, memory usage, and network bandwidth of computing nodes, and adjusting the allocation ratio of computing resources based on the monitoring results; when the load of a single computing node exceeds a preset threshold, some tasks are automatically diverted to backup computing nodes to achieve load balancing.

5. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The improved neural network model employs a mixed-precision training method during training to accelerate the training speed; it uses 16-bit floating-point numbers for calculations during forward propagation and switches to 32-bit floating-point numbers when updating parameters during backpropagation, while setting a gradient clipping threshold to prevent gradient explosion.

6. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The feature extraction employs principal component analysis or independent component analysis. In the principal component analysis process, feature contribution weights are introduced, and principal components with contribution values ​​below a preset threshold are removed to compress data dimensions.

7. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, Before outputting the synthesis result, a quality assessment step is also included. The quality assessment uses peak signal-to-noise ratio and structural similarity index. When the assessment result is lower than the preset standard, the algorithm parameter fine-tuning mechanism is automatically triggered, and the synthesis calculation is re-performed.

8. The instantaneous algorithm for deep synthesis services according to claim 1, characterized in that, The improved neural network model employs a spatiotemporal joint attention module, which can simultaneously assign weights to the spatial and temporal features of the input data, thereby improving the model's accuracy in synthesizing dynamic sequence data.

Citation Information

Cited By

  • Time series data monitoring method and device, electronic equipment and nonvolatile storage medium

    CN116821661A

  • Time series data monitoring method and device, electronic equipment and nonvolatile storage medium

    CN116821661B