Smart card communication method, system and device for music stage performance video and medium
Through the intelligent cartoon model, combined with dynamic lighting adaptation and lightweight U-Net branches, the real-time processing and high-quality cartoonization of music stage videos are solved, efficient visual effects and complex lighting adaptability are achieved, and the artistic expression of music stage videos is enhanced.
Patent Information
- Application Number
- CN202510436972.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-04
AI Technical Summary
The existing cartoonization technology has high computational complexity when processing music stage videos, which is difficult to meet the needs of real-time processing. It also has shortcomings in color style adaptation and subject-background separation accuracy, which is difficult to meet the needs of high-quality video content creation.
The intelligent cartoonization model is adopted, including an encoder, a spatiotemporal attention processor and a decoder. Through optical flow-feature fusion, lightweight U-Net branching and sub-region processing module, combined with dynamic lighting adaptive normalization and dynamic LUT module, it realizes rapid processing of video and high-quality cartoonization.
It significantly improves the visual quality of music stage videos, has strong adaptability to complex lighting conditions, realizes real-time and stable cartoon processing, reduces calculation complexity, and ensures operational efficiency.
Smart Images

Figure CN120259475A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video processing, and particularly relates to an intelligent cartoonization method, system, device and medium for music stage performance videos. Background Art
[0002] In modern video content creation, cartoonization technology is widely used to convert real videos into cartoon styles to enhance visual attraction and artistic expressiveness. As a dynamic and expressive content form, music stage performance videos are particularly suitable for cartoonization to improve the viewing experience of the audience. However, existing cartoonization technologies rely on a large amount of training data when processing music stage videos, with high model computational complexity, making it difficult to meet the real-time processing requirements. Moreover, there are obvious deficiencies in aspects such as color style adaptation and the accuracy of foreground-background separation, making it difficult to meet the needs of high-quality video content creation. Summary of the Invention
[0003] The purpose of the present invention is to provide an intelligent cartoonization method, system, device and medium for music stage performance videos to solve the problems existing in the above-mentioned prior art.
[0004] To achieve the above purpose, the present invention provides an intelligent cartoonization method for music stage performance videos, including: Obtaining a video to be processed; Preprocessing the video to be processed to obtain a video sequence, an optical flow field and an illumination intensity distribution map corresponding to the video to be processed; Inputting the video sequence, the optical flow field and the illumination intensity distribution map into an intelligent cartoonization model for cartoonization processing to obtain a cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor and a decoder connected in sequence, the spatio-temporal attention processor includes an optical flow-feature fusion module, a lightweight U-Net branch and a sub-region processing module connected in sequence, and the sub-region processing module includes a foreground path and a background path arranged in parallel.
[0005] Optionally, the process of obtaining the video sequence specifically includes: Performing frame division on the video to be processed based on a preset frame rate, and eliminating camera jitter based on a time alignment algorithm to ensure smooth temporal alignment of video frames, thus completing the acquisition of the video sequence.
[0006] Optionally, the process of obtaining the optical flow field specifically includes: Extracting key frames from the video sequence based on a preset sampling interval; Calculating the sparse optical flow between key frames based on the Lucas-Kanade algorithm, extracting the motion vector field, and expanding and interpolating the motion vector field to non-key frames to generate a complete optical flow field.
[0007] Optionally, the process of obtaining the light intensity distribution map specifically includes: Convert each frame in the video sequence from an RGB image to an HSV color image; Extract the brightness component of the HSV color image and construct a light intensity distribution map.
[0008] Optionally, the training process of the intelligent cartoonization model specifically includes: Obtain training data, where the training data includes training video data, corresponding preprocessed data, and cartoonization results; Construct an initial intelligent cartoonization model, input the training data into the initial intelligent cartoonization model for initialization, and train based on the target loss function to obtain a trained intelligent cartoonization model.
[0009] Optionally, the processing process of the intelligent cartoonization model specifically includes: Input the video sequence into the encoder, extract low-level features through grouped inverse residual blocks, insert a DIA-Norm module after each group of convolutions, dynamically adjust the mean and variance of normalization according to the light intensity map of the current frame, and complete the acquisition of the low-level feature map; input the extracted low-level feature map into a 3D inverse residual block to process consecutive frames, capture spatio-temporal context information through spatio-temporal convolution, and obtain a high-level feature map; Input the high-level feature map and the optical flow field into a spatio-temporal attention processor, fuse the optical flow field and high-level features, generate a spatio-temporal weight matrix, and calculate channel-spatio-temporal attention weights through lightweight 3D convolution to suppress background noise and output weighted features; Input the low-level features into a lightweight U-Net branch to output a binary mask, and obtain a main body region mask and a background region mask; Input the main body region mask and the weighted main body features into the main body path, enhance the feature details through multi-scale residual blocks, retain clothing textures and facial expressions, and output enhanced main body features; Input the background region mask and the weighted background features into the background path, reduce the feature resolution through adaptive average pooling, and combine edge-preserving filtering to generate a cartoonized abstract background, and output abstracted background features; Input the enhanced main body features and the abstracted background features into the decoder, splice them after bilinear upsampling alignment to obtain spliced features; Input the spliced features into the dynamic LUT module, select a pre-trained style according to the HSV histogram, calculate SE attention for each of the RGB three channels, suppress overexposed regions, complete color enhancement, and output a cartoonization result.
[0010] An intelligent cartoonization system for music stage performance videos, comprising: A data acquisition module, configured to obtain a video to be processed; preprocess the video to be processed to obtain a video sequence, an optical flow field, and a light intensity distribution map corresponding to the video to be processed; A cartoonization processing module, configured to input the video sequence, the optical flow field, and the light intensity distribution map into an intelligent cartoonization model for cartoonization processing to obtain a cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor, and a decoder connected in sequence, the spatio-temporal attention processor includes an optical flow-feature fusion module, a lightweight U-Net branch, and a sub-region processing module connected in sequence, and the sub-region processing module includes a main path and a background path arranged in parallel.
[0011] An electronic device, comprising a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the intelligent cartoonization method for music stage performance videos described above.
[0012] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent cartoonization method for music stage performance videos described above is implemented.
[0013] The technical effects of the present invention are as follows: Through the spatio-temporal-optical flow collaborative attention and dynamic illumination adaptive normalization technologies, the present invention effectively solves the problems of fast motion blur and inconsistent styles in music stage videos, and can significantly improve the visual quality; the dynamic LUT module and the illumination intensity analysis mechanism endow the model with strong adaptability to complex illumination conditions, ensuring natural presentation under various scenarios. Combining the optimized design of the lightweight U-Net branch and the 3D convolution architecture, while reducing the computational complexity, the operation efficiency is guaranteed, forming a resource-efficient solution. In addition, through the hierarchical routing and dynamic parameter interaction mechanism, the system realizes real-time and stable processing of complex scenarios such as dynamic illumination and fast actions, providing a new path with both artistic expressiveness and technical feasibility for the cartoonization creation of music stage performance videos. Description of the Drawings
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings: Figure 1 is the implementation flowchart in the embodiment of the present invention; Figure 2 is the schematic diagram of the model structure in the embodiment of the present invention. Detailed implementation manners
[0016] Now, various exemplary implementation manners of the present invention will be described in detail. This detailed description should not be considered as a limitation to the present invention, but rather as a more detailed description of certain aspects, features, and implementation schemes of the present invention.
[0017] It should be understood that the terms described in the present invention are only for describing specific implementation manners and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded from the range.
[0018] Without departing from the scope or spirit of the present invention, various improvements and changes can be made to the specific implementation manners of the specification of the present invention, which are obvious to those skilled in the art. Other implementation manners obtained from the specification of the present invention are obvious to those skilled in the art. The specification and embodiments of this application are only exemplary.
[0019] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.
[0020] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.
[0021] Such as Figure 1 - Figure 2As shown in the figure, in this embodiment, an intelligent cartoonization method for music stage performance videos is provided, including: obtaining a video to be processed; preprocessing the video to be processed to obtain a video sequence, an optical flow field, and a light intensity distribution map corresponding to the video to be processed; inputting the video sequence, the optical flow field, and the light intensity distribution map into an intelligent cartoonization model for cartoonization processing to obtain a cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor, and a decoder connected in sequence, the spatio-temporal attention processor includes an optical flow-feature fusion module, a lightweight U-Net branch, and a sub-region processing module connected in sequence, and the sub-region processing module includes a main body path and a background path arranged in parallel. This method effectively solves the problems of fast motion blur and inconsistent styles, improves the visual quality, has strong adaptability to complex lighting conditions, and ensures natural presentation in various scenarios. At the same time, the optimized design reduces the computational complexity and guarantees the operation efficiency, providing a new solution for the cartoonization creation of music stage performance videos.
[0022] This embodiment proposes an intelligent cartoonization method for music stage videos, the core of which is to combine the spatio-temporal attention mechanism with the dynamic lighting adaptation technology. First, preprocess the input video, eliminate camera jitter through frame alignment, extract inter-frame motion information based on sparse optical flow estimation, and analyze the light intensity distribution of each frame from the HSV color space. Subsequently, design an improved encoder-decoder architecture. The encoder uses grouped inverted residual blocks to extract low-level features and introduces a dynamic lighting adaptation normalization (DIA-Norm) module to dynamically adjust the feature distribution according to the real-time light intensity to retain details in low light or strong light. In the spatio-temporal attention processor, fuse the optical flow field with high-level features, generate a spatio-temporal weight matrix through lightweight 3D convolution, suppress background noise, and enhance the coherence of the motion area.
[0023] Furthermore, generate a performer main body mask through a lightweight U-Net branch to achieve sub-region processing of the main body and the background: the main body path uses multi-scale residual blocks and dilated convolutions to strengthen details and ensure clear clothing textures and facial expressions; the background path generates an abstract cartoon background through adaptive pooling and edge-preserving filtering. In the decoding stage, introduce a dynamic color lookup table (LUT) technology to adaptively select a preset style according to the hue distribution of the input video, and calibrate the color in combination with the channel attention mechanism to avoid overexposure or color cast. Finally, the model deployment can be optimized through TensorRT, supporting real-time processing of 25fps for 1080p videos, and allowing users to adjust the style intensity through an interactive interface to achieve a smooth transition from a realistic to a cartoon style.
[0024] The specific implementation process of this embodiment includes: Video preprocessing: The video to be processed is framed at a preset frame rate (e.g., 30fps), and a time alignment algorithm is used to eliminate camera jitter, ensuring smooth temporal alignment of video frames. The Lucas-Kanade algorithm is used to calculate the sparse optical flow between key frames, extract the motion vector field, and interpolate the optical flow results to non-key frames to generate a complete optical flow field. Each frame is converted from the RGB color space to the HSV color space, the luminance component (V channel) is extracted, and an illumination intensity distribution map is constructed.
[0025] Intelligent cartoonization model processing: Encoder: Extract low-level features through grouped inverted residual blocks, and insert a dynamic illumination adaptive normalization (DIA-Norm) module after each group of convolutions to dynamically adjust the mean and variance of normalization according to the illumination intensity map. Use 3D inverted residual blocks to process three consecutive frames, capture action coherence through spatio-temporal convolutions, and output a high-level feature map containing spatio-temporal context information.
[0026] Optionally, a feature calibration layer is added before the encoder, bilinear interpolation is used to align the optical flow field resolution, and timestamp matching is used to ensure strict synchronization of the optical flow field with the current frame features; Spatio-temporal attention processor: Fuse the optical flow vectors with the high-level features to generate a spatio-temporal weight matrix. Calculate the channel-spatio-temporal attention weights through lightweight 3D convolutions (kernel size 3×3×3) to suppress background noise.
[0027] Foreground-background separation: Run lightweight U-Net branches in parallel, input low-level features to generate a binary mask of the performer, retain high-resolution features (512×512) for the foreground region, and downsample the background region to 256×256. Among them, an optical flow propagation mechanism is introduced in the U-Net branch, and the mask of the previous frame is deformed through the current optical flow field as a prior constraint for the current frame segmentation to ensure temporal continuity of the foreground region.
[0028] Region-based processing: The foreground path includes multi-scale residual blocks and dilated convolution layers; in the foreground path, multi-scale residual blocks are used to enhance the foreground region, retaining clothing textures and facial expressions; The background path includes a pooling module and a filtering module, and an adaptive mean pooling and edge-preserving filtering are performed on the background region to generate a cartoonized abstract background.
[0029] Color enhancement decoder: Fuse the foreground and background features, input them to the decoder, insert a dynamic LUT module after the deconvolution layer, select the best style LUT from the pre-trained color library according to the hue distribution of the input frame, enhance the saturation and contrast, and adjust the RGB channel weights through channel attention to avoid overexposure or color bias.
[0030] Post-processing and real-time output: Apply the CycleGAN (Cycle-consistent Generative Adversarial Network) to optimize the temporal consistency, ensuring smooth transitions of colors and contours between adjacent frames. Meanwhile, add the temporal coherence loss guided by optical flow to explicitly constrain the temporal smoothness of adjacent frames.
[0031] Optimize the model using TensorRT and deploy it to a GPU server to support real-time processing of 1080p videos at 25fps.
[0032] Provide a user interaction layer to dynamically adjust the LUT weights through a style intensity slider, supporting the gradient effect from "realistic" to "cartoon".
[0033] In summary, through the spatio-temporal - optical flow collaborative attention and dynamic illumination adaptive normalization techniques, this embodiment effectively solves the problems of fast motion blur and style inconsistency in music stage videos, and can significantly improve the visual quality; the dynamic LUT module and the illumination intensity analysis mechanism endow the model with strong adaptability to complex illumination conditions, ensuring natural rendering in various scenarios. Combining the optimized design of the lightweight U-Net branch and the 3D convolution architecture, while reducing the computational complexity, it ensures the operation efficiency, forming a resource-efficient solution. In addition, through the hierarchical routing and dynamic parameter interaction mechanism, the system realizes real-time and stable processing of complex scenarios such as dynamic illumination and fast actions, providing a new path with both artistic expressiveness and technical feasibility for the cartoonization creation of music stage performance videos.
[0034] Implementable, this embodiment also provides an intelligent cartoonization system for music stage performance videos, including: A data acquisition module, used to obtain the video to be processed; preprocess the video to be processed to obtain the video sequence, optical flow field, and illumination intensity distribution map corresponding to the video to be processed; A cartoonization processing module, used to input the video sequence, optical flow field, and illumination intensity distribution map into the intelligent cartoonization model for cartoonization processing to obtain the cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor, and a decoder connected in sequence, the spatio-temporal attention processor includes an optical flow - feature fusion module, a lightweight U-Net branch, and a sub-region processing module connected in sequence, and the sub-region processing module includes a main path and a background path arranged in parallel.
[0035] Implementable, this embodiment also provides an electronic device, including a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the intelligent cartoonization method for music stage performance videos described above.
[0036] Implementable, this embodiment also provides a computer-readable storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the intelligent cartoonization method of a music stage performance video described above.
[0037] As described above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An intelligent cartoonization method for music stage performance videos, characterized in that, Including: Obtain the video to be processed; Preprocess the video to be processed to obtain a video sequence, an optical flow field, and a light intensity distribution map corresponding to the video to be processed; Input the video sequence, the optical flow field, and the light intensity distribution map into an intelligent cartoonization model for cartoonization processing to obtain a cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor, and a decoder connected in sequence, and the spatio-temporal attention processor includes an optical flow-feature fusion module, a lightweight U-Net branch, and a sub-region processing module connected in sequence, and the sub-region processing module includes a main body path and a background path arranged in parallel.
2. The intelligent cartooning method for a music stage performance video according to claim 1, characterized in that The process of obtaining the video sequence specifically includes: Perform frame splitting on the video to be processed based on a preset frame rate, and eliminate camera jitter based on a time alignment algorithm to ensure smooth alignment of video frames in time, and complete the acquisition of the video sequence.
3. The intelligent cartooning method for a music stage performance video according to claim 1, characterized in that, The process of obtaining the optical flow field specifically includes: Extract key frames from the video sequence based on a preset sampling interval; Calculate the sparse optical flow between key frames based on the Lucas-Kanade algorithm, extract the motion vector field, and expand and interpolate the motion vector field to non-key frames to generate a complete optical flow field.
4. The intelligent cartoonization method of a music stage performance video according to claim 1, characterized in that The process of obtaining the light intensity distribution map specifically includes: Convert each frame in the video sequence from an RGB image to an HSV color image; Extract the brightness component of the HSV color image to construct a light intensity distribution map.
5. The intelligent cartoonization method of a music stage performance video according to claim 1, characterized in that The training process of the intelligent cartoonization model specifically includes: Obtain training data, where the training data includes training video data, corresponding preprocessing data, and cartoonization results; Construct an initial intelligent cartoonization model, input the training data into the initial intelligent cartoonization model for initialization, and train based on a target loss function to obtain a trained intelligent cartoonization model.
6. The intelligent cartoonization method of a music stage performance video according to claim 1, characterized in that The processing process of the intelligent cartoonization model specifically includes: Input the video sequence into the encoder, extract low-level features through grouped inverted residual blocks, insert a DIA-Norm module after each group of convolutions, and dynamically adjust the mean and variance of normalization according to the light intensity map of the current frame to complete the acquisition of the low-level feature map; input the extracted low-level feature map into a 3D inverted residual block to process consecutive frames, and capture spatio-temporal context information through spatio-temporal convolution to obtain a high-level feature map; Input the high-level feature map and the optical flow field into the spatio-temporal attention processor, fuse the optical flow field and the high-level features to generate a spatio-temporal weight matrix, and calculate channel-spatio-temporal attention weights through lightweight 3D convolution to suppress background noise and output weighted features; Input the low-level features into the lightweight U-Net branch to output a binary mask to obtain a main body region mask and a background region mask; Input the main body region mask and the weighted main body features into the main body path, enhance the feature details through a multi-scale residual block, retain the clothing texture and facial expressions, and output enhanced main body features; Input the background region mask and weighted background features into the background path, reduce the feature resolution through adaptive average pooling, and generate a cartoonized abstract background in combination with edge-preserving filtering, and output the abstracted background features; Input the enhanced subject features and abstracted background features into the decoder, splice them after bilinear upsampling alignment to obtain the spliced features; Input the spliced features into the dynamic LUT module, select the pre-trained style according to the HSV histogram, calculate the SE attention for each of the RGB three channels respectively, suppress the overexposed regions, complete color enhancement, and output the cartoonized result.
7. An intelligent cartoonization system for music stage performance videos, characterized in that, It includes: A data acquisition module for acquiring the video to be processed; Preprocess the video to be processed to obtain the video sequence, optical flow field and illumination intensity distribution map corresponding to the video to be processed; A cartoonization processing module for inputting the video sequence, optical flow field and illumination intensity distribution map into an intelligent cartoonization model for cartoonization processing to obtain the cartoonization result corresponding to the video to be processed; wherein, the cartoonization model includes an encoder, a spatio-temporal attention processor and a decoder connected in sequence, and the spatio-temporal attention processor includes an optical flow-feature fusion module, a lightweight U-Net branch and a sub-region processing module connected in sequence, and the sub-region processing module includes a main body path and a background path arranged in parallel.
8. An electronic device, characterized in that, It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute an intelligent cartoonization method for a music stage performance video according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements an intelligent cartoonization method for a music stage performance video according to any one of claims 1-6.