AI dynamic graph generation system based on visual communication and cross-media adaptation method

Through the AI ​​dynamic graphics generation system based on visual communication, space-time correlation features are extracted, adaptive gamma curves are constructed, color correction and layered color space conversion are carried out, and the problems of color drift, dynamic range conversion and resolution adaptation in dynamic graphics cross-media adaptation are solved, achieving efficient dynamic graphics adaptation and visual effect improvement.

CN120075629AInactive Publication Date: 2025-05-30SHANXI AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313437.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing dynamic graphics cross-media adaptation solutions have significant shortcomings in the timing color drift of high-speed motion areas, lack of local adaptability in dynamic range conversion, and the resolution adaptation algorithm ignores spatial and temporal correlation, resulting in visual shading, color faults, motion trajectory fractures or blur, and it is difficult to coordinately optimize color fidelity, motion fluency and equipment parameter adaptability.

Method used

An AI dynamic graphics generation system based on visual communication is adopted, which includes an acquisition and processing module, a timing feature extraction module, a color correction module and a dynamic image generation module. By obtaining the continuous frame sequence of the input dynamic graphics and the color space parameters of the target media, extracting the feature tensors of space-time correlation, constructing an adaptive gamma curve for nonlinear color correction, performing hierarchical color space conversion and timing optimization, and generating dynamic graphics that meet the target media format conditions.

Benefits of technology

The cross-media precise adaptation of dynamic graphics is achieved, which significantly improves the color accuracy and visual effects of dynamic graphics, so that it can better adapt to the display characteristics of the target media, enhances the time and space consistency of dynamic graphics, reduces visual interference caused by differences in the acquisition and processing process, and improves the automation and intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075629A_ABST
    Figure CN120075629A_ABST
Patent Text Reader

Abstract

The invention relates to an AI dynamic graph generation system based on visual communication and a cross-media adaptation method. The system comprises an acquisition processing module, a time sequence feature extraction module, a color correction module and a dynamic image generation module. The acquisition processing module acquires a dynamic graph continuous frame sequence and target media parameters; the time sequence feature extraction module generates a space-time correlation feature tensor and captures the motion features of the dynamic elements; the color correction module constructs an adaptive gamma curve and optimizes a dynamic range; and the dynamic image generation module maps the correction frame, and improves the motion coherence through a time sequence optimization algorithm. According to the method, through a cooperative mechanism of optical flow field feature fusion, dynamic gamma correction and layered mapping, multi-dimensional adaptation of a dynamic graph in a color gamut, a dynamic range and a resolution ratio is realized, the problems of visual smear, highlight overexposure and motion blur in a cross-media scene are solved, and the consistency of multi-terminal visual communication is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dynamic graphics processing, and particularly to an AI dynamic graphics generation system based on visual communication and a cross-media adaptation method. Background Art

[0002] With the diversified development of multimedia display terminals, dynamic graphics face adaptation challenges such as color distortion and motion blur when presented across platforms. An AI dynamic graphics generation system based on visual communication needs to be compatible with the display characteristics of different devices (such as color gamut, resolution, dynamic range), and at the same time, it is necessary to ensure the color stability and motion coherence of dynamic elements in the time series dimension. Such a system realizes the full-process optimization of dynamic graphics from content generation to multi-terminal adaptation through intelligent technologies, which is the key to improving the cross-media visual communication effect.

[0003] Existing dynamic graphics cross-media adaptation solutions have significant deficiencies: First, the problem of temporal color drift in high-speed motion areas is prominent. The color gamut difference between devices causes the hue shift of consecutive frames, resulting in visual ghosting and color banding. Second, there is a lack of local adaptability during dynamic range conversion, and global gamma adjustment is likely to cause overexposure of highlights or loss of shadow details. Third, the resolution adaptation algorithm ignores spatio-temporal correlation, resulting in broken or blurred motion trajectories. In addition, existing methods are difficult to synergistically optimize color fidelity, motion smoothness, and device parameter adaptability, severely restricting the visual expressiveness in multiple scenarios. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide an AI dynamic graphics generation system based on visual communication and a cross-media adaptation method that can achieve precise cross-media adaptation of dynamic graphics.

[0005] The purpose of the present invention is achieved by the following solutions:

[0006] In the first aspect, the present invention provides an AI dynamic graphics generation system based on visual communication, and the system includes:

[0007] An acquisition and processing module is used to obtain a continuous frame sequence of input dynamic graphics and color space parameters of the target media;

[0008] A temporal feature extraction module is used to process the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatio-temporally associated feature tensor;

[0009] A color correction module is used for:

[0010] Processing based on the feature tensor and color space parameters to construct an adaptive gamma curve;

[0011] Performing non-linear color correction processing on each frame of the continuous frame sequence based on the adaptive gamma curve to obtain a corrected frame sequence;

[0012] A dynamic image generation module, which is used to input a corrected frame sequence into a hierarchical color space conversion model for mapping processing to generate a mapped frame sequence; and is used to perform temporal optimization processing on the mapped frame sequence to generate an output dynamic graphic, and the output dynamic graphic is used to represent a dynamic graphic that meets the conditions of the target media format.

[0013] In one embodiment, the corrected frame sequence includes a lightness channel, an a channel, and a b channel. Inputting the corrected frame sequence into the hierarchical color space conversion model for mapping processing to generate a mapped frame sequence includes:

[0014] Processing the lightness channel based on a velocity perception compensation formula to generate a compensated lightness channel;

[0015] Processing the a channel based on the deformable convolution technology to generate an adjusted a channel;

[0016] Processing the b channel based on a color temperature matching function to generate a target b channel;

[0017] Performing channel synthesis processing on the compensated lightness channel, the adjusted a channel, and the target b channel to generate a mapped frame sequence.

[0018] In one embodiment, processing the a channel based on the deformable convolution technology to generate an adjusted a channel includes:

[0019] Initializing the sampling point positions of the deformable convolution kernel based on the hue distribution data of the corrected frame sequence;

[0020] Dynamically adjusting the convolution kernel weights based on the displacement data of the feature tensor;

[0021] Performing convolution operation processing on the a channel based on the adjusted convolution kernel to generate an adjusted a channel.

[0022] In one embodiment, performing temporal optimization processing on the mapped frame sequence to generate an output dynamic graphic includes:

[0023] Performing linear combination on the lightness and chromaticity data of the mapped frame sequence to generate observed data;

[0024] Constructing a state transition matrix according to the displacement data of the feature tensor, and the state transition matrix is used to indicate the frame - to - frame state evolution of the feature tensor;

[0025] Constructing a process noise covariance matrix based on the hue gradient data of the mapped frame sequence;

[0026] Performing Kalman filter prediction processing on the mapped frame sequence based on the state transition matrix to generate a predicted state vector;

[0027] Process the predicted state vector and the observed data based on the process noise covariance matrix to generate an output dynamic graph.

[0028] In one embodiment, process the feature tensor and the color space parameters to construct an adaptive gamma curve, including:

[0029] Perform multi-scale normalization processing on the feature tensor based on the spatio-temporal joint convolutional network to generate a normalized feature vector;

[0030] Calculate the dynamic range difference ratio based on the dynamic range value of the target media color space parameters. The calculation formula is:

[0031]

[0032] where δ R is the dynamic range difference ratio, is the normalized feature vector, R target is the target color space reference value, is the maximum luminance value in the source image or source data, is the maximum luminance value in the target image or target data;

[0033] Perform truncation processing on the dynamic range difference ratio based on a preset difference threshold to generate a dynamic range difference parameter;

[0034] Construct an adaptive gamma curve based on the dynamic range difference parameter.

[0035] In one embodiment, perform temporal color extraction processing on the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatio-temporal associated feature tensor, including:

[0036] Estimate the optical flow field of adjacent frames of the continuous frame sequence according to the PWC-Net technology to obtain a motion vector matrix;

[0037] Perform feature fusion processing on the CIE-Lab data and the motion vector matrix according to the spatio-temporal joint convolutional network to obtain a feature response. The formula for calculating the response feature is:

[0038]

[0039] where, is the response feature, σ is a preset temporal attenuation coefficient, is the channel splicing operation, is the motion vector matrix, represents the color data at time t + k, and Conv3D is a three-dimensional convolution operation;

[0040] Process the response feature according to the L2 normalization technology to obtain a spatio-temporal associated feature tensor.

[0041] In one embodiment, the system further includes an image optimization module, and the image optimization module is configured to:

[0042] Calculate the hue difference degree of the output dynamic image according to the hue data of adjacent frames of the output dynamic graphics, and the calculation formula of the hue difference degree is:

[0043]

[0044] where D t is the hue difference degree, is the hue vector of the i-th pixel in the t-th frame, is the hue vector of the i-th pixel in the (t - 1)-th frame;

[0045] Perform backpropagation optimization on the convolution kernel weights of the spatio-temporal joint convolutional network based on the hue difference degree, and update the spatio-temporal joint convolutional network.

[0046] In one embodiment, performing backpropagation optimization on the convolution kernel weights of the spatio-temporal joint convolutional network based on the hue difference degree and updating the spatio-temporal joint convolutional network includes:

[0047] Calculate the network loss value based on the difference between the hue difference degree and a preset threshold;

[0048] Update the 3D convolution kernel parameters of the spatio-temporal joint convolutional network based on the gradient descent algorithm;

[0049] Optimize the spatio-temporal joint convolutional network based on the updated convolution kernel parameters, and update the spatio-temporal joint convolutional network.

[0050] In a second aspect, the present invention provides a cross-media adaptation method for AI dynamic graphics generation based on visual communication, including the following steps:

[0051] Obtain the continuous frame sequence of the continuous graphics of the input dynamic graphics and the color space parameters of the target media;

[0052] Process the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatio-temporally associated feature tensor;

[0053] Process the feature tensor and the color space parameters to obtain a dynamic range difference parameter;

[0054] Construct an adaptive gamma curve based on the dynamic range difference parameter;

[0055] Perform non-linear color correction processing on each frame of the continuous frame sequence based on the adaptive gamma curve to obtain a corrected frame sequence;

[0056] Input the corrected frame sequence into the hierarchical color space conversion model for mapping processing to generate the mapped frame sequence;

[0057] Perform temporal optimization processing on the mapped frame sequence to generate the output dynamic graphic, and the output dynamic graphic is used to represent the dynamic graphic that meets the conditions of the target media format.

[0058] In one embodiment, performing temporal optimization processing on the mapped frame sequence to generate the output dynamic graphic includes:

[0059] Perform a linear combination of the lightness and chromaticity data of the mapped frame sequence to generate the observed data;

[0060] Construct a state transition matrix according to the displacement data of the feature tensor, and the state transition matrix is used to indicate the inter-frame state evolution of the feature tensor;

[0061] Construct a process noise covariance matrix based on the hue gradient data of the mapped frame sequence;

[0062] Perform Kalman filter prediction processing on the mapped frame sequence based on the state transition matrix to generate a predicted state vector;

[0063] Process the predicted state vector and the observed data based on the process noise covariance matrix to generate the output dynamic graphic.

[0064] In summary, an AI dynamic graphic generation system based on visual communication provided by this application realizes a complete dynamic graphic processing solution through the acquisition, feature extraction, color correction, and finally color space mapping and temporal optimization of the input dynamic graphic, thereby significantly improving the color accuracy and visual effect of the dynamic graphic, making it better adapt to the display characteristics of the target media, while enhancing the temporal and spatial consistency of the dynamic graphic and reducing visual interference caused by differences in the acquisition and processing processes. In addition, this process improves the automation and intelligence level of the system, and through the adaptive gamma curve and deep learning model, it can flexibly respond to different scenarios, providing a high-quality data foundation for subsequent dynamic graphic analysis and applications, and contributing to further video understanding, editing, and transmission operations.

[0065] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0066] Figure 1 It is a structural block diagram of an AI dynamic graphic generation system based on visual communication provided by an embodiment of this application;

[0067] Figure 2 It is a flow schematic diagram of the steps executed by the color correction module provided by an embodiment of this application;

[0068] Figure 3 Schematic diagram of the process for generating a mapped frame sequence provided by an embodiment of the present application;

[0069] Figure 4 Schematic diagram of the process for a cross-media adaptation method for AI dynamic graphic generation based on visual communication provided by another embodiment of the present application. Detailed implementation manners

[0070] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the invention more thorough and comprehensive.

[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0072] Embodiment 1: Please refer to Figure 1 , which shows an AI dynamic graphic generation system based on visual communication provided by an embodiment of the present application. In this embodiment, it is exemplified that the system is applied to a terminal. It can be understood that the system can also be applied to a server and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. This embodiment is applicable to multi-platform advertising content delivery, cross-media distribution of film and television content, VR / AR multi-device compatibility, and multi-terminal display of e-commerce live broadcasts. As Figure 1 shown, the AI dynamic graphic generation system 100 based on visual communication provided by the present invention includes an acquisition and processing module 110, a timing feature extraction module 120, a color correction module 130, and a dynamic image generation module 140.

[0073] Exemplarily, the acquisition and processing module 110 is used to obtain a continuous frame sequence of consecutive graphics of the input dynamic graphic and the color space parameters of the target media.

[0074] Specifically, the acquisition and processing module 110 can capture dynamic graphics at a specific frame rate through a high-precision graphics acquisition device (such as a high-speed camera, a graphics sensor, etc.), obtaining a continuous frame sequence. Meanwhile, using a color analyzer or a built-in color detection algorithm, color space parameters of the target media (such as a display, a projection device, etc.) are extracted, including the gamut range, brightness, contrast, etc. For example, in professional film and television production, a high-precision CIE colorimeter is used to measure the color space of a display to ensure color accuracy. Preferably, a multi-angle acquisition device can also be combined to obtain the frame sequence of the dynamic graphics from different perspectives to achieve a more abundant three-dimensional information restoration.

[0075] The temporal feature extraction module 120 is used to process the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence, generating a spatio-temporal associated feature tensor.

[0076] Preferably, the Horn-Schunck algorithm or the Lucas-Kanade algorithm can be used to calculate the optical flow field between adjacent frames to obtain the motion information of pixels. Then, based on this optical flow field information, a spatio-temporal associated feature tensor is generated through a feature extraction algorithm to capture the temporal and spatial associations between frames. Preferably, an improved Lucas-Kanade algorithm can be used to calculate the optical flow field in video analysis to improve the tracking accuracy of fast-moving objects. Meanwhile, the Temporal Convolutional Network in deep learning can be combined to further enhance the feature extraction ability for long time series.

[0077] Preferably, as Figure 2 shown, the color correction module 130 is used to perform the following steps:

[0078] S210. Process based on the feature tensor and the color space parameters to construct an adaptive gamma curve.

[0079] Preferably, the adaptive gamma curve is constructed through the following steps:

[0080] S211. Perform multi-scale normalization processing on the feature tensor based on a spatio-temporal joint convolutional network to generate a normalized feature vector.

[0081] Specifically, the color correction module 130 performs convolution operations on the feature tensor at different scales by constructing a spatio-temporal joint convolution network that includes multiple convolutional layers and pooling layers. After completing the convolution operation, the feature map is normalized through techniques such as Batch Normalization or Layer Normalization to eliminate scale differences, and finally a normalized feature vector is generated. For example, in a video classification task, a 3D convolutional network is used to extract spatio-temporal features from video frames, and the Batch Normalization technique is combined to improve the robustness of the model. Preferably, an attention mechanism can be adopted to assign weights to feature maps at different scales to highlight important features.

[0082] S212. Calculate the dynamic range difference ratio based on the dynamic range value of the target media color space parameters. The calculation formula is:

[0083]

[0084] where, δ R is the dynamic range difference ratio, is the normalized feature vector, R target is the reference value of the target color space, is the maximum brightness value in the source image or source data, is the maximum brightness value in the target image or target data.

[0085] S213. Perform truncation processing on the dynamic range difference ratio based on a preset difference threshold to generate a dynamic range difference parameter.

[0086] Specifically, set a preset difference threshold θ, usually determined according to the characteristics of the target media and visual perception requirements, and limit the dynamic range difference ratio δ R within the preset threshold range. If δ R exceeds the threshold θ, then truncate it to θ; if δ R is less than 0, then truncate it to 0; otherwise, keep δ R unchanged. Finally, use the truncated dynamic range difference ratio as the dynamic range difference parameter for subsequent processing.

[0087] S214. Construct an adaptive gamma curve based on the dynamic range difference parameter.

[0088] Specifically, construct an adaptive gamma curve through the following formula:

[0089]

[0090] where, γ is the gamma value, and δ clipped is the dynamic range difference parameter.

[0091] S220. Perform non - linear color correction processing on each frame of the continuous frame sequence based on an adaptive gamma curve to obtain a corrected frame sequence.

[0092] Specifically, apply this curve to each frame of the continuous frame sequence for non - linear color correction processing. For example, in low - light image enhancement, by dynamically adjusting the gamma curve, the brightness and contrast of the image are improved. Preferably, the color constancy algorithm can be combined to further optimize the color correction effect, making the corrected image more in line with the human visual characteristics.

[0093] The dynamic image generation module 140 is used to input the corrected frame sequence into a hierarchical color space conversion model for mapping processing to generate a mapped frame sequence; and is used to perform temporal optimization processing on the mapped frame sequence to generate an output dynamic graphic, and the output dynamic graphic is used to represent a dynamic graphic that meets the conditions of the target media format.

[0094] Specifically, the dynamic image generation module 140 constructs a hierarchical color space conversion model, which includes color space conversion matrices at multiple levels. Map the corrected frame sequence layer by layer to the target color space to generate a mapped frame sequence. Then, perform temporal optimization processing on the mapped frame sequence through temporal filtering or interpolation algorithms to eliminate temporal discontinuities and jitters, and finally generate an output dynamic graphic that meets the conditions of the target media format. For example, in video transcoding, use a color space conversion matrix to convert the YUV format to the RGB format and optimize the timing through motion compensation techniques. Preferably, the generative adversarial network (GAN) in deep learning can be used to further optimize the mapped frame sequence to improve the quality and realism of the image.

[0095] In summary, an AI dynamic graphic generation system based on visual communication provided by this application realizes a complete dynamic graphic processing solution through the acquisition, feature extraction, color correction of the input dynamic graphic to the final color space mapping and temporal optimization, thereby significantly improving the color accuracy and visual effect of the dynamic graphic, making it better adapt to the display characteristics of the target media, while enhancing the temporal and spatial consistency of the dynamic graphic and reducing visual interference caused by differences in the acquisition and processing processes. In addition, this process improves the automation and intelligence level of the system. Through the adaptive gamma curve and deep learning models, it can flexibly respond to different scenarios, providing a high - quality data basis for subsequent dynamic graphic analysis and applications, and contributing to further video understanding, editing, and transmission operations.

[0096] In one embodiment, the corrected frame sequence includes a lightness channel, an a channel, and a b channel, such as Figure 3As shown, input the correction frame sequence into the hierarchical color space conversion model for mapping processing to generate the mapped frame sequence, including the following steps:

[0097] S310. Process the lightness channel based on the velocity perception compensation formula to generate a compensated lightness channel.

[0098] Specifically, the velocity perception compensation formula usually dynamically adjusts the lightness value based on the motion speed of pixels. For example, the following formula can be used:

[0099] L comp =L original +α*v

[0100] where L comp is the compensated lightness value, L original is the original lightness value, α is the compensation coefficient, and v is the motion speed of the pixel.

[0101] S320. Process the a channel based on the deformable convolution technology to generate an adjusted a channel.

[0102] Preferably, the adjusted a channel is generated through the following steps:

[0103] S321. Initialize the sampling point positions of the deformable convolution kernel based on the hue distribution data of the correction frame sequence.

[0104] Specifically, analyze the hue distribution data of the correction frame sequence to obtain the main color tone and color distribution characteristics of the image. According to these characteristics, initialize the sampling point positions of the deformable convolution kernel. For example, use the k-means clustering algorithm to cluster the hue values to obtain several main hue regions, and then initialize the sampling points of the convolution kernel according to the central positions of these regions. The specific steps are: extract the hue channel data of the correction frame sequence, perform clustering analysis on the hue channel data to obtain the main hue regions, and initialize the sampling point positions of the deformable convolution kernel according to the central positions of the main hue regions.

[0105] S322. Dynamically adjust the convolution kernel weights based on the displacement data of the feature tensor.

[0106] Specifically, the displacement data is usually generated by the optical flow field or the feature matching algorithm, representing the motion information of the pixels. The specific steps are: extract the displacement data in the feature tensor, calculate the offset of each sampling point according to the displacement data, use the bilinear interpolation or cubic interpolation method to calculate the feature values of the offset sampling points, and dynamically adjust the weights of the convolution kernel according to the new feature values.

[0107] S323. Perform convolution operation on the a channel based on the adjusted convolution kernel to generate an adjusted a channel.

[0108] Specifically, the color correction module 130 applies the adjusted convolution kernel to the a channel, performs convolution operations, and normalizes the convolution results to ensure that the values of the a channel are within a reasonable range. After the normalization process, the adjusted a channel is generated. Preferably, multi-scale convolution technology can be combined to extract and fuse a-channel features at different scales to improve the robustness of the processing.

[0109] S330. Process the b channel based on the color temperature matching function to generate the target b channel.

[0110] Specifically, the color temperature matching function usually adjusts the color temperature of the image to the target value based on the definition of color temperature. The color temperature matching function is as follows:

[0111]

[0112] where, b target is the target b-channel value, b original is the original b-channel value, T target is the target color temperature, and T original is the original color temperature. Preferably, a color temperature sensor can be combined to obtain the ambient color temperature in real time to dynamically adjust the color temperature of the image.

[0113] S340. Perform channel synthesis processing on the compensated lightness channel, the adjusted a channel, and the target b channel to generate a mapped frame sequence.

[0114] Specifically, perform channel synthesis processing on the compensated lightness channel, the adjusted a channel, and the target b channel. The specific steps are as follows: First, combine the pixel values of the three channels in sequence to generate new pixel values. Second, normalize the combined pixel values to ensure that they are within a reasonable range. Finally, process the normalized combined pixel values to generate the final mapped frame sequence.

[0115] In one embodiment, perform temporal optimization processing on the mapped frame sequence to generate an output dynamic graphic, including:

[0116] S410. Perform a linear combination of the lightness and chromaticity data of the mapped frame sequence to generate observed data.

[0117] Specifically, the formula for the linear combination can be expressed as:

[0118] y = W * I

[0119] where, y is the observed data, W is the weight matrix, and I is the vector of lightness and chromaticity data.

[0120] S420. Construct a state transition matrix based on the displacement data of the feature tensor. The state transition matrix is used to indicate the inter-frame state evolution of the feature tensor.

[0121] Specifically, the state transition matrix can be expressed as:

[0122]

[0123] where Δt is the time interval, representing the time difference between adjacent frames. The displacement data can be the pixel displacement calculated through the optical flow field or feature matching algorithm. The state transition matrix F is used to predict the state of the next frame, thereby realizing the temporal modeling of dynamic graphics.

[0124] S430. Construct the process noise covariance matrix based on the hue gradient data of the mapped frame sequence.

[0125] Specifically, the hue gradient data reflects the rate and direction of color change in the image and can be used to estimate the magnitude and distribution of process noise. The process noise covariance matrix Q is usually expressed as:

[0126]

[0127] where is the variance of hue, is the variance of chroma.

[0128] S440. Perform Kalman filter prediction processing on the mapped frame sequence based on the state transition matrix to generate a predicted state vector.

[0129] Specifically, the Kalman filter prediction processing predicts the state of the next frame through the state transition matrix F, thereby realizing the temporal modeling and prediction of dynamic graphics. The prediction process can be expressed as:

[0130] x pred = F * x prev

[0131] where x pred is the predicted state vector, and x prev is the state vector of the previous frame.

[0132] S450. Process the predicted state vector and the observed data based on the process noise covariance matrix to generate the output dynamic graphics.

[0133] Specifically, the dynamic image generation module 140 calculates the Kalman gain based on the predicted state vector and the observed data, and uses the Kalman gain and the observed data to update the predicted state vector to generate the final state vector. Finally, the updated state vector is converted into the pixel values of the dynamic graphics to generate the output dynamic graphics.

[0134] In summary, the AI dynamic graphics generation system based on visual communication provided by this application realizes a complete dynamic graphics processing solution through the linear combination of lightness and chromaticity data, the construction of the state transition matrix, the estimation of the process noise covariance matrix, the Kalman filter prediction processing, and the final output of dynamic graphics generation. This can significantly improve the color accuracy and visual effect of dynamic graphics, making it better adapt to the display characteristics of the target media. At the same time, through the adaptive processing method, the temporal and spatial consistency of dynamic graphics is enhanced, and the visual interference caused by differences in the acquisition and processing process is reduced.

[0135] In one embodiment, this embodiment is basically the same as Embodiment 1, except that the sequential color extraction process is performed on the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatio-temporal associated feature tensor, including:

[0136] S510. Estimate the optical flow field of adjacent frames of the continuous frame sequence according to the PWC-Net technology to obtain a motion vector matrix.

[0137] Specifically, PWC-Net accurately estimates the optical flow field between adjacent frames through multi-scale feature extraction and Cost Volume calculation. The specific steps include:

[0138] 1. Feature extraction: Use a convolutional neural network to extract features from the input continuous frame sequence to generate multi-scale feature maps.

[0139] 2. Cost Volume calculation: Generate a Cost Volume by calculating the feature correlation between the reference frame and the target frame. The dimension of the Cost Volume is (Hl, Wl, D2), where Hl and Wl are the width and height of the image data of the l-th layer pyramid, and D is the size of the search window. For example, if the search range is d, then D = 2d + 1.

[0140] 3. Optical flow estimation: Use the optical flow estimation module to process the Cost Volume to generate a motion vector matrix. PWC-Net gradually refines the estimation of the optical flow field through a multi-scale and pyramid structure to improve the accuracy.

[0141] S520. Perform feature fusion processing on the CIE-Lab data and the motion vector matrix according to the spatio-temporal joint convolutional network to obtain a feature response. The formula for calculating the response feature is:

[0142]

[0143] Among them, is the response feature, σ is a preset temporal decay coefficient, is the channel splicing operation, is the motion vector matrix, The color data at time t + k is represented, and Conv3D is a three-dimensional convolution operation.

[0144] S530. Process the response features according to the L2 normalization technique to obtain a spatio-temporal correlation feature tensor.

[0145] Preferably, perform L2 normalization processing on the response features to ensure that the length of the feature vector is 1, and then organize the normalized feature vectors into a feature tensor to generate a spatio-temporal correlation feature tensor. This feature tensor can be used for subsequent dynamic graphics processing and analysis, where L2 normalization can be expressed as:

[0146]

[0147] where F nor is the normalized response feature, is the L2 norm of the response feature.

[0148] In one embodiment, this embodiment is basically the same as Embodiment 1, except that the AI dynamic graphics generation system 100 based on visual communication further includes an image optimization module 150, and the image optimization module 150 is used to perform the following steps:

[0149] S610. Calculate the output dynamic image according to the adjacent frame hue data of the output dynamic graphics to obtain the hue difference degree, and the formula for calculating the hue difference degree is:

[0150]

[0151] where D t is the hue difference degree, is the hue vector of the i-th pixel in the t-th frame, is the hue vector of the i-th pixel in the (t - 1)-th frame;

[0152] S620. Perform backpropagation optimization processing on the convolution kernel weights of the spatio-temporal joint convolution network based on the hue difference degree to update the spatio-temporal joint convolution network.

[0153] Preferably, the spatio-temporal joint convolution network is updated through the following steps:

[0154] S621. Calculate the network loss value based on the difference between the hue difference degree and the preset threshold.

[0155] Specifically, the loss value is used to reflect the difference between the network output and the target, and the loss value is obtained through the following formula:

[0156]

[0157] S622. Update the 3D convolution kernel parameters of the spatio-temporal joint convolution network based on the gradient descent algorithm.

[0158] Preferably, the image optimization module calculates the gradient of the loss value with respect to the 3D convolution kernel parameters through the backpropagation algorithm and updates the 3D convolution kernel parameters based on the gradient descent algorithm. The update formula is as follows:

[0159]

[0160] where, W new is the updated 3D convolution kernel parameter, W old is the 3D convolution kernel parameter before update, and η is the learning rate.

[0161] S623. Optimize the spatio-temporal joint convolution network based on the updated convolution kernel parameters and update the spatio-temporal joint convolution network.

[0162] Specifically, first apply the updated convolution kernel parameters to the network to replace the old parameters to ensure that the network can better capture the spatio-temporal features of dynamic graphics. Then, calculate the network output through forward propagation, evaluate the difference between the output and the true label using the loss function, calculate the gradient through backpropagation, and update the network parameters using optimization algorithms such as gradient descent. This process is repeated on the training data until the network performance reaches the optimal. Finally, evaluate the optimized network on the validation set, verify its performance improvement through metrics such as accuracy, and ensure that the dynamic graphics processing effect meets the expectations. This process effectively improves the color accuracy and visual effect of dynamic graphics, making it better adapt to the display characteristics of the target media, enhances the temporal and spatial consistency of dynamic graphics, reduces visual interference, improves the automation and intelligence level of the system, and provides a high-quality data basis for subsequent dynamic graphics analysis and applications.

[0163] In summary, the AI dynamic graphics generation system based on visual communication provided by the embodiments of the present application realizes a complete dynamic graphics processing solution through the calculation of hue difference degree to the backpropagation optimization processing of the spatio-temporal joint convolution network. Thereby, it can significantly improve the color accuracy and visual effect of dynamic graphics, making it better adapt to the display characteristics of the target media. Through the adaptive processing method, the temporal and spatial consistency of dynamic graphics is enhanced, and visual interference caused by differences in the acquisition and processing process is reduced.

[0164] In one of the embodiments, as Figure 4 shown, the present invention also provides a cross-media adaptation method for AI dynamic graphics generation based on visual communication. This method is applied to the AI dynamic graphics generation system based on visual communication provided in any of the above embodiments. The cross-media adaptation method includes the following steps:

[0165] S710. Obtain a sequence of consecutive frames of the consecutive graphics of the input dynamic graphic and the color space parameters of the target media.

[0166] S720. Process the sequence of consecutive frames based on the inter-frame optical flow field of the sequence of consecutive frames to generate a spatio-temporal correlated feature tensor.

[0167] S730. Process the feature tensor and the color space parameters to obtain a dynamic range difference parameter.

[0168] S740. Construct an adaptive gamma curve based on the dynamic range difference parameter.

[0169] S750. Perform non-linear color correction processing on each frame of the sequence of consecutive frames based on the adaptive gamma curve to obtain a corrected frame sequence.

[0170] S760. Input the corrected frame sequence into a hierarchical color space conversion model for mapping processing to generate a mapped frame sequence.

[0171] S770. Perform temporal optimization processing on the mapped frame sequence to generate an output dynamic graphic, and the output dynamic graphic is used to represent a dynamic graphic that meets the conditions of the target media format.

[0172] Preferably, the output dynamic graphic is generated through the following steps:

[0173] S771. Perform a linear combination on the lightness and chromaticity data of the mapped frame sequence to generate observed data.

[0174] S772. Construct a state transition matrix according to the displacement data of the feature tensor, and the state transition matrix is used to indicate the inter-frame state evolution of the feature tensor.

[0175] S773. Construct a process noise covariance matrix based on the hue gradient data of the mapped frame sequence.

[0176] S774. Perform Kalman filter prediction processing on the mapped frame sequence based on the state transition matrix to generate a predicted state vector.

[0177] S775. Process the predicted state vector and the observed data based on the process noise covariance matrix to generate an output dynamic graphic.

[0178] In summary, a cross-media adaptation method for AI dynamic graphics generation based on visual communication provided by the present application realizes a complete set of dynamic graphics processing solutions through the acquisition, feature extraction, color correction of the input dynamic graphics to the final color space mapping and temporal optimization. Thereby, it significantly improves the color accuracy and visual effects of dynamic graphics, enables them to better adapt to the display characteristics of the target media, enhances the temporal and spatial consistency of dynamic graphics, and reduces visual interference caused by differences in the acquisition and processing processes. In addition, this process improves the automation and intelligence level of the system. Through the adaptive gamma curve and deep learning model, it can flexibly respond to different scenarios, provides a high-quality data basis for subsequent dynamic graphics analysis and applications, and contributes to further operations such as video understanding, editing, and transmission.

[0179] Meanwhile, a cross-media adaptation method for AI dynamic graphics generation based on visual communication provided by the present application realizes a complete set of dynamic graphics processing solutions through the calculation of hue difference degree to the backpropagation optimization processing of the spatio-temporal joint convolutional network. Thereby, it can significantly improve the color accuracy and visual effects of dynamic graphics, enables them to better adapt to the display characteristics of the target media. Through the adaptive processing method, it enhances the temporal and spatial consistency of dynamic graphics and reduces visual interference caused by differences in the acquisition and processing processes.

[0180] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0181] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.

[0182] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. AI dynamic graphics generation system based on visual communication, characterized by: It includes acquisition processing module, time series feature extraction module, color correction module and dynamic image generation module: The acquisition processing module is used to obtain a continuous frame sequence of continuous graphics of the input dynamic graphics and color space parameters of the target media; A temporal feature extraction module, used for processing the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a temporal and spatially associated feature tensor; Color correction module for: Performing processing based on the feature tensor and the color space parameters to construct an adaptive gamma curve; Performing nonlinear color correction processing on each frame of the continuous frame sequence based on the adaptive gamma curve to obtain a corrected frame sequence; A dynamic image generation module, used for inputting the corrected frame sequence into a layered color space conversion model for mapping processing to generate a mapped frame sequence; The method is also used to perform timing optimization processing on the mapping frame sequence to generate output dynamic graphics, wherein the output dynamic graphics are used to express dynamic graphics that meet the target media format conditions.

2. The AI ​​dynamic graphics generation system according to claim 1, characterized in that: The corrected frame sequence includes a brightness channel, an a channel, and a b channel. The corrected frame sequence is input into a layered color space conversion model for mapping processing to generate a mapped frame sequence, including: Processing the brightness channel based on a speed perception compensation formula to generate a compensated brightness channel; Processing the a channel based on a deformable convolution technique to generate an adjusted a channel; Process the b channel based on the color temperature matching function to generate a target b channel; The compensated brightness channel, the adjusted a channel and the target b channel are subjected to channel synthesis processing to generate a mapping frame sequence.

3. The AI ​​dynamic graphics generation system according to claim 2, characterized in that: The step of processing the a channel based on the deformable convolution technology to generate an adjusted a channel includes: Initializing the sampling point positions of the deformable convolution kernel based on the hue distribution data of the corrected frame sequence; Based on the displacement data of the feature tensor, dynamically adjusting the convolution kernel weight; A convolution operation is performed on the a channel based on the adjusted convolution kernel to generate an adjusted a channel.

4. The AI ​​dynamic graphics generation system according to claim 1, characterized in that: The step of performing timing optimization processing on the mapping frame sequence to generate an output dynamic graphic comprises: Linearly combining the brightness and chromaticity data of the mapping frame sequence to generate observation data; Constructing a state transfer matrix according to the displacement data of the feature tensor, wherein the state transfer matrix is ​​used to indicate the inter-frame state evolution of the feature tensor; Based on the hue gradient data of the mapping frame sequence, constructing a process noise covariance matrix; Performing Kalman filter prediction processing on the mapping frame sequence based on the state transfer matrix to generate a predicted state vector; The predicted state vector and the observed data are processed based on the process noise covariance matrix to generate the output dynamic graph.

5. The AI ​​dynamic graphics generation system according to claim 1, characterized in that: The processing of the feature tensor and the color space parameter to construct an adaptive gamma curve includes: Performing multi-scale normalization processing on the feature tensor based on a spatiotemporal joint convolutional network to generate a normalized feature vector; Based on the dynamic range value of the target media color space parameter, the dynamic range difference ratio is calculated using the following formula: Among them, δ R is the dynamic range difference ratio, is the normalized eigenvector, R target is the target color space reference value, is the maximum brightness value in the source image or source data, is the maximum brightness value in the target image or target data; Performing truncation processing on the dynamic range difference ratio based on a preset difference threshold to generate a dynamic range difference parameter; The adaptive gamma curve is constructed based on the dynamic range difference parameter.

6. The AI ​​dynamic graphics generation system according to claim 1, characterized in that: The step of performing temporal color extraction processing on the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatiotemporally associated feature tensor includes: Performing optical flow field estimation processing on adjacent frames of the continuous frame sequence according to the PWC-Net technology to obtain a motion vector matrix; The CIE-Lab data and the motion vector matrix are subjected to feature fusion processing according to the spatiotemporal joint convolutional network to obtain a feature response. The formula for calculating the response feature is: in, is the response characteristic, σ is the preset time series attenuation coefficient, For channel splicing operations, is the motion vector matrix, Represents the color data at time t+k, Conv3D is a three-dimensional convolution operation; The response features are processed according to the L2 normalization technique to obtain a spatiotemporal associated feature tensor.

7. The AI ​​dynamic graphics generation system according to claim 6, characterized in that: The system further comprises an image optimization module, wherein the image optimization module is used to: According to the hue data of adjacent frames of the output dynamic graphics, the output dynamic image is calculated to obtain the hue difference, and the hue difference calculation formula is: Among them, D t is the hue difference, is the hue vector of the i-th pixel in the t-th frame, is the hue vector of the i-th pixel in the t-1-th frame; Based on the hue difference, back-propagation optimization processing is performed on the convolution kernel weights of the spatiotemporal joint convolutional network to update the spatiotemporal joint convolutional network.

8. The AI ​​dynamic graphics generation system according to claim 7, characterized in that: The back-propagation optimization processing is performed on the convolution kernel weights of the spatiotemporal joint convolutional network based on the hue difference to update the spatiotemporal joint convolutional network, including: Calculating a network loss value based on the difference between the hue difference and a preset threshold; Updating the 3D convolution kernel parameters of the spatiotemporal joint convolutional network based on a gradient descent algorithm; The spatiotemporal joint convolutional network is optimized based on the updated convolution kernel parameters, and the spatiotemporal joint convolutional network is updated.

9. A cross-media adaptation method for AI dynamic graphics generation based on visual communication, applied to the AI ​​dynamic graphics generation system based on visual communication as claimed in any one of claims 1 to 8, characterized in that: The following steps are involved: Obtaining a continuous frame sequence of continuous graphics of an input dynamic graphic and a color space parameter of a target medium; Processing the continuous frame sequence based on the inter-frame optical flow field of the continuous frame sequence to generate a spatiotemporally associated feature tensor; Processing the feature tensor and the color space parameter to obtain a dynamic range difference parameter; constructing an adaptive gamma curve based on the dynamic range difference parameter; Performing nonlinear color correction processing on each frame of the continuous frame sequence based on the adaptive gamma curve to obtain a corrected frame sequence; Inputting the corrected frame sequence into a layered color space conversion model for mapping processing to generate a mapped frame sequence; The mapping frame sequence is subjected to timing optimization processing to generate an output dynamic graphic, wherein the output dynamic graphic is used to express a dynamic graphic that meets the target media format condition.

10. The cross-media adaptation method according to claim 9, characterized in that: The step of performing timing optimization processing on the mapping frame sequence to generate an output dynamic graphic comprises: Linearly combining the brightness and chromaticity data of the mapping frame sequence to generate observation data; Constructing a state transfer matrix according to the displacement data of the feature tensor, wherein the state transfer matrix is ​​used to indicate the inter-frame state evolution of the feature tensor; Based on the hue gradient data of the mapping frame sequence, constructing a process noise covariance matrix; Performing Kalman filter prediction processing on the mapping frame sequence based on the state transfer matrix to generate a predicted state vector; The predicted state vector and the observed data are processed based on the process noise covariance matrix to generate the output dynamic graph.