Video mixing and cutting system based on AIGC

By constructing a mapping relationship between semantic slots and learning AI video big data, a priori parameter set is generated. The dynamic compression module performs frequency distribution prediction and processing on AIGC mixed videos, solving the bit rate inflation problem caused by high-frequency signals in AIGC mixed videos and achieving efficient video compression.

CN120956987AActive Publication Date: 2025-11-14XIAMEN XINZHUN TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511259844.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-14
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

AIGC-edited videos may contain unnatural high-frequency signals in the time domain spectrum, leading to abnormal bitrate inflation. Traditional video compressors cannot accurately distinguish between natural and unnatural high-frequency signals, resulting in increased video size and image distortion.

Method used

The semantic mapping module constructs semantic slots, and combines AI video big data to learn semantic temporal spectrum mapping to generate a priori parameter set. The dynamic compression module identifies the causes of high frequencies and takes differentiated processing measures to suppress unnatural high-frequency signals and adjust the bit rate allocation to avoid bit rate inflation.

Benefits of technology

It effectively avoids bitrate inflation caused by non-natural high-frequency signals, maintains video quality, reduces video size, and ensures compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956987A_ABST
    Figure CN120956987A_ABST
Patent Text Reader

Abstract

The invention discloses a video mixing and cutting system based on AIGC, and relates to the technical field of video generation, the video mixing and cutting system comprises a semantic mapping module, a time-frequency prediction module and a dynamic compression module, the semantic mapping module is used for carrying out semantic analysis on cue words and constructing semantic slots, video big data is generated based on existing AI to learn mapping of semantic time-domain frequency spectrums, and the time-frequency prediction module is used for carrying out time-frequency prediction on the cue words; the time-frequency prediction module is used for carrying out time-domain frequency distribution prediction on the current AI mixed cutting video by utilizing the prior parameter set to obtain a motion track characteristic parameter of each picture element and a time-varying property of texture; the dynamic compression module is used for judging high-frequency causes and performing differentiation processing when actually measured frequency spectrum and prediction deviation are too large so as to avoid code rate abnormal expansion, and the semantic mapping module comprises a prompt word acquisition module, a semantic slot construction module, a time domain frequency spectrum reading module, a semantic analysis module and a mapping relation obtaining module. The method has the characteristic of avoiding the abnormal code rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video technology, specifically to a video mixing and editing system based on AIGC. Background Technology

[0002] Video montage refers to cutting, splicing, and recombining two or more video clips to generate a new, complete video. Newer AIGC montage technology can extract specific visual elements from video clips, merge and align them temporally, and then reconstruct them in a new video.

[0003] Compared to naturally shot videos, AIGC (AI Generic Video Editing) is the result of neural networks predicting and synthesizing frame by frame or segment by segment. It does not directly follow the constraints of the physical world and cannot accurately capture kinematic laws. It only guesses the next frame at the pixel level. Unnatural high-frequency signals (jitter, noise) may appear in the temporal spectrum of the edited video. Traditional video compressors will mistakenly believe that the unnatural high-frequency signals mainly come from the pixel differences between adjacent frames caused by the movement of screen elements. If the differences between adjacent frames are huge, more bitrate is needed to save them, resulting in abnormal bitrate expansion and thus a larger video size. If the compressor forcibly reduces the bitrate and suppresses the frequency band of high-frequency signals, the picture will be blurry and distorted.

[0004] Current technologies learn statistical patterns of motion and temporal spectrum of scene elements from large-scale video data. However, AIGC-edited videos often contain many original artistic concepts, which may themselves not conform to statistical laws, making it difficult to accurately define the range between natural and unnatural. Therefore, it is necessary to design an AIGC-based video editing system that avoids abnormal bitrates. Summary of the Invention

[0005] The purpose of this invention is to provide a video mixing and editing system based on AIGC to solve the problems mentioned in the background art.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a video mixing and editing system based on AIGC, comprising a semantic mapping module, a time-frequency prediction module, and a dynamic compression module. The semantic mapping module is used to perform semantic parsing on prompt words and construct semantic slots, learn the semantic temporal spectrum mapping based on existing AI-generated video big data, and output a priori parameter set that can be interpreted by the encoder. The time-frequency prediction module uses the priori parameter set to predict the temporal frequency distribution of the current AI-mixed video, obtaining the motion trajectory feature parameters and time-varying texture of each frame element. The dynamic compression module is used to identify high-frequency causes and differentiate processing when the measured spectrum deviates too much from the prediction, in order to avoid abnormal bitrate inflation.

[0007] According to the above technical solution, the semantic mapping module includes a prompt word acquisition module, a semantic slot construction module, a time-domain spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library. The prompt word acquisition module is electrically connected to the semantic slot construction module, the AI ​​video library is electrically connected to the time-domain spectrum reading module, the semantic analysis module is electrically connected to both the semantic slot construction module and the time-domain spectrum reading module, and the mapping relationship derivation module is electrically connected to the time-domain spectrum reading module. The prompt word acquisition module is used to receive prompt words generated during AI-generated video montage and to complete word segmentation, part-of-speech tagging, and dependency relation determination. The semantic slot construction module is used to organize structured semantics according to slots for subject, action, scene, shot, rhythm, and material. The temporal spectrum reading module is used to read the temporal spectrum of existing AI-generated videos. The semantic analysis module is used to analyze the relationship between the semantics of prompt words in existing AI-generated videos and the range of feature parameters of motion trajectories of screen elements and the range of time-varying textures. The mapping relationship derivation module is used to derive the mapping relationship between the time-varying motion trajectories of screen elements and textures and the temporal spectrum. The AI ​​video library is used to store existing AI-generated videos and the prompt words corresponding to each AI-generated video.

[0008] The time-frequency prediction module includes an element extraction module, a video generation module, a priori parameter reading module, and a time-domain spectrum pre-generation module. The priori parameter reading module is electrically connected to the element extraction module and the time-domain spectrum pre-generation module. The element extraction module is used to extract the screen elements of the user-uploaded video and abstract them into organized structural semantics. The video generation module is used to generate a new video using an AI video generation model based on the prompts input by the user and the extracted screen elements. The priori parameter reading module is used to read the organized structural semantics of the current AI mixed video, as well as the mapping relationship between the time-varying motion trajectory and texture of the screen elements and the time-domain spectrum. The time-domain spectrum pre-generation module is used to pre-generate the time-domain spectrum of the current AI mixed video based on the mapping relationship.

[0009] The dynamic compression module includes a time-domain spectrum comparison module, a causal probability calculation module, a video compression unit, a frequency band suppression module, and a bitrate allocation module. The time-domain spectrum comparison module is used to compare the time-domain spectrum of the pre-generated current AI-edited video with the time-domain spectrum of the current AI-edited video itself. The causal probability calculation module is used to calculate and analyze the causes of high-frequency components based on the difference between the two time-domain spectra. The video compression unit is used to perform video compression. The frequency band suppression module is used to perform frequency band selective suppression preprocessing on unnatural high-frequency signals that are not intended for use. The bitrate allocation module is used to ensure that natural video frames receive a sufficient bitrate during the encoding stage through bitrate control.

[0010] According to the above technical solution, the operation of the video mixing and editing system includes the following steps:

[0011] S0. Receive the user's prompt words and uploaded video materials, read the videos and their corresponding prompt words in the existing AI video library, separate the screen elements related to the prompt words, and capture the motion trajectory feature parameters and texture time-varying properties of the screen elements from beginning to end in the video;

[0012] S1. Perform semantic analysis on the prompt words in the AI ​​video library and construct semantic slots for subject, action, scene, shot, rhythm, and material. Each different combination of semantic slots represents a semantic environment. Extract the specified screen elements that need to be used in the video material uploaded by the user, generate prompt words that can represent the specified screen elements, combine the prompt words input by the user to generate the corresponding semantic slot combination, and find the closest semantic slot combination in the AI ​​video library.

[0013] S2. Based on the closest semantic slot combination, use AI to generate videos in batches, analyze all videos corresponding to this semantic environment in the AI ​​video library, obtain the temporal spectrum of all videos, and then obtain the mapping relationship between the motion trajectory feature parameters of the screen elements and the time-varying nature of the texture and the temporal spectrum to obtain the prior parameter set.

[0014] S3. Based on the prior parameter set, pre-generate the time domain spectrum of the current mixed video, and use video analysis software to obtain the actual time domain spectrum of the current mixed video. Compare it with the pre-generated spectrum segment by segment to find the high-frequency signal in the time domain, and evaluate the probable cause of the high-frequency signal based on the difference between the two.

[0015] S4. Lightweight preprocessing for selective suppression of non-natural high-frequency signals determined to be unintended, and during the encoding stage, the bit rate is tilted to natural high-frequency signals through bit rate allocation and quantization control to suppress bit rate inflation caused by non-natural high frequencies.

[0016] According to the above technical solution, in S0, the characteristic parameter of the motion trajectory of the screen element is the velocity V of the screen element i as a function of time t. ii acceleration a it , Directional change rate Δθ it trajectory curvature r it The time-varying nature of texture is the dominant frequency F of the texture change of image element i. it Amplitude fluctuation range A it Energy frequency distribution uniformity S it .

[0017] According to the above technical solution, in step S1, finding the closest semantic slot combination in the AI ​​video library specifically involves:

[0018] S1-1. For all semantic slot combinations, input them into the pre-trained sentence vector model to transform the semantic slot combinations in natural language form into high-dimensional real-valued vector representations. Each combination corresponds to a fixed-dimensional vector array v. x ={q1,q2,…,q n}, where n is the number of semantic slots, x is the sequence number of semantic slot combinations in the AI ​​video library, and the values ​​of each dimension in the vector are used to characterize the feature distribution of the semantic slot combination in the semantic space;

[0019] S1-2, For the user-inputted prompt word combination v y Similarly, the sentence vector model transforms the information into a corresponding high-dimensional real-valued vector array. The closer the meanings of the two prompt words are, the closer their values ​​of q are. The cosine similarity formula is then used to calculate the similarity. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the degree of similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

[0020] According to the above technical solution, the specific steps for obtaining the prior parameter set in step S2 are as follows:

[0021] S2-1. For all corresponding videos in the AI ​​video library, calculate the time-domain spectrum using their time series s(t). Where T is the length of the time window, the time domain in which the high-frequency signal appears in the time domain spectrum E(f) is found, and the velocity V of each frame element i is determined. it acceleration a it , Directional change rate Δθ it trajectory curvature r it The dominant frequency F of texture changes it Amplitude fluctuation range A it Energy frequency distribution uniformity S it Perform time-domain truncation;

[0022] S2-2. Using a big data analysis model, analyze the combination of these characteristic parameters in the time domain of the high-frequency signal in the time domain spectrum E(f), find all the characteristic parameter combinations that are highly correlated with the occurrence of the high-frequency signal, and thus obtain the prior parameter set corresponding to the high-frequency signal in the AI ​​video generated by this semantic slot combination.

[0023] According to the above technical solution, in step S3, the evaluation of the probable cause of the high-frequency signal specifically involves:

[0024] S3-1. After the current user's AI montage video is generated, the motion trajectory feature parameters and texture time-varying properties of all screen elements in the AI ​​montage video are analyzed over time to identify the feature parameter combinations that are highly correlated with the occurrence of high-frequency signals. The time domain in which these feature parameter combinations appear is the time domain in which high-frequency signals are likely to occur.

[0025] S3-2. Mark the time domain of the high-frequency signal in the actual time domain spectrum of the current mixed video, and compare it with the time domain of the high-frequency signal inferred from the combination of feature parameters. The overlapping time domain is evaluated as a natural high-frequency signal, and the non-overlapping time domain is evaluated as a non-natural high-frequency signal.

[0026] According to the above technical solution, S4 specifically refers to:

[0027] S4-1. In the time domain of high-frequency signals in AI-edited videos, for high-frequency signals evaluated as unnatural and where E(f) > E0, a first-order low-pass filter is used for exponential moving average. The cutoff frequency E0 of the filter is adjusted to suppress the high-frequency signals. The cos(v) of E0 and the semantic slot combination closest to the user prompt word is used to suppress the high-frequency signals. x ,v y The magnitudes are inversely proportional, cos(v) x ,v y The closer the value is to 1, the more the semantics of the current user prompt word matches the existing semantic slot combinations in the AI ​​video library. The more accurate the prior parameter set is, the more high-frequency signals can be filtered.

[0028] S4-2. More coding bit budget is preferentially allocated to the time domain of natural high-frequency signals, and the coding bit budget allocated to the time domain of non-natural high-frequency signals is reduced. For time domain signals identified as natural high-frequency regions, the quantization step size is reduced, i.e., the number of allocated coding bits is increased, in order to maintain detail reproduction. For time domain signals identified as non-natural high-frequency regions, the quantization step size is increased and their bit allocation is reduced. The total bit budget is kept constant through the Lagrange rate distortion optimization model (RDO).

[0029] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention performs multi-granular semantic analysis on the prompt words of historical AI-generated videos to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, obtains the feature parameters of motion trajectory of screen elements and the time-varying nature of texture under various semantic environments, and then obtains the mapping relationship between the time-varying nature of motion trajectory of screen elements and texture and the temporal spectrum, generating a set of prior parameters that can be interpreted by the compressor, and then predicts the frequency distribution in the temporal domain of the current AI-edited video.

[0030] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to natural high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level. Attached Figure Description

[0031] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0032] Figure 1 This is a schematic diagram of the overall modular structure of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Please see Figure 1 This invention provides a technical solution: a video mixing and editing system based on AIGC, including a semantic mapping module, a time-frequency prediction module, and a dynamic compression module. The semantic mapping module is used to perform semantic parsing on prompt words and construct semantic slots. Based on existing AI-generated video big data, it learns the semantic temporal spectrum mapping and outputs a priori parameter set that can be interpreted by the encoder. The time-frequency prediction module uses the priori parameter set to predict the temporal frequency distribution of the current AI-mixed video and obtains the motion trajectory feature parameters and texture time-varying properties of each frame element. The dynamic compression module is used to identify the high-frequency cause and differentiate the processing when the measured spectrum deviates too much from the prediction, so as to avoid abnormal bitrate expansion.

[0035] The semantic mapping module includes a prompt word acquisition module, a semantic slot construction module, a temporal spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library. The prompt word acquisition module is electrically connected to the semantic slot construction module, the AI ​​video library is electrically connected to the temporal spectrum reading module, the semantic analysis module is electrically connected to both the semantic slot construction module and the temporal spectrum reading module, and the mapping relationship derivation module is electrically connected to the temporal spectrum reading module. The prompt word acquisition module receives prompt words during AI-generated video montage and performs word segmentation, part-of-speech tagging, and dependency extraction. The semantic slot construction module organizes structured semantics according to slots for subject, action, scene, shot, rhythm, and material. The temporal spectrum reading module reads the temporal spectrum of existing AI-generated videos. The semantic analysis module analyzes the relationship between the semantics of prompt words in existing AI-generated videos and the range of motion trajectory feature parameters and the time-varying range of texture in the image elements. The mapping relationship derivation module derives the mapping relationship between the motion trajectory and texture of image elements and the temporal spectrum. The AI ​​video library stores existing AI-generated videos and the prompt words corresponding to each AI-generated video.

[0036] The time-frequency prediction module includes an element extraction module, a video generation module, a priori parameter reading module, and a time-domain spectrum pre-generation module. The priori parameter reading module is electrically connected to the element extraction module and the time-domain spectrum pre-generation module. The element extraction module is used to extract the screen elements of the user-uploaded video and abstract them into organized and structured semantics. The video generation module is used to generate a new video using an AI video generation model based on the prompts input by the user and the extracted screen elements. The priori parameter reading module is used to read the organized and structured semantics of the current AI mixed video, as well as the time-varying nature of the motion trajectory and texture of the screen elements and the mapping relationship with the time-domain spectrum. The time-domain spectrum pre-generation module is used to pre-generate the time-domain spectrum of the current AI mixed video based on the mapping relationship.

[0037] The dynamic compression module includes a temporal spectrum comparison module, a causal probability calculation module, a video compression unit, a frequency band suppression module, and a bitrate allocation module. The temporal spectrum comparison module is used to compare the temporal spectrum of the pre-generated current AI mashup video with the temporal spectrum of the current AI mashup video itself. The causal probability calculation module is used to calculate and analyze the causes of high-frequency components based on the difference between the two temporal spectra. The video compression unit is used to perform video compression. The frequency band suppression module is used to perform frequency band selective suppression preprocessing on unnatural high-frequency signals that are not intended for use. The bitrate allocation module is used to ensure that natural video frames receive a sufficient bitrate during the encoding stage through bitrate control.

[0038] The operation of a video editing system includes the following steps:

[0039] S0. Receive the user's prompt words and uploaded video materials, read the videos and their corresponding prompt words in the existing AI video library, separate the screen elements related to the prompt words, and capture the motion trajectory feature parameters and texture time-varying properties of the screen elements from beginning to end in the video;

[0040] S1. Perform semantic analysis on the prompt words in the AI ​​video library and construct semantic slots for subject, action, scene, shot, rhythm, and material. Each different combination of semantic slots represents a semantic environment. Extract the specified screen elements that need to be used in the video material uploaded by the user, generate prompt words that can represent the specified screen elements, combine the prompt words input by the user to generate the corresponding semantic slot combination, and find the closest semantic slot combination in the AI ​​video library.

[0041] S2. Based on the closest semantic slot combination, use AI to generate videos in batches, analyze all videos corresponding to this semantic environment in the AI ​​video library, obtain the temporal spectrum of all videos, and then obtain the mapping relationship between the motion trajectory feature parameters of the screen elements and the time-varying nature of the texture and the temporal spectrum to obtain the prior parameter set.

[0042] S3. Based on the prior parameter set, pre-generate the time domain spectrum of the current mixed video, and use video analysis software to obtain the actual time domain spectrum of the current mixed video. Compare it with the pre-generated spectrum segment by segment to find the high-frequency signal in the time domain, and evaluate the probable cause of the high-frequency signal based on the difference between the two.

[0043] S4. Lightweight preprocessing for selective suppression of non-natural high-frequency signals determined to be unintended, and in the coding stage, the bit rate is tilted to natural high-frequency signals through bit rate allocation and quantization control to suppress bit rate inflation caused by non-natural high frequencies.

[0044] In S0, the characteristic parameter of the motion trajectory of the screen element is the velocity V of the screen element i as a function of time t. it acceleration a it , Directional change rate Δθ it trajectory curvature r it The time-varying nature of texture is the dominant frequency F of the texture change of image element i. it Amplitude fluctuation range A it Energy frequency distribution uniformity S it ;

[0045] In S1, finding the closest semantic slot combination in the AI ​​video library specifically involves:

[0046] S1-1. For all semantic slot combinations, input them into the pre-trained sentence vector model to transform the semantic slot combinations in natural language form into high-dimensional real-valued vector representations. Each combination corresponds to a fixed-dimensional vector array v. x={q1,q2,…,q n}, where n is the number of semantic slots, x is the sequence number of semantic slot combinations in the AI ​​video library, and the values ​​of each dimension in the vector are used to characterize the feature distribution of the semantic slot combination in the semantic space;

[0047] S1-2, For the user-inputted prompt word combination v y Similarly, the sentence vector model transforms the information into a corresponding high-dimensional real-valued vector array. The closer the meanings of the two prompt words are, the closer their values ​​of q are. The cosine similarity formula is then used to calculate the similarity. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

[0048] In S2, the specific steps to obtain the prior parameter set are as follows:

[0049] S2-1. For all corresponding videos in the AI ​​video library, calculate the time-domain spectrum using their time series s(t). Where T is the length of the time window, the time domain in which the high-frequency signal appears in the time domain spectrum E(f) is found, and the velocity V of each frame element i is determined. it acceleration a it , Directional change rate Δθ it trajectory curvature r it The dominant frequency F of texture changes it Amplitude fluctuation range A it Energy frequency distribution uniformity S it Perform time-domain truncation;

[0050] S2-2. Using a big data analysis model, analyze the combination of these characteristic parameters in the time domain of the high-frequency signal in the time domain spectrum E(f), find all the characteristic parameter combinations that are highly correlated with the occurrence of the high-frequency signal, and thus obtain the prior parameter set corresponding to the high-frequency signal in the AI ​​video generated by this semantic slot combination.

[0051] In S3, the specific methods for evaluating the probable causes of high-frequency signals are as follows:

[0052] S3-1. After the current user's AI montage video is generated, the motion trajectory feature parameters and texture time-varying properties of all screen elements in the AI ​​montage video are analyzed over time to identify the feature parameter combinations that are highly correlated with the occurrence of high-frequency signals. The time domain in which these feature parameter combinations appear is the time domain in which high-frequency signals are likely to occur.

[0053] S3-2. Mark the time domain of the high-frequency signal in the actual time domain spectrum of the current mixed video, and compare it with the time domain of the high-frequency signal inferred based on the combination of feature parameters. The overlapping time domain is evaluated as a natural high-frequency signal, and the non-overlapping time domain is evaluated as a non-natural high-frequency signal.

[0054] S4 specifically refers to:

[0055] S4-1. In the time domain of high-frequency signals in AI-edited videos, for high-frequency signals evaluated as unnatural and where E(f) > E0, a first-order low-pass filter is used for exponential moving average. The cutoff frequency E0 of the filter is adjusted to suppress the high-frequency signals. The cos(v) of E0 and the semantic slot combination closest to the user prompt word is used to suppress the high-frequency signals. x ,v y The magnitudes are inversely proportional, cos(v) x ,v y The closer the value is to 1, the more the semantics of the current user prompt word matches the existing semantic slot combinations in the AI ​​video library. The more accurate the prior parameter set is, the more high-frequency signals can be filtered.

[0056] S4-2. More coding bit budget is preferentially allocated to the time domain of natural high-frequency signals, while less is allocated to the time domain of non-natural high-frequency signals. For time-domain signals identified as natural high-frequency regions, the quantization step size is reduced, i.e., the number of allocated coding bits is increased, to maintain detail reproduction. For time-domain signals identified as non-natural high-frequency regions, the quantization step size is increased, reducing their bit allocation. The total bit budget is maintained constant using a Lagrange rate-distortion optimization (RDO) model. The high-frequency signal time domain with a larger coding bit budget is due to pixel differences between adjacent frames caused by the movement of image elements. A smaller quantization step size in video compression allows for more refined preservation of details and high-frequency textures.

[0057] This invention performs multi-granular semantic analysis on the prompts in historical AI-generated videos to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, deriving the motion trajectory feature parameters of screen elements and the time-varying nature of textures under various semantic environments. Then, it derives the mapping relationship between the time-varying nature of the motion trajectory and texture of screen elements and the temporal spectrum, generating a set of prior parameters that can be interpreted by a compressor, and then predicting the temporal frequency distribution of the current AI-edited video.

[0058] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to natural high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level.

[0059] This invention performs multi-granular semantic analysis on user-input prompts to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, deriving the time-varying nature of motion trajectories and textures of scene elements under various semantic environments. Then, it derives the mapping relationship between the time-varying nature of motion trajectories and textures of scene elements and the temporal spectrum, generating a set of prior parameters that can be interpreted by a compressor. This allows for temporal frequency distribution prediction of current AI-edited videos.

[0060] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level.

[0061] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0062] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A video mixing and editing system based on AIGC, characterized in that: The system includes a semantic mapping module, a time-frequency prediction module, and a dynamic compression module. The semantic mapping module is used to perform semantic parsing on prompt words and construct semantic slots. Based on existing AI-generated video big data, it learns the semantic temporal spectrum mapping and outputs a set of prior parameters that can be interpreted by the encoder. The time-frequency prediction module uses the prior parameter set to predict the temporal frequency distribution of the current AI-edited video and obtains the motion trajectory feature parameters and time-varying texture of each frame element. The dynamic compression module is used to identify the high-frequency cause and differentiate the processing when the measured spectrum deviates too much from the prediction.

2. The video mixing and editing system based on AIGC according to claim 1, characterized in that: The semantic mapping module includes a prompt word acquisition module, a semantic slot construction module, a temporal spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library. The prompt word acquisition module receives prompt words during AI-generated video montage and performs word segmentation, part-of-speech tagging, and dependency extraction. The semantic slot construction module organizes structured semantics according to slots for subject, action, scene, shot, rhythm, and material. The temporal spectrum reading module reads the temporal spectrum of existing AI-generated videos. The semantic analysis module analyzes the relationship between the semantics of prompt words in existing AI-generated videos and the range of motion trajectory feature parameters and time-varying texture of screen elements. The mapping relationship derivation module derives the mapping relationship between the time-varying motion trajectory and texture of screen elements and the temporal spectrum. The AI ​​video library stores existing AI-generated videos and the prompt words corresponding to each AI-generated video. The time-frequency prediction module includes an element extraction module, a video generation module, a priori parameter reading module, and a time-domain spectrum pre-generation module. The element extraction module is used to extract the screen elements of the user-uploaded video and abstract them into organized and structured semantics. The video generation module is used to generate a new video using an AI video generation model based on the prompt words input by the user and the extracted screen elements. The priori parameter reading module is used to read the organized and structured semantics of the current AI mixed video, as well as the mapping relationship between the time-varying motion trajectory and texture of the screen elements and the time-domain spectrum. The time-domain spectrum pre-generation module is used to pre-generate the time-domain spectrum of the current AI mixed video based on the mapping relationship. The dynamic compression module includes a time-domain spectrum comparison module, a causal probability calculation module, a video compression unit, a frequency band suppression module, and a bitrate allocation module. The time-domain spectrum comparison module is used to compare the time-domain spectrum of the pre-generated current AI-edited video with the time-domain spectrum of the current AI-edited video itself. The causal probability calculation module is used to calculate and analyze the causes of high-frequency components based on the difference between the two time-domain spectra. The video compression unit is used to perform video compression. The frequency band suppression module is used to perform frequency band selective suppression preprocessing on unnatural high-frequency signals that are not intended for use. The bitrate allocation module is used to ensure that natural video frames receive a sufficient bitrate during the encoding stage through bitrate control.

3. The video mixing and editing system based on AIGC according to claim 2, characterized in that: The video editing system operates by the following steps: S0. Receive the user's prompt words and uploaded video materials, read the videos and their corresponding prompt words in the existing AI video library, separate the screen elements related to the prompt words, and capture the motion trajectory feature parameters and texture time-varying properties of the screen elements from beginning to end in the video; S1. Perform semantic analysis on the prompt words in the AI ​​video library and construct semantic slots for subject, action, scene, shot, rhythm, and material. Each different combination of semantic slots represents a semantic environment. Extract the specified screen elements that need to be used in the video material uploaded by the user, generate prompt words that can represent the specified screen elements, combine the prompt words input by the user to generate the corresponding semantic slot combination, and find the closest semantic slot combination in the AI ​​video library. S2. Based on the closest semantic slot combination, use AI to generate videos in batches, analyze all videos corresponding to this semantic environment in the AI ​​video library, obtain the temporal spectrum of all videos, and then obtain the mapping relationship between the motion trajectory feature parameters of the screen elements and the time-varying nature of the texture and the temporal spectrum to obtain the prior parameter set. S3. Based on the prior parameter set, pre-generate the time domain spectrum of the current mixed video, and use video analysis software to obtain the actual time domain spectrum of the current mixed video. Compare it with the pre-generated spectrum segment by segment to find the high-frequency signal in the time domain, and evaluate the probable cause of the high-frequency signal based on the difference between the two. S4. Lightweight preprocessing for selective suppression of non-natural high-frequency signals determined to be unintended, and during the encoding stage, the bit rate is tilted to natural high-frequency signals through bit rate allocation and quantization control to suppress bit rate inflation caused by non-natural high frequencies.

4. The video mixing and editing system based on AIGC according to claim 3, characterized in that: In S0, the characteristic parameter of the motion trajectory of the screen element is the velocity V of the screen element i as a function of time t. it acceleration a it , Directional change rate Δθ it trajectory curvature r it The time-varying nature of texture is the dominant frequency F of the texture change of image element i. it Amplitude fluctuation range A it Energy frequency distribution uniformity S it .

5. A video mixing and editing system based on AIGC according to claim 4, characterized in that: In step S1, finding the closest semantic slot combination in the AI ​​video library specifically involves: S1-1. For all semantic slot combinations, input them into the pre-trained sentence vector model to transform the semantic slot combinations in natural language form into high-dimensional real-valued vector representations. Each combination corresponds to a fixed-dimensional vector array v. x ={q1,q2,…,q n }, where n is the number of semantic slots, x is the sequence number of semantic slot combinations in the AI ​​video library, and the values ​​of each dimension in the vector are used to characterize the feature distribution of the semantic slot combination in the semantic space; S1-2, For the user-inputted prompt word combination v y Similarly, the sentence vector model transforms the information into a corresponding high-dimensional real-valued vector array. The closer the meanings of the two prompt words are, the closer their values ​​of q are. The cosine similarity formula is then used to calculate the similarity. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the degree of similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

6. A video mixing and editing system based on AIGC according to claim 5, characterized in that: In S2, the specific steps for obtaining the prior parameter set are as follows: S2-1. For all corresponding videos in the AI ​​video library, calculate the time-domain spectrum using their time series s(t). Where T is the length of the time window, the time domain in which the high-frequency signal appears in the time domain spectrum E(f) is found, and the velocity V of each frame element i is determined. it acceleration a it , Directional change rate Δθ it trajectory curvature r it The dominant frequency F of texture changes it Amplitude fluctuation range A it Energy frequency distribution uniformity S it Perform time-domain truncation; S2-2. Using a big data analysis model, analyze the combination of these characteristic parameters in the time domain of the high-frequency signal in the time domain spectrum E(f), find all the characteristic parameter combinations that are highly correlated with the occurrence of the high-frequency signal, and thus obtain the prior parameter set corresponding to the high-frequency signal in the AI ​​video generated by this semantic slot combination.

7. A video mixing and editing system based on AIGC according to claim 6, characterized in that: In step S3, the evaluation of the probable causes of high-frequency signals specifically involves: S3-1. After the current user's AI montage video is generated, the motion trajectory feature parameters and texture time-varying properties of all screen elements in the AI ​​montage video are analyzed over time to identify the feature parameter combinations that are highly correlated with the occurrence of high-frequency signals. The time domain in which these feature parameter combinations appear is the time domain in which high-frequency signals are likely to occur. S3-2. Mark the time domain of the high-frequency signal in the actual time domain spectrum of the current mixed video, and compare it with the time domain of the high-frequency signal inferred from the combination of feature parameters. The overlapping time domain is evaluated as a natural high-frequency signal, and the non-overlapping time domain is evaluated as a non-natural high-frequency signal.

8. A video mixing and editing system based on AIGC according to claim 7, characterized in that: Specifically, S4 is: S4-1. In the time domain of high-frequency signals in AI-edited videos, for high-frequency signals evaluated as unnatural and where E(f) > E0, a first-order low-pass filter is used for exponential moving average. The cutoff frequency E0 of the filter is adjusted to suppress the high-frequency signals. The cos(v) of E0 and the semantic slot combination closest to the user prompt word is used to suppress the high-frequency signals. x ,v y The size is inversely proportional; S4-2. More coding bit budget is preferentially allocated to the time domain of natural high-frequency signals, and the coding bit budget allocated to the time domain of non-natural high-frequency signals is reduced. For time domain signals identified as natural high-frequency regions, the quantization step size is reduced, i.e., the number of allocated coding bits is increased, in order to maintain detail reproduction. For time domain signals identified as non-natural high-frequency regions, the quantization step size is increased and their bit allocation is reduced. The total bit budget is kept constant through the Lagrange rate distortion optimization model (RDO).

Citation Information

Patent Citations

  • Compression code rate prediction method based on video content and clustering analysis

    CN105959685A

  • Video encoding code rate control method, apparatus and system

    CN106254868A

  • Screen content image compression method of two-stage octave convolution based on multi-scale residual error and window attention

    CN117544783A

  • Frequency spectrum prediction-oriented concept drift detection and self-learning method and device, storage medium and computer program product

    CN118249942A

  • Video compression method based on deep learning

    CN120455721A