An AIGC-based video mixing and splicing system

By constructing semantic slots and learning the semantic temporal spectrum mapping of the AI ​​video library, a priori parameter set is generated. The dynamic compression module performs frequency distribution prediction and differential processing on AIGC mashup videos, solving the bit rate inflation problem caused by non-natural high-frequency signals in AIGC mashup videos and achieving efficient video compression.

CN120956987BActive Publication Date: 2026-03-31XIAMEN XINZHUN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

AIGC-edited videos may contain unnatural high-frequency signals in the time domain spectrum, leading to abnormal bitrate inflation. Existing technologies cannot accurately define the range between natural and unnatural signals, and traditional video compressors cannot effectively handle this, resulting in increased video size and image distortion.

Method used

The semantic mapping module constructs semantic slots, learns semantic temporal spectrum mapping by combining AI video library, generates a prior parameter set, uses time-frequency prediction module to predict motion trajectory and texture time-varying properties of scene elements, and the dynamic compression module performs differential processing on high-frequency signals, selectively suppresses non-natural high-frequency signals, and adjusts bitrate allocation.

Benefits of technology

It effectively avoids bitrate inflation caused by non-natural high-frequency signals, maintains video quality, reduces video size, and ensures compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956987B_ABST
    Figure CN120956987B_ABST
Patent Text Reader

Abstract

The application discloses a video mixing and splicing system based on AIGC, and relates to the technical field of video generation, which comprises a semantic mapping module, a time-frequency prediction module and a dynamic compression module. The semantic mapping module is used for semantic analysis of prompt words and construction of semantic slots. The mapping of the semantic time domain spectrum is learned based on existing AI generated video big data. The prior parameter set that can be interpreted by the encoder is output. The time-frequency prediction module uses the prior parameter set to predict the time domain frequency distribution of the current AI mixed and spliced video, so as to obtain the motion trajectory characteristic parameters of each picture element and the time variation of the texture. The dynamic compression module is used for distinguishing the high-frequency causes and differentiating the processing when the measured spectrum and the prediction deviation are too large, so as to avoid abnormal expansion of the code rate. The semantic mapping module comprises a prompt word collection module, a semantic slot construction module, a time domain spectrum reading module, a semantic analysis module and a mapping relationship derivation module. The application has the characteristics of avoiding abnormal code rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video technology, specifically to a video mixing and editing system based on AIGC. Background Technology

[0002] Video montage refers to cutting, splicing, and recombining two or more video clips to generate a new, complete video. Newer AIGC montage technology can extract specific visual elements from video clips, merge and align them temporally, and then reconstruct them in a new video.

[0003] Compared to naturally shot videos, AIGC (AI Generic Video Editing) is the result of neural networks predicting and synthesizing frame by frame or segment by segment. It does not directly follow the constraints of the physical world and cannot accurately capture kinematic laws. It only guesses the next frame at the pixel level. Unnatural high-frequency signals (jitter, noise) may appear in the temporal spectrum of the edited video. Traditional video compressors will mistakenly believe that the unnatural high-frequency signals mainly come from the pixel differences between adjacent frames caused by the movement of screen elements. If the differences between adjacent frames are huge, more bitrate is needed to save them, resulting in abnormal bitrate expansion and thus a larger video size. If the compressor forcibly reduces the bitrate and suppresses the frequency band of high-frequency signals, the picture will be blurry and distorted.

[0004] Current technologies learn statistical patterns of motion and temporal spectrum of scene elements from large-scale video data. However, AIGC-edited videos often contain many original artistic concepts, which may themselves not conform to statistical laws, making it difficult to accurately define the range between natural and unnatural. Therefore, it is necessary to design an AIGC-based video editing system that avoids abnormal bitrates. Summary of the Invention

[0005] The purpose of this invention is to provide a video mixing and editing system based on AIGC to solve the problems mentioned in the background art.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a video mixing and editing system based on AIGC, comprising a semantic mapping module, a time-frequency prediction module, and a dynamic compression module. The semantic mapping module is used to perform semantic parsing on prompt words and construct semantic slots, learn the semantic temporal spectrum mapping based on existing AI-generated video big data, and output a prior parameter set that can be interpreted by the encoder. The time-frequency prediction module uses the prior parameter set to predict the temporal frequency distribution of the current AI-mixed video, obtaining the motion trajectory feature parameters and time-varying texture of each frame element. The dynamic compression module is used to identify the high-frequency causes and perform differentiated processing when the measured spectrum deviates too much from the prediction.

[0007] According to the above technical solution, the semantic mapping module includes a prompt word acquisition module, a semantic slot construction module, a time-domain spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library. The prompt word acquisition module is electrically connected to the semantic slot construction module, the AI ​​video library is electrically connected to the time-domain spectrum reading module, the semantic analysis module is electrically connected to both the semantic slot construction module and the time-domain spectrum reading module, and the mapping relationship derivation module is electrically connected to the time-domain spectrum reading module. The prompt word acquisition module is used to receive prompt words generated during AI-generated video montage and to complete word segmentation, part-of-speech tagging, and dependency relation determination. The semantic slot construction module is used to organize structured semantics according to slots for subject, action, scene, shot, rhythm, and material. The temporal spectrum reading module is used to read the temporal spectrum of existing AI-generated videos. The semantic analysis module is used to analyze the relationship between the semantics of prompt words in existing AI-generated videos and the range of feature parameters of motion trajectories of screen elements and the range of time-varying textures. The mapping relationship derivation module is used to derive the mapping relationship between the time-varying motion trajectories of screen elements and textures and the temporal spectrum. The AI ​​video library is used to store existing AI-generated videos and the prompt words corresponding to each AI-generated video.

[0008] The time-frequency prediction module includes an element extraction module, a video generation module, a priori parameter reading module, and a time-domain spectrum pre-generation module. The priori parameter reading module is electrically connected to the element extraction module and the time-domain spectrum pre-generation module. The element extraction module is used to extract the screen elements of the user-uploaded video and abstract them into organized structural semantics. The video generation module is used to generate a new video using an AI video generation model based on the prompts input by the user and the extracted screen elements. The priori parameter reading module is used to read the organized structural semantics of the current AI mixed video, as well as the mapping relationship between the time-varying motion trajectory and texture of the screen elements and the time-domain spectrum. The time-domain spectrum pre-generation module is used to pre-generate the time-domain spectrum of the current AI mixed video based on the mapping relationship.

[0009] The dynamic compression module includes a time-domain spectrum comparison module, a causal probability calculation module, a video compression unit, a frequency band suppression module, and a bitrate allocation module. The time-domain spectrum comparison module is used to compare the time-domain spectrum of the pre-generated current AI-edited video with the time-domain spectrum of the current AI-edited video itself. The causal probability calculation module is used to calculate and analyze the causes of high-frequency components based on the difference between the two time-domain spectra. The video compression unit is used to perform video compression. The frequency band suppression module is used to perform frequency band selective suppression preprocessing on unnatural high-frequency signals that are not intended for use. The bitrate allocation module is used to ensure that natural video frames receive a sufficient bitrate during the encoding stage through bitrate control.

[0010] According to the above technical solution, the operation of the video mixing and editing system includes the following steps:

[0011] S0. Receive the user's prompt words and uploaded video materials, read the videos and their corresponding prompt words in the existing AI video library, separate the screen elements related to the prompt words, and capture the motion trajectory feature parameters and texture time-varying properties of the screen elements from beginning to end in the video;

[0012] S1. Perform semantic analysis on the prompt words in the AI ​​video library and construct semantic slots for subject, action, scene, shot, rhythm, and material. Each different combination of semantic slots represents a semantic environment. Extract the specified screen elements that need to be used in the video material uploaded by the user, generate prompt words that can represent the specified screen elements, combine the prompt words input by the user to generate the corresponding semantic slot combination, and find the closest semantic slot combination in the AI ​​video library.

[0013] S2. Based on the closest semantic slot combination, use AI to generate videos in batches, analyze all videos corresponding to this semantic environment in the AI ​​video library, obtain the temporal spectrum of all videos, and then obtain the mapping relationship between the motion trajectory feature parameters of the screen elements and the time-varying nature of the texture and the temporal spectrum to obtain the prior parameter set.

[0014] S3. Based on the prior parameter set, pre-generate the time domain spectrum of the current mixed video, and use video analysis software to obtain the actual time domain spectrum of the current mixed video. Compare it with the pre-generated spectrum segment by segment to find the high-frequency signal in the time domain, and evaluate the probable cause of the high-frequency signal based on the difference between the two.

[0015] S4. Lightweight preprocessing for selective suppression of non-natural high-frequency signals determined to be unintended, and during the encoding stage, the bit rate is tilted to natural high-frequency signals through bit rate allocation and quantization control to suppress bit rate inflation caused by non-natural high frequencies.

[0016] According to the above technical solution, in S0, the motion trajectory feature parameters of the screen elements are the screen elements. over time Speed ​​of movement acceleration , rate of change of direction trajectory curvature The time-varying nature of textures is a visual element The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution .

[0017] According to the above technical solution, in step S1, finding the closest semantic slot combination in the AI ​​video library specifically involves:

[0018] S1-1. For all semantic slot combinations, input them into the pre-trained sentence vector model to transform the semantic slot combinations in natural language form into high-dimensional real-number vector representations. Each combination corresponds to a set of vector arrays with fixed dimensions. ,in This refers to the number of semantic slots. This is the sequence number of the semantic slot combination in the AI ​​video library. The values ​​of each dimension in the vector are used to characterize the feature distribution of the semantic slot combination in the semantic space.

[0019] S1-2, For user-input prompt word combinations Similarly, after being transformed into a corresponding high-dimensional real number vector array through the sentence vector model, the closer the meanings of the two prompt words are, the better. The closer the values ​​are, the better the cosine similarity calculation formula is applied. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

[0020] According to the above technical solution, the specific steps for obtaining the prior parameter set in step S2 are as follows:

[0021] S2-1. For all corresponding videos in the AI ​​video library, analyze their time series... The time-domain spectrum was calculated. ,in Find the time-domain spectrum for the time window length. The time domain in which high-frequency signals appear, and the various image elements. speed acceleration , rate of change of direction trajectory curvature The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution Perform time-domain truncation, where The imaginary unit satisfies ;

[0022] S2-2, Utilizing big data analysis models in the time-domain spectrum In the time domain of mid-to-high frequency signals, these feature parameter combinations are analyzed to identify all feature parameter combinations that are highly correlated with the occurrence of high frequency signals, thereby obtaining the prior parameter set corresponding to the high frequency signals in the AI ​​video generated by this semantic slot combination.

[0023] According to the above technical solution, in step S3, the evaluation of the probable cause of the high-frequency signal specifically involves:

[0024] S3-1. After the current user's AI montage video is generated, the motion trajectory feature parameters and texture time-varying properties of all screen elements in the AI ​​montage video are analyzed over time to identify the feature parameter combinations that are highly correlated with the occurrence of high-frequency signals. The time domain in which these feature parameter combinations appear is the time domain in which high-frequency signals are likely to occur.

[0025] S3-2. Mark the time domain of the high-frequency signal in the actual time domain spectrum of the current mixed video, and compare it with the time domain of the high-frequency signal inferred from the combination of feature parameters. The overlapping time domain is evaluated as a natural high-frequency signal, and the non-overlapping time domain is evaluated as a non-natural high-frequency signal.

[0026] According to the above technical solution, S4 specifically refers to:

[0027] S4-1. In the high-frequency signal time domain of AI-edited videos, those evaluated as unnatural and... The high-frequency signal is subjected to exponential moving average using a first-order low-pass filter, and the cutoff frequency of the filter is adjusted. High-frequency signals are suppressed, among which Combined with the semantic slot closest to the user suggestion word Size is inversely proportional. The closer it is to 1, the more the semantics of the current user prompt word matches the existing semantic slot combinations in the AI ​​video library, the more accurate the prior parameter set is, and the more high-frequency signals can be filtered.

[0028] S4-2. More coding bit budget is preferentially allocated to the time domain of natural high-frequency signals, and the coding bit budget allocated to the time domain of non-natural high-frequency signals is reduced. For time domain signals identified as natural high-frequency regions, the quantization step size is reduced, i.e., the number of allocated coding bits is increased, in order to maintain detail reproduction. For time domain signals identified as non-natural high-frequency regions, the quantization step size is increased and their bit allocation is reduced. The total bit budget is kept constant through the Lagrange rate distortion optimization model (RDO).

[0029] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention performs multi-granular semantic analysis on the prompt words of historical AI-generated videos to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, obtains the feature parameters of motion trajectory of screen elements and the time-varying nature of texture under various semantic environments, and then obtains the mapping relationship between the time-varying nature of motion trajectory of screen elements and texture and the temporal spectrum, generating a set of prior parameters that can be interpreted by the compressor, and then predicts the frequency distribution in the temporal domain of the current AI-edited video.

[0030] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to natural high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level. Attached Figure Description

[0031] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0032] Figure 1 This is a schematic diagram of the overall modular structure of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] Please see Figure 1 The present invention provides a technical solution: a video mixing and editing system based on AIGC, including a semantic mapping module, a time-frequency prediction module, and a dynamic compression module. The semantic mapping module is used to perform semantic parsing on prompt words and construct semantic slots. Based on existing AI-generated video big data, it learns the semantic time-domain spectrum mapping and outputs a priori parameter set that can be interpreted by the encoder. The time-frequency prediction module uses the priori parameter set to predict the time-domain frequency distribution of the current AI-mixed video and obtains the motion trajectory feature parameters and time-varying texture of each frame element. The dynamic compression module is used to identify the high-frequency cause and differentiate the processing when the measured spectrum deviates too much from the prediction.

[0035] The semantic mapping module includes a prompt word acquisition module, a semantic slot construction module, a temporal spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library. The prompt word acquisition module is electrically connected to the semantic slot construction module, the AI ​​video library is electrically connected to the temporal spectrum reading module, the semantic analysis module is electrically connected to both the semantic slot construction module and the temporal spectrum reading module, and the mapping relationship derivation module is electrically connected to the temporal spectrum reading module. The prompt word acquisition module receives prompt words during AI-generated video montage and performs word segmentation, part-of-speech tagging, and dependency extraction. The semantic slot construction module organizes structured semantics according to slots for subject, action, scene, shot, rhythm, and material. The temporal spectrum reading module reads the temporal spectrum of existing AI-generated videos. The semantic analysis module analyzes the relationship between the semantics of prompt words in existing AI-generated videos and the range of motion trajectory feature parameters and the time-varying range of texture in the image elements. The mapping relationship derivation module derives the mapping relationship between the motion trajectory and texture of image elements and the temporal spectrum. The AI ​​video library stores existing AI-generated videos and the prompt words corresponding to each AI-generated video.

[0036] The time-frequency prediction module includes an element extraction module, a video generation module, a priori parameter reading module, and a time-domain spectrum pre-generation module. The priori parameter reading module is electrically connected to the element extraction module and the time-domain spectrum pre-generation module. The element extraction module is used to extract the screen elements of the user-uploaded video and abstract them into organized and structured semantics. The video generation module is used to generate a new video using an AI video generation model based on the prompts input by the user and the extracted screen elements. The priori parameter reading module is used to read the organized and structured semantics of the current AI mixed video, as well as the time-varying nature of the motion trajectory and texture of the screen elements and the mapping relationship with the time-domain spectrum. The time-domain spectrum pre-generation module is used to pre-generate the time-domain spectrum of the current AI mixed video based on the mapping relationship.

[0037] The dynamic compression module includes a temporal spectrum comparison module, a causal probability calculation module, a video compression unit, a frequency band suppression module, and a bitrate allocation module. The temporal spectrum comparison module is used to compare the temporal spectrum of the pre-generated current AI mashup video with the temporal spectrum of the current AI mashup video itself. The causal probability calculation module is used to calculate and analyze the causes of high-frequency components based on the difference between the two temporal spectra. The video compression unit is used to perform video compression. The frequency band suppression module is used to perform frequency band selective suppression preprocessing on unnatural high-frequency signals that are not intended for use. The bitrate allocation module is used to ensure that natural video frames receive a sufficient bitrate during the encoding stage through bitrate control.

[0038] The operation of a video editing system includes the following steps:

[0039] S0. Receive the user's prompt words and uploaded video materials, read the videos and their corresponding prompt words in the existing AI video library, separate the screen elements related to the prompt words, and capture the motion trajectory feature parameters and texture time-varying properties of the screen elements from beginning to end in the video;

[0040] S1. Perform semantic analysis on the prompt words in the AI ​​video library and construct semantic slots for subject, action, scene, shot, rhythm, and material. Each different combination of semantic slots represents a semantic environment. Extract the specified screen elements that need to be used in the video material uploaded by the user, generate prompt words that can represent the specified screen elements, combine the prompt words input by the user to generate the corresponding semantic slot combination, and find the closest semantic slot combination in the AI ​​video library.

[0041] S2. Based on the closest semantic slot combination, use AI to generate videos in batches, analyze all videos corresponding to this semantic environment in the AI ​​video library, obtain the temporal spectrum of all videos, and then obtain the mapping relationship between the motion trajectory feature parameters of the screen elements and the time-varying nature of the texture and the temporal spectrum to obtain the prior parameter set.

[0042] S3. Based on the prior parameter set, pre-generate the time domain spectrum of the current mixed video, and use video analysis software to obtain the actual time domain spectrum of the current mixed video. Compare it with the pre-generated spectrum segment by segment to find the high-frequency signal in the time domain, and evaluate the probable cause of the high-frequency signal based on the difference between the two.

[0043] S4. Lightweight preprocessing for selective suppression of non-natural high-frequency signals determined to be unintended, and in the coding stage, the bit rate is tilted to natural high-frequency signals through bit rate allocation and quantization control to suppress bit rate inflation caused by non-natural high frequencies.

[0044] In S0, the motion trajectory characteristic parameters of the screen elements are the screen elements. over time Speed ​​of movement acceleration , rate of change of direction trajectory curvature The time-varying nature of textures is a visual element The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution ;

[0045] In S1, finding the closest semantic slot combination in the AI ​​video library specifically involves:

[0046] S1-1. For all semantic slot combinations, input them into the pre-trained sentence vector model to transform the semantic slot combinations in natural language form into high-dimensional real-number vector representations. Each combination corresponds to a set of vector arrays with fixed dimensions. ,in This refers to the number of semantic slots. This is the sequence number of the semantic slot combination in the AI ​​video library. The values ​​of each dimension in the vector are used to characterize the feature distribution of the semantic slot combination in the semantic space.

[0047] S1-2, For user-input prompt word combinations Similarly, after being transformed into a corresponding high-dimensional real number vector array through the sentence vector model, the closer the meanings of the two prompt words are, the better. The closer the values ​​are, the better the cosine similarity calculation formula is applied. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

[0048] In S2, the specific steps to obtain the prior parameter set are as follows:

[0049] S2-1. For all corresponding videos in the AI ​​video library, analyze their time series... The time-domain spectrum was calculated. ,in Find the time-domain spectrum for the time window length. The time domain in which high-frequency signals appear, and the various image elements. speed acceleration , rate of change of direction trajectory curvature The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution Perform time-domain truncation, where The imaginary unit satisfies ;

[0050] S2-2, Utilizing big data analysis models in the time-domain spectrum In the time domain of mid-to-high frequency signals, these feature parameter combinations are analyzed to identify all feature parameter combinations that are highly correlated with the occurrence of high frequency signals, thereby obtaining the prior parameter set corresponding to the high frequency signals in the AI ​​video generated by this semantic slot combination.

[0051] In S3, the specific methods for evaluating the probable causes of high-frequency signals are as follows:

[0052] S3-1. After the current user's AI montage video is generated, the motion trajectory feature parameters and texture time-varying properties of all screen elements in the AI ​​montage video are analyzed over time to identify the feature parameter combinations that are highly correlated with the occurrence of high-frequency signals. The time domain in which these feature parameter combinations appear is the time domain in which high-frequency signals are likely to occur.

[0053] S3-2. Mark the time domain of the high-frequency signal in the actual time domain spectrum of the current mixed video, and compare it with the time domain of the high-frequency signal inferred based on the combination of feature parameters. The overlapping time domain is evaluated as a natural high-frequency signal, and the non-overlapping time domain is evaluated as a non-natural high-frequency signal.

[0054] S4 specifically refers to:

[0055] S4-1. In the high-frequency signal time domain of AI-edited videos, those evaluated as unnatural and... The high-frequency signal is subjected to exponential moving average using a first-order low-pass filter, and the cutoff frequency of the filter is adjusted. High-frequency signals are suppressed, among which Combined with the semantic slot closest to the user suggestion word Size is inversely proportional. The closer it is to 1, the more the semantics of the current user prompt word matches the existing semantic slot combinations in the AI ​​video library, the more accurate the prior parameter set is, and the more high-frequency signals can be filtered.

[0056] S4-2. More coding bit budget is preferentially allocated to the time domain of natural high-frequency signals, while less is allocated to the time domain of non-natural high-frequency signals. For time-domain signals identified as natural high-frequency regions, the quantization step size is reduced, i.e., the number of allocated coding bits is increased, to maintain detail reproduction. For time-domain signals identified as non-natural high-frequency regions, the quantization step size is increased, reducing their bit allocation. The total bit budget is maintained constant using a Lagrange rate-distortion optimization (RDO) model. The high-frequency signal time domain with a larger coding bit budget is due to pixel differences between adjacent frames caused by the movement of image elements. A smaller quantization step size in video compression allows for more refined preservation of details and high-frequency textures.

[0057] This invention performs multi-granular semantic analysis on the prompts in historical AI-generated videos to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, deriving the motion trajectory feature parameters of screen elements and the time-varying nature of textures under various semantic environments. Then, it derives the mapping relationship between the time-varying nature of the motion trajectory and texture of screen elements and the temporal spectrum, generating a set of prior parameters that can be interpreted by a compressor, and then predicting the temporal frequency distribution of the current AI-edited video.

[0058] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to natural high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level.

[0059] This invention performs multi-granular semantic analysis on user-input prompts to construct multiple semantic slots. Based on the combination of words in these semantic slots, it uses big data to learn semantics from existing AI-generated videos, deriving the time-varying nature of motion trajectories and textures of scene elements under various semantic environments. Then, it derives the mapping relationship between the time-varying nature of motion trajectories and textures of scene elements and the temporal spectrum, generating a set of prior parameters that can be interpreted by a compressor. This allows for temporal frequency distribution prediction of current AI-edited videos.

[0060] When encountering high-frequency signals that deviate significantly from the prediction, the system can more accurately analyze the causes of high-frequency components, adopt different compression strategies for high-frequency signals with different causes, selectively suppress unnatural high-frequency signals beyond the intended purpose, and apply bitrate tilting to high-frequency signals within the intended purpose. This avoids abnormal bitrate expansion caused by unnatural high-frequency signals and minimizes video size while ensuring that the overall compression bitrate remains at a high level.

[0061] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0062] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An AIGC-based video mixing and splicing system, characterized in that: The semantic mapping module, the time-frequency prediction module, and the dynamic compression module are included, the semantic mapping module is used for semantic analysis of prompt words and construction of semantic slots, mapping of a semantic time domain spectrum is learned based on existing AI generated video big data, a set of prior parameters that can be interpreted by an encoder is output, the time-frequency prediction module uses the set of prior parameters to predict time domain frequency distribution of a current AI mixed video, motion trajectory feature parameters and texture time variability of each picture element are obtained, and the dynamic compression module is used for determining high frequency causes and differential processing when a measured spectrum and a prediction deviation are too large; The high probability cause of the high frequency signal is specifically: S3-1, after the current user's AI mixed video is generated, the motion trajectory feature parameters and the texture time variability of all picture elements in the AI mixed video are analyzed with respect to changes over time, a feature parameter combination with high correlation to the occurrence of the high frequency signal is found out, and a time domain in which the feature parameter combination appears is a time domain in which the high frequency signal is likely to occur; S3-2, the high frequency signal time domain in the actual time domain spectrum of the current mixed video is marked, and is compared with the high frequency signal time domain deduced according to the feature parameter combination, a time domain that coincides is evaluated as a natural high frequency signal, and a time domain that does not coincide is evaluated as a non-natural high frequency signal; S4-1. In the high-frequency signal time domain of AI-edited videos, those evaluated as unnatural and... The high-frequency signal is subjected to exponential moving average using a first-order low-pass filter, and the cutoff frequency of the filter is adjusted. To suppress high-frequency signals, among which Combined with the semantic slot closest to the user suggestion word Size is inversely proportional; S4-2, more encoding bit budgets are preferentially allocated to the natural high frequency signal time domain, and the encoding bit budgets allocated to the non-natural high frequency signal time domain are reduced, the quantization step of the time domain signal determined as the natural high frequency region is reduced, that is, the number of allocated encoding bits is increased, to maintain the degree of detail restoration, and the quantization step of the time domain signal determined as the non-natural high frequency region is increased, and the bit allocation is reduced, and the total bit budget is maintained constant through a Lagrange rate distortion optimization model (RDO). 2.The AIGC-based video mixing and cutting system according to claim 1, characterized in that: The semantic mapping module includes a prompt word collection module, a semantic slot construction module, a time domain spectrum reading module, a semantic analysis module, a mapping relationship derivation module, and an AI video library, the prompt word collection module is used for receiving prompt words during AI generated video mixing and cutting, and completing word segmentation, part-of-speech tagging, and dependency relationship extraction, the semantic slot construction module is used for organizing structured semantics according to slot structures of subjects, actions, scenes, shots, rhythms, and materials, the time domain spectrum reading module is used for reading time domain spectra of existing AI generated videos, the semantic analysis module is used for analyzing relationships between prompt word semantics of existing AI generated videos and ranges of picture element motion trajectory feature parameters and texture time variability, the mapping relationship derivation module is used for deriving mapping relationships between picture element motion trajectories and texture time variability and time domain spectra, and the AI video library is used for storing existing AI generated videos and prompt words corresponding to each AI generated video. The time-frequency prediction module comprises an element extraction module, a video generation module, a prior parameter reading module, and a time-domain frequency spectrum pre-generation module. The element extraction module is configured to extract picture elements of a user-uploaded video and abstract the picture elements into organized structured semantics. The video generation module is configured to generate a new video by using an AI video generation model according to a prompt word input by a user and the extracted picture elements. The prior parameter reading module is configured to read organized structured semantics of a current AI mixed video, and a mapping relationship between picture element motion track features and texture time variability and time-domain frequency spectrum. The time-domain frequency spectrum pre-generation module is configured to pre-generate a time-domain frequency spectrum of the current AI mixed video according to the mapping relationship. The dynamic compression module comprises a time-domain frequency spectrum comparison module, a cause probability calculation module, a video compression unit, a frequency band suppression module, and a code rate allocation module. The time-domain frequency spectrum comparison module is configured to compare the pre-generated time-domain frequency spectrum of the current AI mixed video with a time-domain frequency spectrum of the current AI mixed video itself. The cause probability calculation module is configured to calculate and analyze causes of high-frequency components according to differences between the two time-domain frequency spectrums. The video compression unit is configured to perform video compression. The frequency band suppression module is configured to perform a frequency band selective suppression preprocessing on unnatural high-frequency signals outside the intention. The code rate allocation module is configured to tilt code rates for natural high-frequency signals in an encoding stage by code rate control. 3.The AIGC-based video mixing and cutting system according to claim 2, characterized in that: The working process of the video mixing system comprises the following steps: S0, receiving a prompt word and uploaded video materials of a user, reading videos in an existing AI video library and corresponding prompt words of the videos, separating picture elements related to the prompt words, and capturing motion track feature parameters and texture time variability of the picture elements from beginning to end in the videos; S1, performing semantic analysis on prompt words in the AI video library and constructing semantic slot positions of subjects, actions, scenes, shots, rhythms, and materials. Each different combination of semantic slot positions represents a semantic environment. Extracting specified picture elements needed in the uploaded video materials of the user, generating prompt words representing the specified picture elements, generating corresponding semantic slot position combinations of the prompt words input by the user, and finding the closest semantic slot position combinations in the AI video library; S2, generating videos by using AI according to the closest semantic slot position combinations, analyzing all videos corresponding to the semantic environment in the AI video library, obtaining time-domain frequency spectrums of all the videos, obtaining mapping relationships between picture element motion track feature parameters and texture time variability and time-domain frequency spectrums, and obtaining a prior parameter set; S3, pre-generating a time-domain frequency spectrum of a current mixed video according to the prior parameter set, obtaining an actual time-domain frequency spectrum of the current mixed video by using video analysis software, comparing the actual time-domain frequency spectrum with the pre-generated frequency spectrum in segments, finding high-frequency signals appearing in the time domain, and evaluating probable causes of the high-frequency signals according to differences between the two; S4, performing a frequency band selective suppression preprocessing on unnatural high-frequency signals outside the intention, and tilting code rates for natural high-frequency signals in an encoding stage by code rate allocation and quantization control to suppress code rate expansion caused by unnatural high frequencies.

4. The AIGC-based video mixing and cutting system according to claim 3, characterized in that: In S0, the motion trajectory feature parameters of the screen elements are the screen elements over time Speed ​​of movement acceleration , rate of change of direction trajectory curvature The time-varying nature of textures is a visual element The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution .

5. The AIGC-based video mixing and cutting system according to claim 4, characterized in that: In the S1, the closest semantic slot combination in the AI video library is found, and the specific process is as follows: S1-1, inputting all semantic slot combinations into a pre-trained sentence vector model respectively, converting the semantic slot combinations in natural language form into high-dimensional real number vector representations, each combination corresponding to a fixed-dimensional vector array wherein is the number of semantic slots, is the semantic slot combination sequence number in the AI video library, and the dimension value in the vector is used to depict the feature distribution of the semantic slot combination in the semantic space. S1-2, For user-input prompt word combinations Similarly, after being transformed into a corresponding high-dimensional real number vector array through the sentence vector model, the closer the meanings of the two prompt words are, the better. The closer the values ​​are, the better the cosine similarity calculation formula is applied. The closer the calculation result is to 1, the more similar they are. The cosine value of the angle between the user's semantic vector and each semantic vector in the library is calculated to quantify the similarity between the two at the semantic level. The semantic slot combinations in the AI ​​video library are sorted according to the similarity score, and the semantic slot combination that is closest to the user's input is selected to find the corresponding video.

6. The AIGC-based video mixing and cutting system according to claim 5, characterized in that: In the S2, the specific steps for obtaining the prior parameter set are as follows: S2-1. For all corresponding videos in the AI ​​video library, analyze their time series... The time-domain spectrum was calculated. ,in Find the time-domain spectrum for the time window length. The time domain in which high-frequency signals appear, and the various image elements. speed acceleration , rate of change of direction trajectory curvature The main frequency of texture changes Amplitude fluctuation range Uniformity of energy frequency distribution Perform time-domain truncation, where The imaginary unit satisfies ; S2-2, using big data analysis model in time domain frequency spectrum In the time domain of the medium and high frequency signal, analyze these feature parameter combinations, find all feature parameter combinations with high correlation with the occurrence of the high frequency signal, and thus obtain the prior parameter set corresponding to the high frequency signal in the AI video generated by this semantic slot combination.

Citation Information

Patent Citations

  • Compression code rate prediction method based on video content and clustering analysis

    CN105959685A

  • Video encoding code rate control method, apparatus and system

    CN106254868A