Video production special effect display device and method for enhancing and user interactivity

By analyzing video clips and special effects frame data, the problem of insufficient intelligence in video special effects technology has been solved, enabling efficient and flexible special effects design and user interactivity, and quickly responding to diverse creative needs.

CN122293908APending Publication Date: 2026-06-26WENZHOU POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing video special effects technologies lack sufficient intelligence, with most relying on manual adjustment of effect parameters and frame sequences. This is time-consuming and labor-intensive, making it difficult to meet the high-efficiency production needs of large-scale video content. The lack of automated semantic understanding and dynamic adaptation capabilities results in a disconnect between special effects and video content semantics, making it unable to quickly respond to diverse creative needs and scene changes.

Method used

By acquiring special effects requirements and raw video data, the video data is parsed into multiple video segments, the special effects frame data of each segment is determined, and frames that do not meet the requirements are removed to generate a target video special effects display scheme, thereby improving processing efficiency and flexibility and quickly responding to creative needs and scene changes.

Benefits of technology

It enables refined processing of video special effects, improves processing efficiency, reduces resource consumption, enhances the flexibility of special effects design and user engagement, and quickly responds to diverse creative needs and scene changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122293908A_ABST
    Figure CN122293908A_ABST
Patent Text Reader

Abstract

This application relates to the field of video special effects technology, and particularly to a video production special effects display device and method that enhances and improves user interactivity. The method includes: acquiring special effects requirement information and original video data to provide prerequisites for subsequent analysis; parsing the original video data according to the special effects requirement information to determine multiple video segment data corresponding to the original video data; determining a first special effects frame data corresponding to each video segment data based on the multiple video segment data corresponding to the original video data; removing the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data to determine a second special effects frame data corresponding to each video segment data; and generating a target video special effects display scheme based on each video segment data and the first and second special effects frame data corresponding to each video segment data, quickly responding to diverse creative needs and scene changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of video special effects technology, and in particular relates to a video production special effects display device and method that enhances and improves user interactivity. Background Technology

[0002] Video special effects background technology refers to various technical means and methods used in video production, editing, and content creation to achieve goals such as visual enhancement, scene reconstruction, and element compositing. Through the integration of technologies from computer graphics, image processing, and artificial intelligence, it endows videos with richer expressiveness, stronger visual impact, and more immersive narrative capabilities.

[0003] Existing video special effects technologies lack sufficient intelligence, with most relying on manual adjustment of effect parameters and frame sequences. This approach is time-consuming and labor-intensive, making it difficult to meet the high-efficiency production needs of large-scale video content. The lack of automated semantic understanding and dynamic adaptation capabilities leads to a disconnect between special effects and video content semantics, and special effects production cannot quickly respond to diverse creative needs and scene changes. Summary of the Invention

[0004] This application provides a video production special effects display device and method that enhances user interactivity. It can solve the technical problems of existing video special effects technology lacking automatic semantic understanding and dynamic adaptation capabilities, resulting in a semantic disconnect between special effects and video content, and special effects production being unable to quickly respond to diverse creative needs and scene changes.

[0005] In a first aspect, embodiments of this application provide a method for enhancing and improving the display of video production effects, including: Obtain information on special effects requirements and raw video data; The original video data is parsed based on the special effects requirements information to determine multiple video segment data corresponding to the original video data; Based on the multiple video segment data corresponding to the original video data, determine the first special effects frame data corresponding to each video segment data; Remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and determine the second special effects frame data corresponding to each video segment data; Based on each video segment data and the corresponding first and second special effects frame data, a target video special effects display scheme is generated.

[0006] The technical solutions described in this application embodiment have at least the following technical effects: The video production and special effects display method for enhanced user interactivity provided in this application provides prerequisites for subsequent semantic analysis and special effects adaptation by acquiring special effects requirement information and original video data. The original video data is parsed based on the special effects requirement information to determine multiple video segment data corresponding to the original video data, achieving refined processing, improving processing efficiency, and reducing resource consumption. Based on the multiple video segment data corresponding to the original video data, the first special effects frame data corresponding to each video segment data is determined, enhancing the flexibility and adjustability of special effects design. Special effects frame data corresponding to the special effects requirement information is removed from the first special effects frame data corresponding to each video segment data, determining the second special effects frame data corresponding to each video segment data, achieving reverse optimization under changing requirements. Based on each video segment data and the corresponding first and second special effects frame data, a target video special effects display scheme is generated, improving user participation and satisfaction with the final scheme, and quickly responding to diverse creative needs and scene changes.

[0007] Secondly, embodiments of this application provide a video production effects display device that enhances and improves user interactivity, including: The acquisition unit is used to acquire special effects requirement information and raw video data; The parsing unit is used to parse the original video data according to the special effects requirement information and determine multiple video segment data corresponding to the original video data; A segment unit is used to determine the first special effects frame data corresponding to each video segment data based on multiple video segment data corresponding to the original video data; The removal unit is used to remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and to determine the second special effects frame data corresponding to each video segment data; The generation unit is used to generate a target video effects display scheme based on each video segment data and the first special effects frame data and the second special effects frame data corresponding to each video segment data.

[0008] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any of the foregoing aspects.

[0009] Fourthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to perform the method described in any of the above aspects.

[0010] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the above aspects, and will not be repeated here. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating a method for enhancing video production effects display and user interactivity according to an embodiment of this application; Figure 2 This is a schematic diagram illustrating the operation of a video production effects display method that enhances user interactivity according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a video production effects display device that enhances user interactivity according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determination" or "if the described condition or event is detected" may be interpreted, depending on the context, as "once determination," "in response to determination," "once the described condition or event is detected," or "in response to the detection of the described condition or event."

[0017] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0018] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0019] Existing video special effects technologies lack sufficient intelligence, with most relying on manual adjustment of effect parameters and frame sequences. This approach is time-consuming and labor-intensive, making it difficult to meet the high-efficiency production needs of large-scale video content. The lack of automated semantic understanding and dynamic adaptation capabilities leads to a disconnect between special effects and video content semantics, low visual integration, and the generation of multiple versions of solutions relies on manual trial and error, resulting in low decision-making efficiency and an inability to quickly respond to diverse creative needs and scene changes.

[0020] To address the aforementioned issues, this application provides a video production special effects display device and method that enhances user interactivity. In this method, special effects requirement information and original video data are acquired to provide prerequisites for subsequent semantic analysis and special effects adaptation. The original video data is parsed based on the special effects requirement information to determine multiple video segment data corresponding to the original video data, achieving refined processing, improving processing efficiency, and reducing resource consumption. Based on the multiple video segment data corresponding to the original video data, a first special effects frame data corresponding to each video segment data is determined, enhancing the flexibility and adjustability of special effects design. Special effects frame data corresponding to the special effects requirement information is removed from the first special effects frame data corresponding to each video segment data, determining a second special effects frame data corresponding to each video segment data, enabling reverse optimization in scenarios of changing requirements. Based on each video segment data and its corresponding first and second special effects frame data, a target video special effects display scheme is generated, improving user engagement and satisfaction with the final scheme, and quickly responding to diverse creative needs and scene changes.

[0021] The video production effects display method for enhancing user interactivity provided in this application embodiment can be applied to electronic devices. In this case, the electronic device is the executing entity of the video production effects display method for enhancing user interactivity provided in this application embodiment. This application embodiment does not impose any restrictions on the specific type of electronic device.

[0022] For example, electronic devices can be mobile phones, tablets, laptops, ultra-mobile personal computers (UMPCs), netbooks, desktop computers, smart screens, smart TVs, handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, vehicle networking terminals, computers, laptops, communication devices, computing devices, etc.

[0023] To better understand the video production effects display method with enhanced user interactivity provided in the embodiments of this application, the specific implementation process of the video production effects display method with enhanced user interactivity provided in the embodiments of this application will be described by way of example below.

[0024] Figure 1 This illustration shows a schematic flowchart of a video production effects display method that enhances user interactivity according to an embodiment of this application. Figure 2 The diagram illustrates the operation flowchart of the video production effects display method for enhancing user interactivity provided in this application embodiment. The method for enhancing user interactivity in video production effects display includes: S100, acquires special effects requirements and raw video data.

[0025] It's understandable that special effects requirements and raw video data can be obtained through user input or interfacing with relevant channels. Special effects requirements encompass the user's specific expectations for the effects, such as effect type (e.g., dynamic stickers, lighting effects, particle effects), application scenarios (e.g., character close-ups, scene transitions, object movement paths), style and tone (e.g., realistic, cartoon, cyberpunk), timeline planning (start and end times of the effects), and details of the effect elements (e.g., specific color parameters, animation trajectories, text content). Raw video data consists of unprocessed source material, which can come from camera footage, imported video files, or third-party platforms. It includes continuous video frame sequences, audio tracks, and metadata (e.g., resolution, frame rate, encoding format). Special effects requirements and raw video data can be obtained through API calls, file uploads, or database reads, providing fundamental resources for subsequent effects analysis and processing, ensuring that subsequent processes revolve around clearly defined requirements and materials.

[0026] S200 parses the original video data based on the special effects requirements information to determine multiple video segment data corresponding to the original video data.

[0027] Understandably, a complete original video can be broken down into multiple independent video segments based on special effects requirements. By analyzing key information such as time points, scene changes, or object motion trajectories in the special effects requirements, the specific location in the original video where the special effects need to be applied can be located. For example, if the requirement is "to add a flame effect when the main character waves," then image recognition or frame analysis technology is needed to identify the start and end frames of the main character's wave, and extract the continuous frame sequence within that time period as an independent segment. If there are multiple different types of special effects requirements (such as needing transition effects and text animations simultaneously), then the video interval corresponding to each special effect can be located separately, and the original video can be cut into multiple segments. Each segment corresponds to the application scope of one or a group of related special effects, which facilitates targeted special effects processing for different segments, improving processing efficiency and accuracy.

[0028] In one possible implementation, S200 parses the original video data based on the special effects requirement information to determine multiple video segment data corresponding to the original video data, including: S210, perform special effects semantic classification based on special effects requirement information to obtain multiple special effects semantic information corresponding to the special effects requirement information.

[0029] It's understandable that special effects requests described in natural language can be transformed into a recognizable semantic tagging system. For example, the request "add a gradient filter and overlay a flying bird animation in a sunset scene" can be broken down into semantic units such as "scene - sunset," "effect type - filter," "filter attribute - gradient," and "effect element - flying bird animation." By extracting keywords using Natural Language Processing (NLP) technology and combining them with pre-defined semantic classification rules (such as classification by effect type, application object, visual effect, etc.), each request item is mapped to a specific semantic tag, forming a structured set of special effects semantic information. This semantic information not only facilitates understanding user needs but also provides clear guidance for subsequent frame sequence parsing, effect matching, and parameter configuration, ensuring that each special effects request can be accurately identified and processed.

[0030] S220: Based on multiple special effects semantic information, the original video data is parsed to determine the independent video frame sequence corresponding to each special effects semantic information.

[0031] It's understandable that semantic tags can be matched with original video frames to locate specific frame sequences. For example, for the semantic tag "person-running-motion blur," a human pose recognition algorithm scans the original video frames to identify the running motion range of the person, extracting consecutive frames within that range as the corresponding sequence. For the tag "scene-night-light spot effect," image color analysis and scene classification models are used to filter out nighttime scene frames with brightness below a threshold and a cool color tone. Alternatively, computer vision algorithms can be used to perform feature matching between each semantic tag and the video content, ensuring that each semantic piece of information corresponds to a real, suitable frame sequence in the original video, providing accurate processing targets for subsequent frame interpolation and effects generation.

[0032] S230: Based on the independent video frame sequence corresponding to each special effect semantic information, multiple video segment data corresponding to the original video data are obtained.

[0033] It's understandable that each independent video frame sequence can be encapsulated into an independent video segment unit. Each frame sequence is continuous on the timeline and contains complete contextual information (such as the motion trends of preceding and following frames, scene transitions), facilitating the maintenance of visual continuity during special effects processing. For example, the frame sequence of "character jumping" can be cut from the original video and saved as a segment, while simultaneously recording metadata such as the segment's time position, frame rate, and resolution within the original video. If multiple semantic tags correspond to frame sequences within the same time period (such as "character-jumping" and "special effects-slow motion"), they are merged into a composite segment, ensuring synchronized processing when multiple special effects are overlaid. Through fragmentation, the original video is transformed into multiple independently operable units, laying the foundation for subsequent staged and categorized special effects processing.

[0034] Optionally, in step S230, based on the independent video frame sequence corresponding to each special effect semantic information, multiple video segment data corresponding to the original video data are obtained, including: S231, Based on the independent video frame sequence corresponding to each special effect semantic information, determine the blank frame data corresponding to each independent video frame sequence.

[0035] It's understandable that the position and number of blank frames required for inserting special effects can be calculated. Blank frames are used to create transition space between special effects and the original video, avoiding abrupt visuals caused by direct overlay of effects. For example, if a particle explosion effect is planned to be inserted into a frame, several blank frames need to be inserted before and after that frame for fade-in and fade-out processing of the effect. By analyzing the motion speed and effect type of the frame sequence, the insertion rules for blank frames are determined: high-speed motion scenes may require more blank frames to ensure smooth effects, while static scenes can have fewer. Blank frame data includes the insertion position (e.g., the 2nd frame before the effect insertion frame, the 3rd frame after the effect insertion frame) and the number (e.g., a total of 5 blank frames inserted), providing specific parameters for subsequent frame interpolation operations to ensure seamless integration of special effects with the original video.

[0036] For example, S231, based on the independent video frame sequence corresponding to each special effect semantic information, determines the blank frame data corresponding to each independent video frame sequence, including: S2311, each independent video frame sequence is input into a preset frame semantic model to obtain special effects insertion frames for each independent video frame sequence; wherein, the frame semantic model is a machine learning model pre-trained.

[0037] Understandably, AI models can be used to automatically identify the best frames for inserting special effects. Frame semantic models (such as Convolutional Neural Networks (CNNs) or Transformers) are trained on large amounts of labeled data and can identify semantic features in video frames (such as facial expressions, object positions, and scene changes). After inputting independent frame sequences into the model, the model outputs an "effect fit score" for each frame. The higher the score, the more suitable the frame is for inserting special effects (e.g., keyframes with clear facial features are suitable for adding stickers, and scene transition frames are suitable for adding transition effects). Through model inference, high-scoring frames in each sequence are automatically selected as effect insertion frames, reducing manual annotation costs, improving processing efficiency, and leveraging the model's generalization ability to adapt to different types of video content and special effects requirements.

[0038] For example, the training method of the frame semantic model can adopt a supervised learning framework, which is built based on a large amount of labeled video data. By pre-collecting video clips containing various semantic tags (such as "person-running", "scene-night", "action-waving"), single-frame images are extracted and their semantic categories are manually labeled. Then, a convolutional neural network (CNN) or Transformer architecture is used as the backbone network. The RGB pixel values ​​or feature maps of the input frame images are used to extract spatial features (such as edges, textures, and object contours) through multi-layer nonlinear transformations. Then, temporal information (such as optical flow vectors of adjacent frames) is introduced to enhance the model's understanding of dynamic semantics (optional). Then, the model parameters are optimized through the cross-entropy loss function to minimize the error between the output semantic tag distribution and the real tags. Finally, the accuracy of the model in classifying the semantics of unseen video frames is tested on the validation set. The generalization ability of the model is improved by adjusting the network structure, data augmentation strategies (such as rotation, scaling, and color dithering), or learning rate decay mechanism, and finally, a frame semantic model that can accurately identify the semantic content of video frames is obtained.

[0039] S2312, count the number of special effects insertion frames in each independent video frame sequence, and determine the blank frame data corresponding to each independent video frame sequence.

[0040] It's understandable that the number and distribution of inserted effect frames can be used to calculate the required number of blank frames. For example, if a frame sequence contains 3 inserted effect frames, and the preset rule is to insert 2 blank frames before and after each inserted frame, then the sequence needs to insert a total of (3×4) = 12 blank frames. By counting the number of inserted frames and combining this with a preset blank frame generation strategy (such as fixed-interval insertion, dynamic adjustment based on frame spacing, etc.), the position and number of blank frames corresponding to each inserted frame can be determined.

[0041] S232, perform independent frame interpolation on each independent video frame sequence based on the blank frame data to obtain multiple video segment data corresponding to the original video data.

[0042] It's understandable that blank frames can be inserted at specified positions to create space for special effects processing. Blank frames can be black frames, transparent frames, or copies of adjacent frames, depending on the desired effect. For example, if the effect is to overlay semi-transparent text, a transparent blank frame can be inserted; if a pause effect is needed, a copy of the previous blank frame can be inserted. Using video editing libraries (such as FFmpeg) or self-developed tools, blank frames can be inserted frame-by-frame according to their position and number in the data to generate new video clips. The inserted clips add blank frame intervals to the original frame sequence, forming a structure of "original frame - blank frame - effect insertion frame - blank frame - original frame," providing a usable time window for subsequent special effects generation.

[0043] For example, in S232, each independent video frame sequence is independently interpolated based on the blank frame data to obtain multiple video segment data corresponding to the original video data, including: S2321, based on the blank frame data, blank frames are inserted before and after the special effects insertion frame in each independent video frame sequence to obtain multiple video segment data corresponding to the original video data.

[0044] It's understandable that inserting a blank frame before the effect's insertion frame can be used for pre-effect setup (such as an effect element entering the frame from off-screen); inserting a blank frame after it can be used for subsequent extensions of the effect (such as the gradual disappearance of residual effects). For example, for an explosion effect, the preceding blank frame can be used to show the environmental setup before the explosion, and the following blank frame can be used to show the smoke spreading after the explosion. This bidirectional insertion ensures that the effect has sufficient display space on the timeline, avoiding the cramped effect caused by a single insertion position. After the insertion operation is completed, the frame sequence structure of each segment becomes more complex, but it provides richer temporal support for the natural presentation of the effect.

[0045] S300 determines the first special effects frame data corresponding to each video segment data based on multiple video segment data corresponding to the original video data.

[0046] Understandably, fragment data and blank frame information can be integrated to generate a set of frame data for initial effects processing. The first set of effects frame data contains a complete sequence of original frames, blank frames, and effect insertion frames for each fragment, serving as the foundational material for effects generation. For example, a fragment sequence containing blank frames might be "Frame 1 - Frame 2 (blank) - Frame 3 (inserted frame) - Frame 4 (blank) - Frame 5," where frames 2 and 4 are blank frames, and frame 3 is the effect insertion frame to be processed. By arranging the fragment data chronologically and labeling the type of each frame (original frame, blank frame, inserted frame), a structured frame data list is formed, facilitating subsequent batch processing of specific frame types, such as extracting blank frame positions and locating the context of inserted frames.

[0047] In one possible implementation, S300, based on multiple video segment data corresponding to the original video data, determines the first special effects frame data corresponding to each video segment data, including: S310, perform frame extraction on each video segment data to obtain all special effects frame data packets corresponding to each video segment data; wherein, the special effects frame data packets include special effects insertion frames and the corresponding blank frames before and after the special effects insertion frames.

[0048] It's understandable that each video segment's data can be broken down into smaller processing units—effect frame data packets. Each data packet centers on the inserted effect frame and includes the blank frames before and after it, forming a local processing window. For example, if the inserted frame is frame 10, the preceding blank frame is frame 9, and the following blank frame is frame 11, then the data packet includes these three frames. This structure facilitates centralized processing of the transition area between the effect and the original video, ensuring a natural connection between the effect and the preceding and following frames when it's inserted. By traversing the frame sequence of each segment and extracting data packets according to a preset window size (e.g., one frame before and after the inserted frame), modular processing of complex segments is achieved, reducing the computational complexity of effect generation.

[0049] S320, perform frame copying based on the special effects insertion frame in the special effects frame data packet, and determine the preceding and following special effects insertion frames corresponding to the special effects insertion frame.

[0050] It's understandable that transitional footage is generated by copying the inserted frame. The purpose of copying is to fill the blank frame positions with the same content as the inserted frame, forming a continuous visual foundation, which facilitates visual consistency when subsequent effects are overlaid. For example, the inserted frame (frame 3) is copied as Frame 3_Before and Frame 3_After, replacing the previous blank frame (frame 2) and the subsequent blank frame (frame 4) respectively, making the segment sequence "Frame 1 - Frame 3_Before - Frame 3 - Frame 3_After - Frame 5". At this point, the previous and subsequent blank frames are replaced with the same content as the inserted frame, providing a unified background for the effects and avoiding abrupt effects due to significant differences between the blank frame content and the inserted frame. It also simplifies the calculation of blending effects with the original video.

[0051] Optionally, S320, frame duplication is performed based on the special effects insertion frame in the special effects frame data packet, determining the preceding and following special effects insertion frames corresponding to the special effects insertion frame, including: S321, perform frame copying based on the special effect insertion frame in the special effect frame data packet, and determine the first copy frame and the second copy frame corresponding to the special effect insertion frame.

[0052] It can be understood that the frame duplication product of the inserted effect frame is two independent copies: a first copy frame and a second copy frame. The first copy frame (e.g., before frame 3_) is used to fill the blank frame position before the inserted effect frame, and the second copy frame (e.g., after frame 3_) is used to fill the blank frame position after the inserted effect frame. The duplication process can be achieved through pixel-level copying, ensuring that the copies are completely identical to the original inserted frame. For example, if the inserted frame contains a close-up of a person's face, both copy frames will also present the same close-up of the face, providing the basic material for adding effects with different directions (such as left-hand and right-hand lighting effects) to these two frames later. By generating two copies, the areas before and after the inserted frame can be differentiated, increasing the flexibility and layering of the effects.

[0053] S322, replace the blank frame corresponding to the effect insertion frame based on the first copied frame, and replace the blank frame corresponding to the effect insertion frame based on the second copied frame to obtain the front effect insertion frame and the back effect insertion frame corresponding to the effect insertion frame.

[0054] It's understandable that the first and second duplicate frames can be replaced in the blank frame position to construct the pre-processing structure for special effects. After the replacement, the area of ​​the first blank frame becomes the first duplicate frame (the same as the inserted effect frame), and the area of ​​the second blank frame becomes the second duplicate frame (again, the same as the inserted effect frame), forming a three-frame structure of "duplicate frame-inserted frame-duplicate frame". For example, the dark background in the original blank frame position is replaced with the bright scene in the inserted frame. At this time, adding a gradually darkening effect to the duplicate frame can simulate the transition of the scene from bright to dark, while the inserted frame itself remains unchanged, forming a visual focus. This structure provides a carrier for the gradual presentation of special effects, allowing the effects to extend forward and backward from the inserted frame as the center, enhancing the dynamic effect of the picture.

[0055] S330, perform special effects insertion on the pre-effect insertion frame and the post-effect insertion frame respectively to obtain the pre-effect frame and post-effect frame corresponding to the special effects insertion frame.

[0056] It's understandable that user-required effects can be added to the copied frames. The pre-effect frame (first copied frame) and the post-effect frame (second copied frame) serve as the carriers of the effects, undergoing differentiated processing based on semantic tags and parameter configurations. For example, the pre-effect frame can add a particle effect moving to the left, while the post-effect frame can add a blur effect moving to the right, so that the inserted frame (such as the moment a character jumps) is surrounded by bidirectional effects, highlighting the impact of the action. Effect insertion methods include pixel rendering, filter overlay, and animation compositing, which can be accelerated through the graphics processing unit (GPU) to improve processing efficiency. The processed pre-effect frame, post-effect frame, and inserted frame together constitute a complete effect display unit, achieving a smooth transition from the original video to the effects.

[0057] S340, determine the first effect frame data corresponding to each video segment data by determining the previous and next effect frames of all effect insertion frames for each video segment data.

[0058] This approach allows for the integration of all processed effect frames within a single video clip. For example, if a video clip contains three inserted effect frames, each generating corresponding before and after effect frames, then the first effect frame data of that video clip contains six effect frames (3×2). These frame data are arranged in the chronological order of the original clip, forming a complete video stream together with the original frames, blank frames, etc. By binding effect frame data to video clip data, it ensures accurate tracing of the original position and context of each effect in subsequent processes, facilitating debugging, modification, and optimization, while providing a complete set of materials for generating the final effect scheme.

[0059] S400, remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and determine the second special effects frame data corresponding to each video segment data.

[0060] It is understandable that the special effects frame data can be optimized by removing the required parts. The reverse filtering mechanism removes special effects frame data that is highly related to the special effects requirements. This is intended to address scenarios such as changes in requirements, visual redundancy, comparison of multiple versions, or algorithm misjudgment. In other words, based on preset rules, frames that highly meet the original requirements are removed from the first special effects frame data, while frames with medium to low relevance are retained as the optimized second special effects frame data. This ensures that the special effects presentation meets the new priority or avoids over-rendering, while maintaining the logical coherence and functional integrity of the frame sequence through effect verification.

[0061] In one possible implementation, the first effects frame data includes a previous effects frame and a subsequent effects frame; S400, the effects frame data corresponding to the effects requirement information is removed from the first effects frame data corresponding to each video segment data, and the second effects frame data corresponding to each video segment data is determined, including: S410, calculate the correlation between the first special effects element and the special effects requirement information corresponding to the previous special effects frame based on the first special effects frame data, and calculate the correlation between the second special effects element and the special effects requirement information corresponding to the subsequent special effects frame based on the first special effects frame data.

[0062] It is understandable that the matching degree between special effects frames and required elements can be evaluated through feature extraction and quantization algorithms. Specifically, for the first special effects frame, visual features (such as RGB values, pixel distribution, and texture patterns) related to the special effects requirement information (such as specified color, shape, and position parameters) are extracted, and the correlation score between the two is calculated using cosine similarity or mean square error algorithms. At the same time, for the second special effects frame, dynamic features (such as optical flow vectors, displacement curves, and alpha channel values) related to the special effects requirement information (such as motion trajectory, speed, and transparency parameters) are extracted, and the correlation score is calculated using dynamic time warping (DTW) or Euclidean distance algorithms. Finally, quantifiable requirement matching indicators are generated for the first and second special effects frames, namely the correlation between the first special effects frame and the first special effects element corresponding to the special effects requirement information, and the correlation between the second special effects frame and the second special effects element corresponding to the special effects requirement information, providing data support for subsequent comparison decisions.

[0063] S420: Compare the correlation between the first special effect element and the second special effect element to determine the removal frame corresponding to the first special effect frame data.

[0064] It is understandable that targeted selection of special effects frames can be achieved by comparing quantitative indicators through comparative algorithms, that is, by comparing the numerical values ​​based on the calculated relevance of the first special effects element in the preceding special effects frame and the relevance of the second special effects element in the following special effects frame.

[0065] Optionally, S420, the relevance of the first special effect element is compared with the relevance of the second special effect element to determine the removal frame corresponding to the first special effect frame data, including: S421, when the correlation of the first special effect element is greater than that of the second special effect element, the previous special effect frame is determined as the removal frame corresponding to the data of the first special effect frame.

[0066] It is understandable that special effect frames can be filtered through reverse priority logic. That is, when the relevance of the first special effect element in the current special effect frame is greater than the relevance of the second special effect element in the subsequent special effect frame, the previous special effect frame is marked as a removed frame, so as to retain the subsequent special effect frame with lower relevance but more in line with the current needs. In this way, the priority of different elements is balanced in the presentation of special effects, and the overemphasis of a single dimension is avoided from affecting the overall effect.

[0067] S422, when the relevance of the first effect element is not greater than that of the second effect element, the later effect frame is determined as the removal frame corresponding to the data of the first effect frame. It is understandable that special effects frame filtering can be achieved through reverse priority logic. That is, if the relevance of the first special effects element in the current special effects frame is not greater than the relevance of the second special effects element in the subsequent special effects frame, the subsequent special effects frame is identified as the object to be removed, so as to retain the previous special effects frame with higher relevance. This ensures that the special effects sequence eliminates secondary or redundant parts under the guidance of requirements, highlights the presentation effect of key elements, and achieves the unity of visual focus and functional requirements.

[0068] S430, based on the removal of frames, remove the effect frame data corresponding to the effect requirement information from the first effect frame data corresponding to each video segment data, and determine the second effect frame data corresponding to each video segment data.

[0069] It's understandable that video stream processing algorithms (such as FFmpeg's select filter) can be used to remove frames by timestamp, regenerating segments with consistent frame rates. Based on the identified removed frames (previous or subsequent effect frames), the corresponding frame sequences are located and deleted from the timeline of the first effect frame data, while unmarked effect frames in the other direction are retained, and these are then reassembled into structurally complete second effect frame data. By synchronously adjusting the timestamps and metadata of the frame sequences, it's ensured that the removed segments are unaffected in terms of duration, frame rate, and logical coherence. Ultimately, optimized frame data that meets the priority requirements of the effects while avoiding redundancy is generated, laying the foundation for subsequent generation of the target video effect display scheme.

[0070] S500 generates a target video effects display scheme based on each video segment data and the corresponding first and second effects frame data. It is understandable that a complete special effects implementation plan can be generated by integrating multi-dimensional data and visual design. This involves aligning the original video clips with their corresponding first and second special effects frame data along the timeline, labeling each clip with its special effects type, parameters (such as duration, intensity, and position), and version differences (such as whether previous special effects frames are retained). The frame sequence distribution of different versions of special effects can be displayed using timeline views, split-screen comparisons, etc., and special effects elements can be distinguished by color coding or labels (such as red marking highly relevant frames and blue marking low-relevance frames). Finally, by combining audio track synchronization strategies, output format configurations (such as resolution and encoding parameters), and user interaction instructions (such as adjustable parameter nodes), a previewable and editable target video special effects display plan can be formed. This provides users with intuitive effect references and decision-making basis, ensuring that the final special effects presentation meets the expected visual logic and functional requirements.

[0071] Corresponding to the video production effects display method with enhanced user interactivity described in the above embodiments, this application also provides a video production effects display device with enhanced user interactivity. Each unit of the device can implement each step of the video production effects display method with enhanced user interactivity. Figure 3 This diagram illustrates the structure of a video production effects display device that enhances user interactivity according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiments of this application are shown.

[0072] Reference Figure 3 The enhanced and user-interactive video production effects demonstration device includes: The acquisition unit is used to acquire special effects requirement information and raw video data; The parsing unit is used to parse the original video data according to the special effects requirement information and determine multiple video segment data corresponding to the original video data; A segment unit is used to determine the first special effects frame data corresponding to each video segment data based on multiple video segment data corresponding to the original video data; The removal unit is used to remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and to determine the second special effects frame data corresponding to each video segment data; The generation unit is used to generate a target video effects display scheme based on each video segment data and the first special effects frame data and the second special effects frame data corresponding to each video segment data.

[0073] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0074] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit module can exist physically separately, or two or more unit modules can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0075] This application also provides an electronic device. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 6 of this embodiment includes: at least one processor 60 ( Figure 4 Only one is shown in the image), at least one memory 61 ( Figure 4 (Only one is shown in the image) and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60, wherein when the processor 60 executes the computer program 62, it causes the electronic device 6 to perform the steps in any of the above embodiments of the enhanced and user-interactive video production effects display methods, or causes the electronic device 6 to perform the functions of each unit in the above embodiments of the apparatus.

[0076] For example, the computer program 62 may be divided into one or more units, which are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 62 in the electronic device 6.

[0077] Electronic device 6 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. This electronic device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that... Figure 4This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.

[0078] The processor 60 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0079] In some embodiments, the memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. In other embodiments, the memory 61 may be an external storage device of the electronic device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 6. Furthermore, the memory 61 may include both internal and external storage units of the electronic device 6. The memory 61 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 61 can also be used to temporarily store data that has been output or will be output.

[0080] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0081] This application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the steps in any of the above method embodiments.

[0082] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0083] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0084] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0085] In the embodiments provided in this application, it should be understood that the disclosed enhanced and user-interactive video production effects display device / electronic device and method can be implemented in other ways. For example, the embodiments of the enhanced and user-interactive video production effects display device / electronic device described above are merely illustrative. For instance, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.

[0086] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0087] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for enhanced and user-interactive video production special effect presentation, characterized by, include: Obtain information on special effects requirements and raw video data; The original video data is parsed based on the special effects requirements information to determine multiple video segment data corresponding to the original video data; Based on the multiple video segment data corresponding to the original video data, determine the first special effects frame data corresponding to each video segment data; Remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and determine the second special effects frame data corresponding to each video segment data; Based on each video segment data and the corresponding first and second special effects frame data, a target video special effects display scheme is generated.

2. The video production effects display method for enhancing user interactivity as described in claim 1, characterized in that, The step of parsing the original video data based on the special effects requirement information to determine multiple video segment data corresponding to the original video data includes: Based on the special effects requirement information, special effects semantic classification is performed to obtain multiple special effects semantic information corresponding to the special effects requirement information; The original video data is parsed based on multiple special effects semantic information to determine the independent video frame sequence corresponding to each special effects semantic information; Based on the independent video frame sequence corresponding to each of the special effects semantic information, multiple video segment data corresponding to the original video data are obtained.

3. The video production effects display method for enhancing user interactivity as described in claim 2, characterized in that, The process of obtaining multiple video segment data corresponding to the original video data based on the independent video frame sequence corresponding to each of the special effects semantic information includes: Based on the independent video frame sequence corresponding to each of the special effects semantic information, determine the blank frame data corresponding to each of the independent video frame sequences; Based on the blank frame data, each independent video frame sequence is independently interpolated to obtain multiple video segment data corresponding to the original video data.

4. The video production effects display method for enhancing user interactivity as described in claim 3, characterized in that, The step of determining the blank frame data corresponding to each independent video frame sequence based on the semantic information of each special effect includes: Each of the independent video frame sequences is input into a preset frame semantic model to obtain special effects insertion frames for each of the independent video frame sequences; wherein, the frame semantic model is a machine learning model that has been pre-trained. The number of effect insertion frames in each independent video frame sequence is counted to determine the blank frame data corresponding to each independent video frame sequence.

5. The video production special effects display method for enhancing user interactivity as described in claim 4, characterized in that, The step of independently interpolating each independent video frame sequence based on the blank frame data to obtain multiple video segment data corresponding to the original video data includes: Based on the blank frame data, blank frames are inserted before and after the special effects insertion frame in each independent video frame sequence to obtain multiple video segment data corresponding to the original video data.

6. The video production special effects display method for enhancing user interactivity as described in claim 1, characterized in that, The step of determining the first special effects frame data corresponding to each of the multiple video segment data corresponding to the original video data includes: Frame extraction is performed on each video segment data to obtain all special effects frame data packets corresponding to each video segment data; wherein, the special effects frame data packets include special effects inserted frames and the pre-blank frames and post-blank frames corresponding to the special effects inserted frames; Based on the special effects inserted frame in the special effects frame data packet, perform frame duplication to determine the number of the preceding and following special effects inserted frames corresponding to the special effects inserted frame; Special effects are inserted into the pre-effect insertion frame and the post-effect insertion frame respectively to obtain the pre-effect frame and post-effect frame corresponding to the special effects insertion frame; The first effect frame data corresponding to each video segment data is determined by the first effect frame data of each video segment data, which consists of the first effect frame and the second effect frame of all effect insertion frames of each video segment data.

7. The video production special effects display method for enhancing user interactivity as described in claim 6, characterized in that, The step of performing frame duplication based on the special effects inserted frame in the special effects frame data packet, and determining the number of preceding and following special effects inserted frames corresponding to the special effects inserted frame, includes: Based on the special effect inserted frame in the special effect frame data packet, perform frame duplication to determine the first and second copied frames corresponding to the special effect inserted frame; The first copied frame replaces the blank frame corresponding to the effect insertion frame, and the second copied frame replaces the blank frame corresponding to the effect insertion frame, to obtain the number of the first and second effect insertion frames corresponding to the effect insertion frame.

8. The video production effects display method for enhancing user interactivity as described in claim 1, characterized in that, The first special effects frame data includes a previous special effects frame and a subsequent special effects frame; the step of removing the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and determining the second special effects frame data corresponding to each video segment data, includes: Calculate the correlation between the first special effects frame and the first special effects element corresponding to the special effects requirement information based on the first special effects frame data, and calculate the correlation between the second special effects frame and the second special effects element corresponding to the special effects requirement information based on the first special effects frame data. The correlation between the first special effect element and the correlation between the second special effect element are compared to determine the removal frame corresponding to the first special effect frame data; Based on the removal of the effect frame data corresponding to the effect requirement information from the first effect frame data corresponding to each video segment data, the second effect frame data corresponding to each video segment data is determined.

9. The video production special effects display method for enhancing user interactivity as described in claim 8, characterized in that, The step of comparing the relevance of the first special effects element with the relevance of the second special effects element to determine the removal frame corresponding to the first special effects frame data includes: When the relevance of the first special effect element is greater than that of the second special effect element, the previous special effect frame is determined as the removal frame corresponding to the first special effect frame data; When the relevance of the first special effect element is not greater than that of the second special effect element, the subsequent special effect frame is determined as the removal frame corresponding to the first special effect frame data.

10. A video production special effects display device that enhances and enhances user interactivity, characterized in that, include: The acquisition unit is used to acquire special effects requirement information and raw video data; The parsing unit is used to parse the original video data according to the special effects requirement information and determine multiple video segment data corresponding to the original video data; A segment unit is used to determine the first special effects frame data corresponding to each video segment data based on multiple video segment data corresponding to the original video data; The removal unit is used to remove the special effects frame data corresponding to the special effects requirement information from the first special effects frame data corresponding to each video segment data, and to determine the second special effects frame data corresponding to each video segment data; The generation unit is used to generate a target video effects display scheme based on each video segment data and the first special effects frame data and the second special effects frame data corresponding to each video segment data.