Intelligent video editing optimization method

An AI-driven video editing system dynamically predicts audience characteristics, optimizes content through automated engines, and iteratively improves adaptation, addressing inefficiencies and cultural mismatches to enhance video content matching and user conversion.

CN120321429APending Publication Date: 2025-07-15BOYUBO INTERNET CROSS-BORDER COMMERCIAL TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510359271.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Traditional video editing methods are inefficient and have poor cross-cultural adaptability, and cannot optimize video content in real time, resulting in unreasonable display of key information, insufficient matching accuracy for target audiences, and high error rate for identification and replacement of sensitive elements in cross-cultural scenarios, slow response speed, and it is difficult to continuously optimize through the closed-loop feedback mechanism.

Method used

The artificial intelligence model is used to predict target demographic characteristics and generate structured portraits, and the video content is dynamically optimized through a programmatic decision-making engine, and iterative optimization is carried out in combination with a closed-loop feedback mechanism, including lens selection, duration allocation and cultural adaptation. The YOLOv5 and CLIP models are used to detect sensitive symbols and generate neutral alternative content to achieve seamless synergy of automated editing strategies.

Benefits of technology

It significantly improves the accuracy of cross-cultural adaptation, strategy response speed and user conversion effect, reduces the cross-cultural error rate, and realizes personalized matching and continuous optimization of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321429A_ABST
    Figure CN120321429A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent video editing optimization system and method based on target audience prediction, and belongs to the technical field of artificial intelligence and digital marketing. According to the method, dynamic analysis of target audience features and real-time adaptation of a video editing strategy are realized through a programmed automatic process, and the problems of high manual dependency, poor cross-culture adaptability, strategy updating lagging and the like in the traditional technology are solved. The method specifically comprises the following steps: calling a third-party platform application program interface to obtain user behavior data, and generating a dynamic audience portrait by using an artificial intelligence large model; shot selection, duration distribution and culture adaptation of video content are optimized through a reinforcement learning algorithm on the basis of portraits, a shot selection module dynamically distributes technical parameter close-up and cost comparison content according to user role types, and a culture adaptation module calls a taboo library to automatically replace sensitive elements; a synchronous test is carried out by generating multiple versions of videos, closed-loop feedback iteration is realized based on a user click rate and a conversion rate index, and an editing strategy is continuously optimized. Compared with a traditional method, the system has the advantages that the cross-culture adaptation precision, the strategy response speed and the user conversion effect of the video content are remarkably improved through cooperation of the artificial intelligence model and the programmed decision engine, and the system is suitable for precision marketing requirements of multiple scenes such as cross-border e-commerce and industrial equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and digital marketing, and specifically to an intelligent video editing optimization system and method based on target audience prediction, which is particularly suitable for the precise marketing needs of scenarios such as cross-border e-commerce and industrial equipment. The system realizes the dynamic analysis of user behavior data and the automatic adaptation of video editing strategies through artificial intelligence technology, solves problems such as poor cross-cultural adaptability and lagging strategy updates in traditional methods, and improves the personalized matching efficiency and communication effect of video content. Background Art

[0002] Traditional video editing relies on manual analysis of user data and content adjustment, with defects such as low efficiency, insufficient cross-cultural adaptability, and strategy solidification. Existing automated editing tools mostly generate content based on fixed rules and cannot dynamically optimize video elements according to real-time user portraits, resulting in unreasonable display rhythms of key information and insufficient matching accuracy for target audiences. In cross-cultural scenarios, the identification and replacement of sensitive elements rely on manual screening, with problems such as high error rates and slow response speeds. In addition, static editing strategies are difficult to continuously optimize content adaptability through a closed-loop feedback mechanism, and the improvement of user conversion effects is limited.

[0003] The present invention aims to provide an intelligent video editing optimization system and method to solve the above problems through the following technical improvements: Based on artificial intelligence large models, dynamically predict the characteristics of target audiences and generate structured portraits, replacing manual experience judgment; Dynamically optimize the shot selection, duration allocation, and cultural adaptation of video content through a programmatic decision-making engine to improve cross-scenario adaptation capabilities; Construct a closed-loop feedback mechanism, continuously iterate editing strategies based on user behavior data, and break through the limitations of traditional static strategies.

[0004] Compared with traditional methods, the system of the present invention has significant improvements in cross-cultural error rate control, strategy response speed, and user conversion effect. Summary of the Invention

[0005] The present invention provides an intelligent video editing optimization method, which realizes the coordinated operation of target audience prediction, editing strategy optimization, and closed-loop feedback iteration through a programmatic automation process. The specific technical solutions are as follows: In the target audience prediction stage, the system automatically collects user behavior data and industry reports by programmatically calling the application program interfaces of third-party self-media platforms. Subsequently, an artificial intelligence large model deeply analyzes the data, extracts user interest tags, geographical features, and decision-making preferences, and generates a dynamically updated structured audience portrait. This process realizes real-time portrait updates through an automated data pipeline, providing continuous input for subsequent editing strategies.

[0006] In the clip strategy optimization phase, multi-dimensional automated adjustments are executed through a programmed decision-making engine. Based on the audience profile, the system automatically allocates the presentation ratio of technical parameter close-ups and cost comparison content: for audience groups with a high proportion of technical personnel, the program automatically increases the weight of technical shots and extends the parsing duration; for groups dominated by procurement decision-makers, the cost comparison content is automatically strengthened and dynamically generated adaptive subtitles are inserted. The duration allocation module optimizes the shot transition rhythm in a programmed manner by integrating a reinforcement learning algorithm to ensure that the error rate of the dwell time of key information is stably controlled within a preset threshold. The cultural adaptation module detects sensitive symbols (such as religious totems) in video frames through the YOLOv5 model, and performs semantic matching in combination with the CLIP model. If the confidence level > 90%, the Stable Diffusion model is called to generate neutral alternative content, and the replacement area coverage is ≤ 30% and does not damage the main structure of the video.

[0007] In the closed-loop feedback iteration phase, continuous optimization is achieved through a programmed testing framework. The system automatically generates multiple versions of the video and synchronously launches them, and real-time monitors click-through rate, conversion rate, and completion rate indicators through the data embedding program. Based on the data updated daily, the programmed parameter adjustment module automatically fine-tunes the prompting words of the artificial intelligence model and the algorithm weights. For example, when the user dwell time of technical content is lower than expected, the system automatically triggers the shot weight coefficient adjustment program and optimizes the display logic of key frames.

[0008] Each module realizes collaboration through a standardized data interface and an event-driven mechanism. The structured profile output by the target audience prediction module automatically triggers the parameter configuration of the clip strategy engine; multiple versions of the video generated during the strategy execution process automatically enter the A / B test queue; the feedback data is transmitted back to the model training module in real time through the event bus, forming a closed-loop optimization link. The programmed process ensures the full-link automated execution from data collection, strategy generation to effect feedback, eliminates the delay of manual intervention, and realizes efficient collaboration across modules.

[0009] Compared with traditional methods, the present invention realizes seamless connection of the prediction, optimization, and feedback links through a programmed automated architecture, and has a significant improvement in cross-cultural adaptation accuracy, strategy response speed, and system iteration efficiency. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0011] Figure 1: System architecture flowchart, showing the complete closed-loop process from data collection to strategy iteration.

[0012] Figure 2: Flowchart for optimizing the editing strategy, describing the collaborative logic of shot selection, duration allocation, and cultural adaptation. Detailed implementation manners

[0013] The detailed implementation of the intelligent video editing optimization method of the present invention is described as follows with reference to FIG. 1 and FIG. 2: Target audience prediction (corresponding to FIG. 1) As shown in FIG. 1, the system automatically accesses the third-party self-media platform API through a programmatic interface to obtain user behavior data in real time, including viewing duration, interaction records, and geographical distribution information. The data is input into the AI large model for feature extraction to generate a dynamically updated audience portrait. The portrait contains structured tags of user role types, interest preferences, and decision-making characteristics, providing a decision-making basis for subsequent editing strategies.

[0014] Optimization of editing strategy (corresponding to FIG. 2) As shown in FIG. 2, the editing strategy optimization process includes the following core modules: Shot selection module: Dynamically allocate the presentation weights of close-up and cost comparison content of technical parameters according to the proportion of user roles in the audience portrait. For example, for technical audiences, the priority of technical analysis shots is increased, and multilingual adapted subtitles are automatically generated.

[0015] Duration allocation module: Optimize the display rhythm of key information through a reinforcement learning algorithm to ensure that the dwell time error rate of core parameters is stably controlled within a set threshold.

[0016] Cultural adaptation module: Call a preset taboo library to automatically detect and replace sensitive elements in the video (such as religious symbols, animal metaphors) with neutral content that conforms to the cultural norms of the target region.

[0017] The outputs of each module are integrated by the video synthesis module to generate multiple versions of videos adapted to different audiences.

[0018] Closed-loop feedback iteration (corresponding to FIG. 1) As shown in FIG. 1, the system conducts an A / B test on the generated videos, synchronously launches multiple versions and monitors user behavior metrics in real time (such as click-through rate, conversion rate). Based on the feedback of the test data, the system automatically optimizes the AI model parameters and editing strategies. For example, when the user dwell time on technical content is lower than expected, it triggers adjustments to the shot weights and optimization of the key frame display logic.

[0019] Programmatic automation collaboration (linkage between FIG. 1 and FIG. 2) The system realizes efficient collaboration between modules through a standardized data interface: The parameter configuration of the clip strategy engine is automatically triggered by the audience portrait data; After video synthesis, it automatically enters the A / B test queue; User behavior metrics are transmitted back to the model training end in real time through the event bus, forming a closed-loop optimization link.

[0020] Implementation example Taking the industrial equipment export scenario as an example, the system generates technical analysis version and business comparison version videos according to the predicted audience characteristics. The technical version focuses on the close-up display of technical certification parameters, the business version strengthens the analysis of supply chain cost advantages, and at the same time automatically replaces potential cultural taboo elements. The test results show that this system significantly improves the user conversion effect and reduces the cross-cultural communication risk.

[0021] AB test data comparison: The average conversion rate of the traditional manual editing group (100 videos) is 12%, and the cross-cultural error rate is 18%; The average conversion rate of the editing group of this system (100 videos) is 22% (+83%), and the cross-cultural error rate is 3% (-83%).

[0022] Technical parameters: The mAP@0.5 of sensitive element detection is 95%, and the word error rate of subtitle generation is <2%.

Claims

1. An intelligent video editing optimization method, characterized in that Including the following steps: Obtain user behavior data through the application programming interface of a third-party self-media platform, and call an artificial intelligence large model to generate a dynamic audience portrait; Based on the audience portrait, dynamically adjust the video content through an artificial intelligence model, including allocating the presentation ratio of technical parameter close-ups and cost comparison content according to the user role type, optimizing the display duration of key information, and calling a taboo library based on an image semantic segmentation model (such as U-Net) and an object detection model (such as YOLOv5) to detect sensitive elements (religious symbols, animal metaphors) in the video frame in real time, and replacing them with neutral content that conforms to the cultural norms of the target region through a generative adversarial network (GAN). Generate multiple versions of the same product video for synchronous testing, and optimize the model parameters and iterative editing strategies based on the user behavior index data.

2. The intelligent video editing optimization method according to claim 1, wherein: The dynamic adjustment of the video content includes enhancing the weight of technical parameter type shots through a shot selection module, and generating adapted subtitles through the multi-language translation API of the Transformer architecture, supporting languages such as Chinese, English, and Arabic, with a subtitle generation delay ≤ 50ms.

3. The intelligent video editing optimization method according to claim 1, wherein: The optimization of the display duration of key information is achieved through a reinforcement learning model, which controls the rhythm of shot switching to stabilize the residence duration of key content.

4. An intelligent video editing optimization method according to claim 1, characterized in that: The call to the taboo library to replace sensitive elements includes automatically identifying religious symbols and animal metaphors and replacing them with neutral content.

Citation Information

Cited By

  • Short video clip automatic generation method based on neural network model

    CN121486655A