AI video real-time editing system based on natural language interaction

Through the multimodal AI fusion physics engine and the generation and adversarial network optimization video content, combined with search enhanced generation and dynamic template libraries, the AI ​​video editing tools have achieved significant improvements in authenticity, accuracy, personalization and global adaptation, and solved the limitations of the existing technology.

CN120583281AInactive Publication Date: 2025-09-02YUNMU TECHNOLOGY (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510425242.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-09-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AI video editing tools lack an in-depth understanding of physical laws, resulting in errors in the generated video logic, insufficient content accuracy, rigid editing templates, and inability to personalize adjustments, limiting the application of tools in professional fields and multilingual scenarios.

Method used

The multimodal AI fusion physics engine and the generative adversarial network are used for screen optimization, combined with search enhancement generation technology to verify the accuracy of content, and a dynamic template library is built to adjust clips based on user portraits, realize text-video bidirectional mapping editing, and support multilingual processing.

Benefits of technology

The generated video content is more realistic and accurate, and can be personalized to adapt to different user needs, improving the authenticity, accuracy and global adaptability of the video, and reducing production costs and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120583281A_ABST
    Figure CN120583281A_ABST
Patent Text Reader

Abstract

The invention discloses an AI video real-time editing system based on natural language interaction, which relates to the technical field of artificial intelligence AI and comprises a user input module used for receiving multi-modal data, including texts, voices or pictures, input by a user and converting the multi-modal data into standardized demand description; the multi-modal analysis module is used for analyzing the standardized demand description, extracting keywords and semantic information in the standardized demand description and generating corresponding sub-shot scripts; picture detail optimization is carried out by using a multi-modal AI fusion physical engine and a generative adversarial network; the motion trail of the object in the generated video content better conforms to the real physical law, such as elastic collision and parabolic motion of a sphere; the video can provide a more real and credible visual effect in scientific education, product demonstration and other scenes requiring high physical authenticity, so that the acceptance and persuasion of audiences are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence (AI) technology, and specifically to an AI video real-time editing system based on natural language interaction. Background Art

[0002] With the rapid development of artificial intelligence technology, AI video editing tools have gradually become an important auxiliary means for media production and content creation. These tools improve the efficiency and quality of video editing through automation and intelligence, lowering the threshold for professional video production. However, existing AI video editing tools still have the following technical limitations, which limit their potential application in a wider range of scenarios:

[0003] Insufficient understanding of physical laws: Current AI video generation models, such as Sora and Runway, primarily rely on imitating training data and lack a deep understanding of physical laws (such as object collisions and parabolic motion). This results in logical errors in the generated videos in untrained scenarios, such as abnormal sphere motion trajectories, which reduces the authenticity and credibility of the videos.

[0004] Content accuracy issues: AI-generated content, such as scripts or subtitles, often contains fabricated data, misquotes, and lacks reliable source identification. This inaccuracy limits the application of AI video editing tools in professional fields (such as law and medicine) and increases the cost of content correction for users.

[0005] Rigid editing templates: Existing AI editing tools, such as Yibingmiaochuang, typically rely on fixed templates and are unable to dynamically adjust editing rhythm, transition effects, or background music based on user profiles (such as audience age and region). These rigid templates limit the degree of video personalization and make it difficult to meet the differentiated needs of different industries (such as e-commerce promotions and education).

[0006] In response to the limitations of the above-mentioned existing technologies, the present invention proposes an AI video real-time editing system based on natural language interaction; this system systematically solves the core defects of existing AI video tools in authenticity, accuracy, personalization, interaction efficiency and global adaptation through multimodal AI fusion physics engine, RAG enhanced generation, dynamic template library and two-way editing technology; the technical effect of the present invention covers the whole process optimization from script generation to final editing, which can significantly reduce the production cost of corporate promotional videos and improve the content quality. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention provides an AI video real-time editing system based on natural language interaction, which solves the core problems of existing AI video tools in authenticity, accuracy, personalization, interaction efficiency and global adaptation proposed in the background technology.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: an AI video real-time editing system based on natural language interaction, comprising:

[0009] A user input module is used to receive multimodal data input by the user, including text, voice or picture, and convert the multimodal data into a standardized demand description;

[0010] A multimodal parsing module, configured to parse the standardized requirement description, extract keywords and semantic information therein, and generate corresponding storyboard scripts;

[0011] The physics simulation module combines the physics engine and the generative adversarial network (GAN) to verify the physical laws of the object motion scenes involved in the storyboard script and correct the scenes that do not conform to the physical laws;

[0012] A dynamic template library, which uses reinforcement learning to dynamically generate adaptive editing templates based on user profiles and industry labels, including editing parameters such as transition speed and music style;

[0013] The AI ​​generation engine is used to call video clips, pictures, and music materials from the material library based on the storyboard script and dynamic templates, generate the first version of the video, and add automatic dubbing and subtitles;

[0014] An interactive editing engine that enables bidirectional mapping between text and video. When users modify text content, the corresponding images, subtitles, and dubbing in the video are automatically adjusted synchronously.

[0015] Multi-language processing module, used to generate subtitles in multiple languages ​​and support localized interface adaptation;

[0016] The finished product output module is used to export the final video product and synchronize it to the cloud platform or generate data reports.

[0017] Preferably, the multimodal analysis module includes a natural language processing unit and an image recognition unit, wherein the natural language processing unit is used to extract keywords from text, and the image recognition unit is used to analyze key elements in images.

[0018] Preferably, the physical law simulation module includes a physical engine interface and a GAN correction unit, wherein the physical engine interface is used to call the physical engine to verify the object's motion trajectory, and the GAN correction unit is used to generate a correction frame to correct the image that does not conform to the physical law, wherein the verification of the object's motion trajectory can be expressed as:

[0019]

[0020] Among them, S(t) represents the state of the object at time t, S0 is the initial state, F(S(τ),u(τ)) is the dynamic function simulated by the physics engine, u(τ) is the control input, and G(S(t)) is the correction function generated by GAN.

[0021] Preferably, the dynamic template library includes a user portrait analysis unit and a reinforcement learning unit. The user portrait analysis unit is used to extract user portrait information, and the reinforcement learning unit is used to dynamically generate an adaptive editing template based on the user portrait and industry label. The generation process of the editing template can be expressed as follows:

[0022]

[0023] Among them, π * (a|s) represents the optimal strategy, s is the state of user portrait and industry label, a is the action of clipping template, r(s t ,a t ) is the reward function and γ is the discount factor.

[0024] Preferably, the interactive editing engine includes a text-video mapping unit and a synchronization adjustment unit, wherein the text-video mapping unit is used to establish a bidirectional mapping relationship between text and video, and the synchronization adjustment unit is used to automatically adjust the video content according to the text modification, wherein the bidirectional mapping between text and video can be expressed as:

[0025] V new =M·V old +b

[0026] Among them, V new is the adjusted video feature matrix, V old is the original video feature matrix, M is the mapping matrix, and b is the offset vector.

[0027] Preferably, the multi-language processing module includes a subtitle generation unit and a localization adaptation unit, wherein the subtitle generation unit is used to generate subtitles in multiple languages, and the localization adaptation unit is used to adapt a localization interface.

[0028] Preferably, the AI ​​generation engine includes a material calling unit and a content synthesis unit. The material calling unit is used to call video clips, pictures and music from the material library. The content synthesis unit is used to generate a first version of the video according to the storyboard script and dynamic template, and add automatic dubbing and subtitles.

[0029] Preferably, the finished product output module includes a video export unit and a cloud platform synchronization unit, the video export unit is used to export the final video product, and the cloud platform synchronization unit is used to synchronize the video product to the cloud platform or generate a data report.

[0030] Preferably, the AI ​​generation engine further includes a content optimization unit based on deep learning, which uses the user's historical editing data and preferences to train a personalized video content generation model, and is used to adjust the video content generation strategy according to the user's editing history and preferences. The optimization process of the content optimization unit can be expressed as follows:

[0031]

[0032] Where: V opt Represents the optimized video content feature vector; represents the initially generated video content feature vector; N is the number of user editing histories; T i is the feature vector of the i-th user’s editing history; ω i is the weight of the i-th editing history, which is determined by the user's satisfaction with the editing result; Loss i is the loss function of the i-th editing history, which is used to measure the difference between the video content and the user preference; Represents the gradient of the video content feature vector; α is the learning rate, which controls the optimization step size.

[0033] Preferably, the multilingual processing module further includes a context-aware subtitle generation unit, which utilizes natural language processing technology and machine translation technology to optimize subtitle generation to improve the accuracy and naturalness of subtitles. The analysis process of the context-aware subtitle generation unit can be expressed as follows:

[0034] C ctx =Transformer(V,C,S)

[0035] Where: C ctx represents the subtitle feature vector after considering the context; Represents the video content feature vector; C represents the initially generated subtitle feature vector; S represents the context feature vector of the video content, including scene description, character emotion, and other information; Transformer is a function based on the Transformer model that is used to generate more accurate subtitles based on video content and context;

[0036] The context-aware subtitle generation unit optimizes subtitle generation by analyzing the contextual information of the video content, such as scene descriptions and character emotions, so that the subtitles not only accurately reflect the video content but are also more natural and in line with the context.

[0037] The present invention provides an AI video real-time editing system based on natural language interaction.

[0038] Beneficial effects:

[0039] (1) Simulation of physical laws and enhancement of realism: Multimodal AI is used to integrate physics engines (such as the Bullet physics engine) and generative adversarial networks (GANs) to optimize image details. The motion trajectories of objects in the generated video content are more consistent with real physical laws, such as elastic collisions and parabolic motion of spheres. This enables videos to provide more realistic and credible visual effects in scenarios that require a high degree of physical realism, such as scientific education and product demonstrations, thereby improving audience acceptance and persuasiveness.

[0040] (2) Eliminate AI content "illusions" and improve generation accuracy: Integrate retrieval-augmented generation (RAG) technology to call authoritative databases (such as academic literature and corporate knowledge bases) in real time to verify generated content and automatically annotate data sources; through real-time verification and annotation, the accuracy of scripts and subtitles is significantly improved to over 99%, greatly reducing the need and cost of manual verification; this enables the system to provide more reliable and accurate video content in professional fields such as law and medicine that have extremely high requirements for content accuracy;

[0041] (3) Realize dynamic personalized editing and multi-scene adaptation: Build a dynamic template library based on user portraits and use reinforcement learning to optimize editing parameters; the system can automatically generate highly matching video content based on the user's specific needs and preferences, such as generating fast-paced promotional videos for e-commerce customers and slow-paced analytical videos for education customers. This personalized video content can better attract the target audience and improve the content conversion rate and user satisfaction;

[0042] (4) Provide real-time interactive editing and reverse adjustment functions: Develop a text-video bidirectional mapping engine. When the user modifies the text, the AI ​​will synchronously adjust the corresponding picture, subtitles, and dubbing. In commercial video production, users can quickly and conveniently perform multiple rounds of modifications and iterations without complicated operations. This efficient editing method can significantly shorten the production cycle, improve work efficiency, and meet the client's needs for high-frequency adjustments.

[0043] (5) Strengthening multi-language support and deep localization adaptation: Integrating multi-language processing technology, supporting multi-language subtitle generation and Indian local language interface, and cooperating with local operators to optimize distribution; the system can generate subtitles in multiple languages ​​and adapt to local interfaces in different regions, reducing the cost of multi-language content production by 60% and increasing enterprise user coverage by 50%. This helps companies better expand into the international market and reach a wider audience.

[0044] Through these specific technical solutions, the AI ​​video real-time editing system of the present invention has achieved significant improvements and enhancements in authenticity, accuracy, personalization, interactive efficiency and global adaptation, providing enterprises and individual users with more efficient, convenient and high-quality video editing tools. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION

[0046] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] Example 1: Production of a promotional video for a smartwatch

[0048] background:

[0049] A smartwatch manufacturer wanted to create a promotional video showcasing the watch's water resistance and heart rate monitoring capabilities, targeting sports enthusiasts.

[0050] Operation process:

[0051] 1. User input: The user uploads the product parameters of the smart watch (waterproof 50 meters, heart rate monitoring) through the system interface and enters a text description "Demonstrating the waterproof performance and heart rate monitoring function of the smart watch, aimed at sports enthusiasts."

[0052] 2. Multimodal analysis: The system parses the input text and product parameters, extracts the keywords "smartwatch," "waterproof," "heart rate monitoring," and "sports enthusiast," and uses the industry knowledge base to generate storyboards, including "underwater testing," "sports scene monitoring," and "health data display."

[0053] 3. Physical Law Verification: The system corrects the problem of bubble trajectory not conforming to fluid mechanics when the watch enters water in the physical law simulation module to ensure the authenticity of the image.

[0054] 4. Dynamic template matching: Based on the target audience of "sports enthusiasts", the system selects the "fast-paced editing + electronic music" template from the dynamic template library and sets the transition interval to 1.2 seconds.

[0055] 5. Interactive Editing: After previewing the initial version of the video, if the user decides to delete the "Night Mode Description", the system will automatically delete the corresponding night scene clip and connect the subsequent scenes.

[0056] 6. Multi-language output: The system generates 4K videos with English and Spanish subtitles.

[0057] 7. Final product output: The final video product will be pushed to social media platforms simultaneously.

[0058] Effect:

[0059] Through the system of the present invention, smart watch manufacturers can quickly generate high-quality, multilingual promotional videos to meet the needs of sports enthusiasts in different regions and enhance the market competitiveness of their products.

[0060] Example 2: Online Education Course Video Production

[0061] background:

[0062] A teacher on an online education platform needs to produce a chemistry experiment course video, which includes a demonstration of the experimental steps and an explanation of the principles of chemical reactions. The target audience is high school students.

[0063] Operation process:

[0064] 1. User input: Teachers input course content text, including experimental steps and chemical reaction principles, through the system interface and upload pictures of experimental operations.

[0065] 2. Multimodal analysis: The system analyzes the input text and images, extracts the keywords "chemical experiment," "experimental steps," and "reaction principle," and generates a storyboard script, including "experimental preparation," "operation demonstration," and "principle explanation."

[0066] 3. Physical law verification: The system verifies the physical rationality of experimental operations in the physical law simulation module, such as flame combustion, liquid flow, etc.

[0067] 4. Dynamic template matching: Based on the target audience of "high school students", the system selects the "slow-paced editing + educational style music" template from the dynamic template library, and sets the transition interval to 2 seconds.

[0068] 5. Interactive editing: After previewing the initial version of the video, the teacher decides to add more experimental safety tips, and the system automatically inserts the corresponding pictures and subtitles into the video.

[0069] 6. Multi-language output: The system generates high-definition videos with Chinese and English subtitles.

[0070] 7. Product output: The final video product is synchronized to the online education platform for students to learn.

[0071] Effect:

[0072] Through the system of the present invention, teachers can efficiently produce professional and accurate chemistry experiment course videos, improve teaching quality, and support multi-language output to meet the needs of students in different regions.

[0073] Example 3: Corporate promotional video production

[0074] background:

[0075] A technology company plans to create a promotional video to showcase its latest smart home product. The video needs to include a product introduction, functional demonstrations, and user experience feedback, with the target audience being potential consumers.

[0076] Operation process:

[0077] 1. User input: Technology companies enter product introduction text through the system interface and upload product images and demonstration video clips.

[0078] 2. Multimodal Parsing: The system parses input text and images, extracting keywords such as "smart home," "product introduction," "function demonstration," and "user experience." It also uses image recognition to analyze uploaded video clips for key elements, such as product appearance and user interface.

[0079] 3. Physical law verification: For physical interactions involved in product demonstrations, such as lighting control and automatic opening and closing of curtains, the system verifies the physical rationality of these actions in the physical law simulation module and makes necessary image corrections.

[0080] 4. Dynamic template matching: Based on the target audience of "potential consumers" and the industry label of "technology products", the system selects the "modern editing + relaxing background music" template from the dynamic template library and sets the transition interval to 1.5 seconds to attract the audience's attention.

[0081] 5. AI generation: The AI ​​generation engine calls relevant video clips, pictures, and music from the material library based on the storyboard script and dynamic template to generate a preliminary version of the video that includes product introduction, function demonstration, and user experience feedback.

[0082] 6. Interactive Editing: After previewing an initial video, a tech company might decide to enhance the details of a particular feature demonstration. Using the interactive editing engine, the system automatically inserts the corresponding images and subtitles into the video, eliminating the need for complex manual editing.

[0083] 7. Multi-language output: Taking into account the multilingual needs of potential consumers, the system generates high-definition videos with subtitles in Chinese, English and Spanish.

[0084] 8. Final product output: The final video product will be synchronized to the company's official website and social media platforms for potential consumers around the world to watch.

[0085] Effect:

[0086] The system of the present invention enables technology companies to quickly and efficiently produce professional-level promotional videos, which not only saves time and costs but also expands the scope of promotion through multi-language support. The high quality and attractiveness of the videos significantly increase potential consumers' interest in and willingness to purchase smart home products.

[0087] Working principle:

[0088] graph TD

[0089] A[User Input Module]-->B[Multimodal Parsing Module]

[0090] B-->C[Physical Law Simulation Module]

[0091] C-->D[Dynamic Template Library]

[0092] D-->E[AI Generation Engine]

[0093] E-->F[Interactive Editing Engine]

[0094] F-->G[Multi-language processing module]

[0095] G-->H[Finished product output module]

[0096] 1. User input module:

[0097] Users input multimodal data through the system interface, including text descriptions, voice commands or pictures.

[0098] The system automatically identifies the type of input data and converts it into a standardized requirement description, providing a unified data format for subsequent processing.

[0099] 2. Multimodal parsing module:

[0100] This module includes a natural language processing unit and an image recognition unit.

[0101] The natural language processing unit uses advanced NLP technologies (such as the BERT model) to extract keywords and semantic information from text.

[0102] The image recognition unit uses deep learning models such as CNN to analyze the key elements in the image.

[0103] Combining text and image information, a corresponding storyboard script is generated, providing a blueprint for the generation of video content.

[0104] 3. Physical law simulation module:

[0105] Combine the physics engine (such as the Bullet engine) and the GAN network to verify the physical laws of the object motion scenes involved in the storyboard script.

[0106] The physics engine previews the object's motion trajectory, and the GAN network generates corrected frames to ensure that the generated images conform to physical laws such as gravity and collision.

[0107] 4. Dynamic Template Library:

[0108] Based on user portraits (such as age and region) and industry labels, editing templates are dynamically generated through reinforcement learning.

[0109] The template includes editing parameters such as transition speed and music style to meet the needs of different users and scenarios.

[0110] 5. AI Generation Engine:

[0111] Integrate resources in the material library (video clips, pictures, music, etc.) and synthesize videos according to storyboard scripts and dynamic templates.

[0112] Add automatic dubbing and subtitles to make the video content richer and more complete.

[0113] 6.Interactive editing engine:

[0114] Realize two-way mapping between text and video. When the user modifies the text, the system automatically and synchronously adjusts the corresponding pictures, subtitles and dubbing in the video.

[0115] Supports multi-version backtracking to improve editing efficiency and flexibility.

[0116] 7. Multi-language processing module:

[0117] Generate subtitles in multiple languages, support 11 local Indian language interfaces, and call local operator APIs to optimize distribution.

[0118] Reduce multilingual video production costs and increase reach for non-English speaking audiences.

[0119] 8. Finished product output module:

[0120] Export finished videos (4K / 1080P), generate data reports (such as conversion rate predictions), and synchronize them to cloud platforms (such as AWS S3).

[0121] Provide finished product files and analysis reports to help users evaluate video effects and optimize subsequent production.

[0122] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, from all points of view, the embodiments should be regarded as illustrative and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and range of equivalents of the claims are included in the present invention. Any reference signs in the claims should not be construed as limiting the claim to which they relate.

[0123] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. An AI video real-time editing system based on natural language interaction, characterized in that: include: A user input module is used to receive multimodal data input by the user, including text, voice or picture, and convert the multimodal data into a standardized demand description; A multimodal parsing module, configured to parse the standardized requirement description, extract keywords and semantic information therein, and generate corresponding storyboard scripts; The physics simulation module combines the physics engine and the generative adversarial network (GAN) to verify the physical laws of the object motion scenes involved in the storyboard script and correct the scenes that do not conform to the physical laws; A dynamic template library, which uses reinforcement learning to dynamically generate adaptive editing templates based on user profiles and industry labels, including editing parameters such as transition speed and music style; The AI ​​generation engine is used to call video clips, pictures, and music materials from the material library based on the storyboard script and dynamic templates, generate the first version of the video, and add automatic dubbing and subtitles; An interactive editing engine that enables bidirectional mapping between text and video. When users modify text content, the corresponding images, subtitles, and dubbing in the video are automatically adjusted synchronously. Multi-language processing module, used to generate subtitles in multiple languages ​​and support localized interface adaptation; The finished product output module is used to export the final video product and synchronize it to the cloud platform or generate data reports.

2. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The multimodal analysis module includes a natural language processing unit and an image recognition unit. The natural language processing unit is used to extract keywords from text, and the image recognition unit is used to analyze key elements in pictures.

3. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The physical law simulation module includes a physical engine interface and a GAN correction unit. The physical engine interface is used to call the physical engine to verify the object's motion trajectory. The GAN correction unit is used to generate a correction frame to correct the image that does not conform to the physical law. The verification of the object's motion trajectory can be expressed as: Among them, S(t) represents the state of the object at time t, S0 is the initial state, F(S(τ),u(τ)) is the dynamic function simulated by the physics engine, u(τ) is the control input, and G(S(t)) is the correction function generated by GAN.

4. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The dynamic template library includes a user portrait analysis unit and a reinforcement learning unit. The user portrait analysis unit is used to extract user portrait information, and the reinforcement learning unit is used to dynamically generate an adaptive editing template based on the user portrait and industry label. The generation process of the editing template can be expressed as follows: Among them, π * (a|s) represents the optimal strategy, s is the state of user portrait and industry label, a is the action of clipping template, r(s t ,a t ) is the reward function and γ is the discount factor.

5. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The interactive editing engine includes a text-video mapping unit and a synchronization adjustment unit. The text-video mapping unit is used to establish a bidirectional mapping relationship between text and video, and the synchronization adjustment unit is used to automatically adjust the video content according to the text modification. The bidirectional mapping between text and video can be expressed as follows: V new =M·V old +b Among them, V new is the adjusted video feature matrix, V old is the original video feature matrix, M is the mapping matrix, and b is the offset vector.

6. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The multi-language processing module includes a subtitle generation unit and a localization adaptation unit. The subtitle generation unit is used to generate subtitles in multiple languages, and the localization adaptation unit is used to adapt the localization interface.

7. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The AI ​​generation engine includes a material calling unit and a content synthesis unit. The material calling unit is used to call video clips, pictures and music from the material library. The content synthesis unit is used to generate a first version of the video based on the storyboard script and dynamic template, and add automatic dubbing and subtitles.

8. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The finished product output module includes a video export unit and a cloud platform synchronization unit. The video export unit is used to export the final video product, and the cloud platform synchronization unit is used to synchronize the video product to the cloud platform or generate a data report.

9. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The AI ​​generation engine further includes a deep learning-based content optimization unit, which uses the user's historical editing data and preferences to train a personalized video content generation model to adjust the video content generation strategy based on the user's editing history and preferences. The optimization process of the content optimization unit can be expressed as follows: Where: V opt Represents the optimized video content feature vector; represents the initially generated video content feature vector; N is the number of user editing histories; T i is the feature vector of the i-th user’s editing history; ω i is the weight of the i-th editing history, which is determined by the user's satisfaction with the editing result; Loss i is the loss function of the i-th editing history, which is used to measure the difference between the video content and the user preference; Represents the gradient of the video content feature vector; α is the learning rate, which controls the optimization step size.

10. The AI ​​video real-time editing system based on natural language interaction according to claim 1, characterized in that: The multilingual processing module further includes a context-aware subtitle generation unit, which utilizes natural language processing technology and machine translation technology to optimize subtitle generation to improve the accuracy and naturalness of subtitles. The analysis process of the context-aware subtitle generation unit can be expressed as follows: C ctx =Transformer(V,C,S) Where: C ctx represents the subtitle feature vector after considering the context; Represents the video content feature vector; C represents the initially generated subtitle feature vector; S represents the context feature vector of the video content, including scene description, character emotion and other information; Transformer is a function based on the Transformer model, which is used to generate more accurate subtitles based on video content and context.