Content creation editing system
Through the multimodal editing and dynamic interface modules of the content creation and editing system, combined with artificial intelligence models, the integration difficulties and format mismatch problems in multimodal content creation and publishing are solved, and efficient multi-channel content publishing and user-friendly creation experience are achieved.
Patent Information
- Application Number
- CN202510626347.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing technology, the process of creating and publishing multimodal content is difficult to integrate and inefficient, and cross-channel publishing is cumbersome and prone to format mismatches.
It adopts a content creation and editing system that integrates multimodal editing modules, dynamic interface modules and publishing modules, uses artificial intelligence models for multimodal content processing and automatic publishing, combines text, image and video processing options, and adjusts the interface layout based on user operating habits to achieve cross-platform content format adaptation and monitoring.
It achieves efficient integration and processing of multimodal content, improves creation efficiency, ensures the accuracy and consistency of content publishing, meets user needs, and improves the efficiency and quality of cross-channel publishing.
Smart Images

Figure CN120602728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of content generation, and in particular to a content creation and editing system. Background Art
[0002] Currently, content creation and publishing often face the following challenges: Multimodal content integration is difficult. Traditional tools require separate processing of text, images, and video data within different software programs, fragmenting the entire creative process and leading to inefficiencies. Furthermore, cross-channel publishing is cumbersome and prone to mismatches between content formats and platforms. Summary of the Invention
[0003] In response to the deficiencies in the prior art, the present invention provides a content creation and editing system that can effectively integrate and process multimodal content, automatically match the content requirements of the publishing platform, and improve creation efficiency.
[0004] The present application provides a content creation and editing system, which is associated with an artificial intelligence model and includes:
[0005] A multimodal editing module, wherein the multimodal editing module provides at least a text processing option, an image processing option, and a video processing option;
[0006] A dynamic interface module, which adjusts the layout of the multimodal editing module based on the user's operating habits;
[0007] A publishing module is connected to multiple publishing platforms. The publishing module automatically identifies the specifications of the publishing platforms, publishes the created content to the multiple publishing platforms, and monitors its publishing status.
[0008] In one aspect, the text processing option receives input text and automatically generates matching images and / or videos based on the artificial intelligence model;
[0009] The image processing option receives the selected background image and generates a background image that meets the text requirements based on the artificial intelligence model;
[0010] The video processing option is used to select a digital human image and generate a digital human broadcast video based on the artificial intelligence model according to the input dialogue text.
[0011] In one aspect, the multimodal editing module further comprises:
[0012] A dialogue generation option, wherein the dialogue generation option receives input keywords and generates dialogue text based on the artificial intelligence model;
[0013] Speech synthesis options, which are used to adjust speech pauses, speaking speed, and intonation;
[0014] A storyboard adding option is used to receive product images uploaded by users, embed the product images into storyboards, and support the replacement of background images of product images.
[0015] In one aspect, the editing system further comprises:
[0016] Line adjustment options, which are based on modifying the line text and updating the digital human's broadcast content and speech synthesis in real time;
[0017] Storyboard optimization options, which adjust the position, size, and background of the product image in the storyboard.
[0018] In one aspect, the editing system further comprises:
[0019] A partial redraw option is provided, wherein the partial redraw option is based on a partial area in the video selected by the user to redraw the partial area.
[0020] In one aspect, the content creation and editing system further includes: a content collaboration module, which collaborates to align the matching between semantics and visual features.
[0021] In one aspect, the content creation and editing system includes: a preview window, based on which the editing video effect is previewed in real time.
[0022] In one aspect, the content creation and editing system further includes: a feedback module, which provides operational feedback based on an artificial intelligence model.
[0023] In one aspect, the multimodal editing module also includes an intelligent optimization option, which analyzes multimodal content and automatically recommends optimization solutions.
[0024] In one aspect, the content creation and editing system also includes a learning module, which collects user behavior data, analyzes and optimizes interaction processes and function recommendations.
[0025] The beneficial effects of the present invention are as follows: the content creation and editing system includes a multimodal editing module, and the multimodal editing module integrates text processing options, image processing options, and video processing options. Therefore, it can be seen that at least three functions of text processing, image processing, and video processing can be completed through the same content creation and editing system. This completes the integration of multimodal content processing and automatically matches the content requirements of the publishing platform through the publishing module. In addition, the dynamic interface module can adjust the layout of the multimodal editing module based on the user's operating habits, thereby making the use of the content creation and editing system more in line with the user's needs and improving creation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.
[0027] Figure 1 A schematic diagram of the prediction process of the content creation and editing system of this application;
[0028] Figure 2 A schematic diagram of the interface of the editing system for content creation of this application;
[0029] Figure 3 This is a schematic diagram of the interface for text processing options and image processing options in the multimodal editing module of the content creation and editing system of this application;
[0030] Figure 4 Create and edit the content for this application Figure 3 Schematic diagram of selecting product images;
[0031] Figure 5 Create and edit the content for this application Figure 4 Schematic diagram of selecting a theme style for product images;
[0032] Figure 6 Create and edit the content for this application Figure 5 Schematic diagram of generating promotional pictures from product pictures;
[0033] Figure 7 Create and edit the content for this application Figure 6 Schematic diagram of the product image before the background image is changed;
[0034] Figure 8 Create and edit the content for this application Figure 6 Schematic diagram of the product image after the background image is replaced;
[0035] Figure 9A schematic diagram of the channel publishing in the content creation and editing system of this application;
[0036] Figure 10 A schematic diagram of monitoring the publishing status in the content creation and editing system for this application;
[0037] Figure 11 A schematic diagram of adding options for storyboards in the content creation and editing system of this application;
[0038] Figure 12 Schematic diagram of selecting a digital human in the content creation and editing system for this application;
[0039] Figure 13 This is a schematic diagram of background generation on a product image in the content creation and editing system of this application.
[0040] Description of the drawings: 10. Multimodal editing module; 20. Dynamic interface module; 30. Publishing module; 40. Content collaboration module; 50. Preview window; 60. Feedback module; 70. Learning module;
[0041] 100. Text processing options; 110. Image processing options; 120. Video processing options; 130. Dialogue generation options; 140. Speech synthesis options; 150. Storyboard addition options; 160. Dialogue adjustment options; 170. Storyboard optimization options; 180. Partial redrawing options; 190. Intelligent optimization options; 101. Artificial intelligence model; 310. Publishing platform. DETAILED DESCRIPTION
[0042] The following embodiments of the technical solution of the present invention will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore only examples and are not intended to limit the scope of protection of the present invention.
[0043] It should be noted that, unless otherwise specified, the technical or scientific terms used in this application should have the common meanings understood by those skilled in the art to which the present invention belongs.
[0044] like Figures 1 to 13 As shown, the present application provides a content creation and editing system, which is associated with an artificial intelligence model 101. The content creation and editing system in this application can be connected to the artificial intelligence model 101, such as Deepseek or Manus. Alternatively, the artificial intelligence model 101 can be embedded in the content creation and editing system, and through iterative training of the artificial intelligence model 101, it can be applied to content creation. The content creation and editing system includes: a multimodal editing module 10, a dynamic interface module 20, and a publishing module 30.
[0045] The multimodal editing module 10 provides at least a text processing option 100, an image processing option 110, and a video processing option 120. The text processing option 100 can be a text input box, and by entering content into the text input box, the desired image, voice, or video can be generated. The image processing option 110 allows users to upload an image or select from the image library to complete image processing. The video processing option 120 allows users to perform video processing.
[0046] The dynamic interface module 20 adjusts the layout of the multimodal editing module 10 based on the user's operating habits; for example, it displays the tools commonly used by the user at the top of the interface, or calls out the commonly used options corresponding to the user role.
[0047] The publishing module 30 connects to multiple publishing platforms 310. It automatically identifies the specifications of the publishing platforms 310, publishes the created content to the multiple publishing platforms 310, and monitors its publishing status. Examples include Weibo, WeChat, and short video platforms such as Xiaohongshu. The publishing module 30 can also receive feedback from the publishing platforms 310 and monitor the content on the publishing platforms 310 in real time.
[0048] To give a further example, for the above solution, after the user enters the text in the text processing option 100, the system automatically generates matching visual content (such as images or videos) based on the artificial intelligence model 101, and supports real-time editing. The dynamic interface module 20 dynamically adjusts the interface layout by analyzing the user's historical operation data (such as the frequency of commonly used tools and the order of editing), such as placing the frequently used "storyboard addition" function at the top, or displaying customized toolbars based on the user's role (such as designer, marketer). The publishing module 30 connects to multiple platforms (such as WeChat, Tik Tok, and Xiaohongshu) through a preset API interface, automatically identifies the content specifications of each platform (such as video resolution, image-text ratio), completes format adaptation and batch publishing, and monitors the publishing status in real time (such as success / failure prompts, pageview statistics) to ensure the efficiency and accuracy of cross-channel distribution.
[0049] In this embodiment, the content creation and editing system includes a multimodal editing module 10, which integrates text processing options 100, image processing options 110, and video processing options 120. Therefore, the same content creation and editing system can perform at least three functions: text processing, image processing, and video processing. This completes the integration of multimodal content processing, and the publishing module 30 automatically matches the content requirements of the publishing platform 310. Furthermore, the dynamic interface module 20 can adjust the layout of the multimodal editing module 10 based on the user's operating habits, thereby making the content creation and editing system more in line with the user's needs and improving creation efficiency.
[0050] In one embodiment of the present application, the text processing option 100 receives the input text and automatically generates matching images and / or videos based on the artificial intelligence model 101; the image processing option 110 receives the selected background image and generates a background picture that meets the text requirements based on the artificial intelligence model 101; for example, if you need to recommend a product, you can select the product image and background image separately, and then use the artificial intelligence model 101 to superimpose the product image on the background image to form the image that needs to be promoted.
[0051] The video processing option 120 is used to select a digital human image and generate a digital human broadcast video based on the artificial intelligence model 101 according to the input dialogue text. It can be to generate images or videos, or to generate images and videos at the same time. For example, if you enter the copy of "Spring New Product Launch", the system will automatically generate a spring scene background image and recommend dynamic text animation effects. In the image processing option 110, after the user uploads the product image, the artificial intelligence model 101 combines the semantics of the copy (such as "technological sense") to intelligently generate a matching virtual background (such as a futuristic city scene) and supports background replacement. The video processing option 120 allows the user to select a digital human image (such as a virtual anchor), enter the dialogue text, and the artificial intelligence model 101 drives the digital human to perform voice broadcast and generate a lip-synced broadcast video. The user can preview the effect in real time and adjust the video rhythm by dragging the timeline to achieve seamless integration of multimodal content.
[0052] In one embodiment of the present application, the multimodal editing module 10 further includes: a dialogue generation option 130 , a speech synthesis option 140 and a storyboard adding option 150 .
[0053] The dialogue generation option 130 receives input keywords and generates dialogue text based on the AI model 101. The speech synthesis option 140 adjusts the pauses, speaking speed, and intonation of the speech. The storyboard addition option 150 receives user-uploaded product images and embeds them into the storyboard, allowing users to change the background image of the product image. In the dialogue generation option 130, users enter keywords (such as "sale," "limited-time discount"), and the AI model 101 automatically generates dialogue text that suits the marketing scenario based on natural language generation technology (such as GPT-4), providing multiple versions for users to choose from. The speech synthesis option 140 allows users to adjust speech parameters such as speaking speed and intonation, and generates audition audio in real time to ensure that it matches the rhythm of the video. In the storyboard addition option 150, after users upload product images, the system uses image segmentation technology to extract the main product and embed it into a preset storyboard template (such as a product close-up or a scene display). Users can also customize the background (such as replacing it with a dynamic starry sky or a real-life scene), enhancing the flexibility and creativity of content production.
[0054] In one embodiment of the present application, the editing system further includes: a dialogue adjustment option 160 and a storyboard optimization option 170 .
[0055] The dialogue adjustment option 160 updates the digital human's broadcast content and voice synthesis in real time based on the modified dialogue text; the storyboard optimization option 170 adjusts the position, size and background of the product image in the storyboard. The editing system provides dialogue adjustment and storyboard optimization functions. When the user modifies the dialogue text, the system updates the digital human's broadcast content in real time and synchronously adjusts the voice synthesis's intonation and pauses to ensure consistency between sound and picture. For example, if the user changes the dialogue from "limited to 3 days" to "limited to 24 hours", the digital human's voice type and voice rhythm will automatically adapt. In the storyboard optimization option 170, users can adjust the position and size of the product in the storyboard by dragging and dropping, and preview the effects of different backgrounds in real time (such as switching from indoor to outdoor scenes). The system also provides intelligent alignment auxiliary lines to ensure beautiful composition.
[0056] In one embodiment of the present application, the editing system also includes: a local redraw option 180, which redraws the local area based on the user's selection of the local area in the video. The local redraw function allows the user to perform fine-grained editing of specific areas in the video. For example, the user selects a background somewhere in the video (such as the sky area), and the system generates a variety of replacement options (such as sunset glow, starry sky) through the image restoration algorithm. After the user selects, the system automatically merges the new background with the original video to ensure a natural transition. This function can also support manual drawing. The user can accurately define the area to be modified and optimize the details by adjusting the brush size and transparency, greatly improving the flexibility and professionalism of video production.
[0057] In one embodiment of the present application, the content creation and editing system also includes: a content collaboration module 40, which coordinates the matching between semantics and visual features. In addition, the artificial intelligence model 101 will also detect the visual coordination between the product and the background through the content collaboration module 40. If a mismatch is found, an optimization suggestion will pop up. For example, if a modern product is placed on a retro background, a recommended minimalist style background will pop up. The content collaboration module 40 can also ensure the semantic consistency of text, images, and videos through multimodal semantic alignment technology. For example, if the text entered by the user is "environmentally friendly materials", the system will automatically detect whether the picture contains green elements (such as plants, recyclable signs), and if there is no match, it will recommend replacing it with a related material library picture. At the same time, if the digital human lines in the video mention "lightweight design", the system will synchronously highlight the corresponding product details in the storyboard (such as a close-up of the thin and light body) to achieve cross-modal content collaboration and avoid information fragmentation.
[0058] In one embodiment of the present application, the content creation and editing system includes: a preview window 50, based on which the editing video effect is previewed in real time. The preview window 50 supports real-time rendering of the editing effect. When the user adjusts the text, image or video parameters, the preview window 50 immediately displays the final effect. For video editing, the user can view the animation transition effect frame by frame and locate the keyframes through the timeline zoom function. The preview window 50 also provides a multi-view mode (such as split-screen comparison of the original and optimized versions) to help users quickly evaluate the editing effect and reduce the time cost of repeated export and viewing.
[0059] In one embodiment of the present application, the content creation and editing system further includes a feedback module 60 that provides operational feedback based on the artificial intelligence model 101. When a user frequently adjusts the speech rate of the same line, the system prompts: "Multiple modifications detected. Do you want to enable automatic speech rate optimization?" If the user agrees, the artificial intelligence model 101 automatically generates recommended parameters based on the video's tempo.
[0060] In one embodiment of the present application, the multimodal editing module 10 also includes an intelligent optimization option 190, which performs multimodal content analysis and automatically recommends optimization solutions. The intelligent optimization option 190 automatically generates an optimization solution through multimodal content analysis (such as text emotion recognition, image composition scoring, and video fluency detection). For example, when the system detects that the video transition is abrupt, it recommends adding a gradient effect; if the text conflicts with the background style, it recommends changing the color scheme or font style. After the user clicks "One-click Optimization", the system automatically applies the best solution and retains the authority to make manual adjustments to achieve a balance between efficiency and creativity.
[0061] In one embodiment of the present application, the content creation and editing system also includes a learning module 70, which collects user behavior data, analyzes and optimizes interaction processes and function recommendations. The learning module 70 continuously collects user behavior data (such as function usage frequency, editing preferences), and optimizes system interactions through machine learning algorithms. For example, if most users tend to manually adjust the storyboard position rather than use automatic alignment, the system will lower the recommendation priority of the alignment function and enhance the response speed of manual operations. At the same time, the module will dynamically update the function library (such as the new "live slice generation" tool) based on the common needs of the user group (such as e-commerce users need to quickly generate product videos) to achieve self-evolution and personalized adaptation of the system.
[0062] The specific embodiments and beneficial effects of the deep learning-based cell data integration method in this application can be found in the above-mentioned content creation and editing system, which will not be repeated here.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.
Claims
1. A content creation and editing system, characterized in that: The content creation and editing system is associated with an artificial intelligence model, and the content creation and editing system includes: A multimodal editing module, wherein the multimodal editing module provides at least a text processing option, an image processing option, and a video processing option; A dynamic interface module, which adjusts the layout of the multimodal editing module based on the user's operating habits; A publishing module is connected to multiple publishing platforms. The publishing module automatically identifies the specifications of the publishing platforms, publishes the created content to the multiple publishing platforms, and monitors its publishing status.
2. The content creation and editing system according to claim 1, characterized in that: The text processing option receives input text and automatically generates matching images and / or videos based on the artificial intelligence model; The image processing option receives the selected background image and generates a background image that meets the text requirements based on the artificial intelligence model; The video processing option is used to select a digital human image and generate a digital human broadcast video based on the artificial intelligence model according to the input dialogue text.
3. The content creation and editing system according to claim 2, characterized in that: The multimodal editing module further includes: A dialogue generation option, wherein the dialogue generation option receives input keywords and generates dialogue text based on the artificial intelligence model; Speech synthesis options, which are used to adjust speech pauses, speaking speed, and intonation; A storyboard adding option is used to receive product images uploaded by users, embed the product images into storyboards, and support the replacement of background images of product images.
4. The content creation and editing system according to claim 3, characterized in that: The editing system further includes: Line adjustment options, which are based on modifying the line text and updating the digital human's broadcast content and speech synthesis in real time; Storyboard optimization options, which adjust the position, size, and background of the product image in the storyboard.
5. The content creation and editing system according to claim 3, characterized in that: The editing system further includes: A partial redraw option is provided, wherein the partial redraw option is based on a partial area in the video selected by the user to redraw the partial area.
6. The content creation and editing system according to claim 1, characterized in that: The content creation and editing system further includes a content collaboration module, which collaborates to align the matching between semantics and visual features.
7. The content creation and editing system according to claim 1, characterized in that: The content creation and editing system includes: a preview window, based on which the editing effect is previewed in real time.
8. The content creation and editing system according to claim 1, characterized in that: The content creation and editing system also includes: a feedback module, which provides operation feedback based on an artificial intelligence model.
9. The content creation and editing system according to claim 1, characterized in that: The multimodal editing module also includes an intelligent optimization option, which analyzes multimodal content and automatically recommends optimization solutions.
10. The content creation and editing system according to claim 1, characterized in that: The content creation and editing system also includes a learning module, which collects user behavior data, analyzes and optimizes interaction processes and function recommendations.